A target detection method based on frame image and event stream feature fusion

CN120953942BActive Publication Date: 2026-09-08TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511058260.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2026-09-08
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

自动驾驶车辆在行驶过程中需要保持一定速度,以满足实际交通需求,但这也带来了图像模糊和抖动的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953942B_ABST
    Figure CN120953942B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of image processing, and discloses a target detection method based on frame image and event stream feature fusion, which comprises the following steps: collecting RGB images and event stream data, and constructing a multi-modal data pair with time and space synchronization; constructing a double-flow feature extraction backbone network to extract event and RGB features; based on a cross attention mechanism, using multi-head attention to realize semantic alignment; according to statistical distribution, adaptively adjusting the fusion weight; constructing a multi-level fusion network, using an FPN pyramid for cross-scale fusion, and through a spatial semantic aggregation module, synergistically strengthening target positioning and classification feature expression. The application uses the above-mentioned target detection method based on frame image and event stream feature fusion, through a multi-modal data collaborative perception and adaptive feature optimization mechanism, significantly improves the detection robustness in complex scenes, effectively reduces the dynamic target missed detection rate and misidentification rate, and provides an all-weather high-precision environmental perception capability for an automatic driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a target detection method based on the fusion of frame images and event stream features. Background Technology

[0002] Target detection in autonomous driving is a hot research topic in computer vision. Its core objective is to use deep learning technology and various sensors to comprehensively detect and identify targets such as vehicles, pedestrians, and traffic signs in road scenes, ensuring the safe and stable operation of autonomous driving systems. Real-time perception of the surrounding environment is essential for ensuring the safe operation of autonomous vehicles. One of the main tasks of this perception work is to identify and detect various obstacles and traffic participants on the road, thereby preventing traffic accidents caused by identification errors. Therefore, target detection in autonomous driving scenarios is of great significance for ensuring driving safety and reliability.

[0003] Currently, most autonomous driving target detection relies on traditional vision sensors such as onboard cameras for environmental perception. This method has certain advantages, such as wide coverage, relatively low cost, and the ability to provide rich scene information. However, traditional vision sensors also face many challenges in practical applications. First, due to the limitations of the working principle and technical characteristics of traditional vision sensors, their performance is often unsatisfactory in complex lighting environments. For example, in strong light or low light conditions at night, the lighting changes drastically, potentially resulting in strong exposure and backlighting. In such situations, it is difficult for the sensor to capture a clear image of the road scene, leading to reduced detection accuracy.

[0004] Furthermore, the high-speed movement of vehicles is also a significant factor. Autonomous vehicles need to maintain a certain speed to meet real-world traffic demands, but this introduces problems such as image blurring and jitter. Especially in complex road conditions and inclement weather, the vehicle's motion constantly changes, further exacerbating image blurring and jitter. These combined issues mean that target detection models based on traditional vision sensors lack sufficient maturity and reliability in practical applications, severely impacting their practical value. Summary of the Invention

[0005] The purpose of this invention is to provide a target detection method based on the fusion of frame images and event stream features. Through multimodal data collaborative perception and adaptive feature optimization mechanism, it improves the detection robustness in complex scenarios, effectively reduces the false detection rate and false recognition rate of dynamic targets, and provides all-weather high-precision environmental perception capability for autonomous driving systems.

[0006] To achieve the above objectives, this invention provides a target detection method based on the fusion of frame images and event stream features, comprising the following steps:

[0007] Step S1: Acquire RGB image and event stream data, and use the field of view (FOV) spatial alignment algorithm to achieve spatial matching between the event stream and the RGB image, and construct a spatiotemporally synchronized multimodal data pair;

[0008] Step S2: Construct a dual-stream feature extraction backbone network to extract event features and RGB features respectively to obtain semantic features;

[0009] Step S3: Establish cross-modal feature association based on cross-attention mechanism, and achieve semantic alignment using multi-head attention; dynamically optimize fused features using mean-standard deviation method, and adaptively adjust fusion weights according to statistical distribution;

[0010] Step S4: Construct a multi-level fusion network, input the optimized features into the FPN pyramid for cross-scale fusion, and through the spatial semantic aggregation module, collaboratively enhance the target localization and classification feature representation;

[0011] Step S5: Train the detection network using a spatiotemporally aligned multimodal dataset and optimize the parameter update process by combining a dynamic hard sample mining strategy;

[0012] Step S6: Input the real-time aligned event stream and RGB image into the training network, and output the target detection box with confidence and motion parameters to provide environmental perception decision support for the autonomous driving system.

[0013] Preferably, in step S1, RGB images and event stream data are acquired, and spatiotemporally synchronized multimodal data pairs are constructed using a field of view (FOV) spatial alignment algorithm. The specific process is as follows:

[0014] Step S11: Acquire RGB images from a traditional vision sensor and event stream data from a neuromorphic vision sensor, respectively;

[0015] Step S12: Eliminate multi-camera field-of-view differences in the RGB image through geometric correction, and calculate the transformation matrix M between cameras. r1_r0 As shown below:

[0016]

[0017] Among them, K r1 K represents the intrinsic parameter matrix of the left camera; r0 R represents the intrinsic parameter matrix of the right camera; rect1 Represents the left camera correction rotation matrix; R rect0 Indicates the right camera correction rotation matrix; T 10 This represents the transformation matrix from camera 0 to camera 1;

[0018] Step S13: Implement frequency-adaptive representation for the event stream data and dynamically adjust the event accumulation time, as shown below:

[0019] Δt∈[10ms,50ms](2);

[0020] Where Δt represents the event accumulation time;

[0021] Step S14: For each coordinate point (u,v) in the source image, calculate the mathematical expression for its coordinate mapping table:

[0022] P r0 =[u,v,1] T (3);

[0023] P r1 =M r1_r0 ×P r0 (4);

[0024]

[0025] Among them, P r0 P represents the homogeneous coordinates in the source image; r1 P represents the homogeneous coordinates of the reprojected point; r1 [0] represents the unnormalized x-coordinate; P r1 [1] represents the unnormalized y-coordinate; P r1 [2] represents the scaling factor of homogeneous coordinates; mapx represents the x-coordinate mapping table used for remapping and mapy represents the y-coordinate mapping table used for remapping;

[0026] Step S15: Calculate the mathematical expression for image remapping using the coordinate mapping table:

[0027] I remap =Remap(I src ,mapx,mapy) (7);

[0028] Among them, I src Represents the source image; Remap represents the remapping process; I remap This represents the remapped image.

[0029] Preferably, in step S2, a dual-stream feature extraction backbone network is constructed to extract event features and RGB features respectively to obtain semantic features;

[0030] Among them, the RGB branch enhances the target contour representation through deformable convolution to obtain rich texture information; the event branch enhances cross-channel features through AFRM grouped convolution spatial attention to obtain clear contour features.

[0031] Preferably, the dual-stream feature extraction backbone network is divided into 4 layers. Each layer has the same structure except for the channel dimension. Feature extraction is achieved through residual operations. The specific process is as follows:

[0032] Step S21: A single residual unit consists of a 3×3 convolutional layer, batch normalized normalization (BN), and ReLU activation; therefore, the output feature map F of a single residual unit... out As shown below:

[0033] F out =ReLU(BN(Conv) 3*3 (ReLU(BN(Conv 1*1 (F in ))))+F in (8);

[0034] Among them, F in For input features; Conv 1*1 Represents a 1×1 convolution operation; BN is batch normalization; ReLU is the activation function for residual units; Conv 3*3 This represents a 3×3 convolution operation;

[0035] Step S22: After standardization preprocessing, the RGB image is input into the dual-stream feature extraction backbone network. Multi-scale features are extracted using the ResNet-50 architecture, with its first-layer 7×7 convolutional kernel expanding the spatial receptive field to achieve initial feature modeling and obtain the initial RGB features F. f ;

[0036] Step S23: The event branch enhances cross-channel features through AFRM grouped convolution spatial attention to obtain clear contour features;

[0037] First, the event frame sequence is tensor-quantized and then input into the dual-stream feature extraction backbone network. The first layer uses a 5×5 convolutional kernel to enhance the modeling capability of spatiotemporal event distribution, and its output channels are 64.

[0038] Subsequently, hierarchical feature extraction is performed through a four-level residual module, which gradually reduces the spatial resolution of the feature map to 1 / 16 of the original size.

[0039] Preferably, in the event branch, an AFRM module is embedded after each residual unit to enhance the event feature representation capability. The specific process is as follows:

[0040] Step S231: Divide the channel dimension into 8 groups using grouped convolution, and perform spatial attention calculation on the individual features after grouping, as shown below:

[0041] F h =PoolH(F group (9);

[0042] F w =PoolW(F group ) T (10);

[0043] HW = Conv 1*1 ([F h ,F w ]) (11);

[0044] Among them, F h F represents the feature map after height-oriented pooling; w The first part represents the feature map after pooling in the width direction; PoolH represents adaptive average pooling in the height direction; PoolW represents adaptive average pooling in the width direction; F group The original grouped features are represented by T; T represents the transpose operation; HW represents the merged feature map; [F h ,F w ] represents connecting F along the channel dimension h ,F w ;

[0045] Step S232: The spliced ​​tensor can be divided into F using the Split operation. h ′,F w And the global features are obtained using the Sigmoid activation function, as shown below:

[0046] F h ′,F w ′=Split(HW,[h,w],dim=2) (12);

[0047]

[0048] Among them, F h ′,F w ′ represents the height attention feature and width attention feature after segmentation, respectively; F1 represents the global feature output of the spatial attention branch; h represents the height dimension; w represents the width dimension; dim represents the dimension; GN represents the group normalization operation; σ represents the Sigmoid activation function;

[0049] Step S233: Extract local features through 3×3 convolution, multiply them with the obtained global features F1 by cross-pixel multiplication, and obtain the weights, as shown below:

[0050] F2 = Conv 3*3 (F group (14);

[0051] F 11=Softmax(AvgPool(F1)) T (15);

[0052] F 21 =Softmax(AvgPool(F2)) T (16);

[0053] F 12 =reshape(F2),F 22 =reshape(F1) (17);

[0054] Weights=(F 11 ×F 12 +F 21 ×F 22 (18);

[0055] Among them, F 11 Represents global features; F 21 Indicates local features; F 12 and F 22 These represent the reshaped global and local features, respectively; `reshape` indicates changing the shape of the parameters; `AvgPool` represents global average pooling; `Softmax` is the normalization function; and `Weights` represents the weights calculated by cross-attention.

[0056] Step S234: Multiply the original grouped features by the Weights to obtain the initial event features F. e As shown below:

[0057] F e =(F group ×σ(Weights)) (19);

[0058] Where σ represents the Sigmoid activation function.

[0059] Preferably, in step S3, the initial feature F of the event is... e With RGB initial feature F f Convolution processing is performed, and global attention is calculated through element-wise multiplication to fuse bimodal features. A spatial-channel hybrid attention mechanism is then used to enhance feature representation, as shown below:

[0060]

[0061] in, and These represent the RGB image features and event features after the convolution operation, respectively; Conv represents the convolution operation. and These represent the RGB enhancement features and event enhancement features after initial fusion, respectively;

[0062] Step S32: Linearly project the fused and enhanced RGB features and event features onto the query, key, and value respectively, as shown below:

[0063]

[0064] Among them, Q f Q represents the projection of the enhanced RGB features onto the query; e K represents the projection of the enhanced event features onto the query; f K represents the projection of the enhanced RGB features onto the key. e V represents the projection of the enhanced event features onto the key; f V represents the projection of the enhanced RGB features onto the value; e This represents the projection of the enhanced event features onto the value; and These represent the learnable linear projections of RGB and event features in the query, respectively; and These represent the learnable linear projections of RGB and event features in the key, respectively; and These represent the learnable linear projections of RGB and event features in the value, respectively;

[0065] Step S33: By calculating the dual-branch cross-attention mechanism, the multimodal features are fused as follows:

[0066]

[0067] in, and These represent the results of RGB-dominated fusion and event-dominated fusion, respectively; D represents the channel dimension of the feature map; Softmax represents the normalization function.

[0068] Step S34: Analyze the fusion results. The matching style alignment refines the fusion process, employs the mean-standard deviation method to dynamically optimize the fusion features, and utilizes Concat to connect the alignment features, yielding the output result F. output As shown below:

[0069]

[0070] in, and These represent the RGB features and event features after style alignment refinement and fusion, respectively. and The mean of the content features; and The standard deviation representing content characteristics; and The mean of the style characteristics; and The standard deviation represents the style feature; Concat represents the join alignment feature operation;

[0071]

[0072] Where, μ x This represents the formula for calculating the mean; σ x The formula for calculating the standard deviation is given; H represents the height of the feature map; W represents the width of the feature map; ∈ is a variable between 10 and 10. -8 Up to 10 -3 Positive numbers between these ranges are used to avoid division by zero errors.

[0073] Preferably, in step S4, a multi-level fusion network is constructed, and the optimized features are input into the FPN pyramid for cross-scale fusion. The spatial semantic aggregation module is used to collaboratively enhance the target localization and classification feature expression.

[0074] The spatial semantic aggregation module constructs a pyramid structure P2-P6 containing rich semantic and spatial information by fusing multi-scale feature maps of the backbone network. First, 1×1 convolution is used to unify all input feature channels to 256 dimensions. Then, 3×3 convolution feature integration is performed on the base layers P2-P5, and the P6 layer is specially extended to adapt to the needs of large-scale target detection.

[0075] Among them, features P3-P5 are spatially adapted using the nearest neighbor interpolation method, ultimately forming a five-level scale optimized feature pyramid P2 / P3 / P4 / P5 / P6.

[0076] Preferably, the P5 feature is generated as follows:

[0077] P5 temp =Conv 1*1 (C5) (31);

[0078] P5 upsampled =Upsample(P5) temp (32);

[0079] P5 = Conv 3×3 (P5 temp (33);

[0080] Where C5 represents the fusion result of layer 4 in the backbone network; P5 tempP5 represents the unified result of features; Upsample represents nearest neighbor upsampling; P5 upsampled P5 represents the upsampling result; P5 represents the P5 feature representation after convolution enhancement.

[0081] P4 and P3 features are generated as follows:

[0082] Pi temp =Conv 1×1 (Ci) (34);

[0083] Pi fused =Pi temp +Pi+1 upsampled (35);

[0084] Pi upsampled =Upsample(Pi temp (36);

[0085] Pi = Conv 3×3 (Pi temp (37);

[0086] Where i takes values ​​in the range [3,4]; Ci represents the fusion result of the (i-1)th layer in the backbone network; Pi temp Indicates the result of feature unification; Pi+1 upsampled Pi represents the upsampling result from the previous layer. fused Indicates the feature fusion result; Pi upsampled Pi represents the upsampling result; Pi represents the feature representation after convolution enhancement.

[0087] In the feature pyramid structure, the generation mechanism of P2 is similar to that of P3 / P4, but its bottom feature C2 is located in the shallowest layer of the network, so there is no need to upsample and connect to the next level of features.

[0088] Layer P6 directly processes the C5 feature map through 3×3 convolutions, generating a larger-scale feature representation, as shown below:

[0089] P6 = Conv 3×3,stride=2 (C5) (38).

[0090] Preferably, in step S5, the detection network is trained using a spatiotemporally aligned multimodal dataset, and the parameter update process is optimized by combining a dynamic hard sample mining strategy. The specific process is as follows:

[0091] First, a neuromorphic vision sensor, namely an event camera, is introduced as the key data source. By constructing a dual-branch feature extraction network for event streams and RGB data, event features are used to capture dynamic target edge information, and RGB features are used to extract texture semantic information, thereby achieving cross-modal feature complementarity.

[0092] Secondly, a novel unilateral event feature enhancement scheme is designed, which enhances the target edge information by using an asymmetric feature enhancement module to strengthen the event features extracted by the dual-branch backbone network. Then, a novel dynamic dual-branch feature fusion module is designed to perform pixel-level pre-fusion of the enhanced event feature map and the RGB feature map, and to deeply fuse multimodal features through a dual-branch cross-attention module, and to perform style alignment operation on the dual-branch fused features.

[0093] Then, after the fused features are enhanced by the attention module, they are input into the decoupled detection head and output the target classification probability, bounding box coordinates, and motion state vector simultaneously.

[0094] Finally, accuracy, recall, and mAP parameters are generated during training to evaluate the model's performance; when the loss function during training becomes smooth, the training of the autonomous driving target detection network based on multimodal feature fusion is complete.

[0095] Therefore, this invention adopts the above-mentioned target detection method based on the fusion of frame images and event stream features. It introduces the event stream collected by a high dynamic range, low latency neuromorphic vision sensor as one of the key data sources. Based on the multimodal fusion technology of event features and RGB image features, environmental perception and target detection are performed. This enables the trained perception model to maintain excellent performance under various complex lighting, weather and motion conditions, significantly improves the adaptability of the autonomous driving system to dynamic environmental changes, effectively reduces false detection, missed detection and perception delay, thereby greatly improving the driving safety guarantee coefficient and providing solid technical support for the application of advanced autonomous driving technology.

[0096] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0097] Figure 1 This is a schematic diagram of an autonomous driving target detection structure based on multimodal feature fusion according to an embodiment of the present invention;

[0098] Figure 2 This is a schematic diagram of the dual-branch feature extraction structure according to an embodiment of the present invention;

[0099] Figure 3 This is a schematic diagram of the grouped convolutional cross-channel feature enhancement structure according to an embodiment of the present invention;

[0100] Figure 4 This is a schematic diagram of the adaptive feature fusion structure according to an embodiment of the present invention. Detailed Implementation

[0101] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0102] like Figure 1 As shown, the present invention provides a target detection method based on the fusion of frame images and event stream features, comprising the following steps:

[0103] Step S1: Acquire RGB image and event stream data, and use the field of view (FOV) spatial alignment algorithm to achieve spatial matching between the event stream and the RGB image, and construct a spatiotemporally synchronized multimodal data pair;

[0104] Step S2: Construct a dual-stream feature extraction backbone network to extract event features and RGB features respectively to obtain semantic features;

[0105] Step S3: Establish cross-modal feature association based on cross-attention mechanism, and achieve semantic alignment using multi-head attention; dynamically optimize fused features using mean-standard deviation method, and adaptively adjust fusion weights according to statistical distribution;

[0106] Step S4: Construct a multi-level fusion network, input the optimized features into the FPN pyramid for cross-scale fusion, and strengthen the target localization and classification feature expression through the spatial semantic aggregation module.

[0107] Step S5: Train the detection network using a spatiotemporally aligned multimodal dataset and optimize the parameter update process by combining a dynamic hard sample mining strategy;

[0108] Step S6: Input the real-time aligned event stream and RGB image into the training network, and output the target detection box with confidence and motion parameters to provide environmental perception decision support for the autonomous driving system.

[0109] Example

[0110] Step S1: Acquire RGB image and event stream data, and use the field of view (FOV) spatial alignment algorithm to achieve spatial matching between the event stream and the RGB image, and construct a spatiotemporally synchronized multimodal data pair.

[0111] Traditional camera lenses suffer from radial and tangential distortion, leading to curved or disproportionate image edges. In uncorrected images, the position, shape, and size of target objects are shifted due to distortion. Therefore, a distortion model is established through geometric correction, mapping the original image to an ideal imaging plane to restore the geometry of the real scene. This invention utilizes a field-of-view (FOV) spatial alignment algorithm to obtain spatiotemporal alignment experimental data, sourced from the publicly available datasets DSEC and PKU-DDD17-CAR.

[0112] Step S11: Acquire RGB images from a traditional vision sensor and event stream data from a neuromorphic vision sensor, respectively.

[0113] Step S12: Eliminate multi-camera field-of-view differences in the RGB image through geometric correction, and calculate the transformation matrix M between cameras. r1_r0 As shown below:

[0114]

[0115] Among them, K r1 K represents the intrinsic parameter matrix of the left camera; r0 R represents the intrinsic parameter matrix of the right camera; rect1 Represents the left camera correction rotation matrix; R rect0 Indicates the right camera correction rotation matrix; T 10 This represents the transformation matrix from camera 0 to camera 1.

[0116] Step S13: Implement frequency-adaptive representation for the event stream data and dynamically adjust the event accumulation time, as shown below:

[0117] Δt∈[10ms,50ms] (2).

[0118] Where Δt represents the event accumulation time.

[0119] Step S14: Since the neuromorphic event stream and the RGB image come from different sensors, differences in field of view (FOV) and imaging principles can lead to spatial misalignment. Therefore, unifying the coordinate system through an FOV alignment algorithm, so that the two types of data are matched under the same spatial reference, can provide a foundation for subsequent cross-modal fusion.

[0120] For each coordinate point (u, v) in the source image, calculate the mathematical expression for its coordinate mapping table:

[0121] P r0 =[u,v,1] T (3);

[0122] P r1 =M r1_r0 ×P r0 (4);

[0123]

[0124] Among them, P r0 P represents the homogeneous coordinates in the source image; r1 P represents the homogeneous coordinates of the reprojected point; r1 [0] represents the unnormalized x-coordinate; P r1 [1] represents the unnormalized y-coordinate; P r1 [2] represents the scaling factor of homogeneous coordinates; mapx represents the x-coordinate mapping table used for remapping and mapy represents the y-coordinate mapping table used for remapping.

[0125] Step S15: Calculate the mathematical expression for image remapping using the coordinate mapping table:

[0126] I remap =Remap(I src ,mapx,mapy) (7);

[0127] Among them, I src Represents the source image; Remap represents the remapping process; I remap This represents the remapped image.

[0128] Step S2: Construct a dual-stream feature extraction backbone network to extract event features and RGB features respectively to obtain semantic features.

[0129] To address the issue that single-sensor modalities (such as visible light cameras) are susceptible to interference from factors like lighting changes and motion blur in complex traffic scenarios, a dual-stream feature extraction backbone network based on events and RGB is constructed, referencing a dual-stream backbone network architecture. Figure 2 As shown, an inconsistent dual-branch structure design is adopted, extracting edge features from the event modality and texture features from the RGB modality respectively to obtain semantic features.

[0130] The RGB branch enhances target contour representation and obtains rich texture information through deformable convolution; the event branch enhances cross-channel features and obtains clear contour features through AFRM grouped convolutional spatial attention. The dual-stream feature extraction backbone improves the robustness of the target detection model in extreme scenes by fusing dynamic event streams from the event camera and static image data from the RGB camera.

[0131] Step S21: The dual-stream feature extraction backbone network is divided into 4 layers. Each layer has the same structure except for the number of channels. The feature extraction effect is improved through residual operations.

[0132] Each residual unit consists of a 3×3 convolutional layer, batch normalization (BN), and ReLU activation; therefore, the output feature map F of a single residual unit is... out As shown below:

[0133] F out =ReLU(BN(Conv) 3*3 (ReLU(BN(Conv 1*1 (F in ))))+F in (8);

[0134] Among them, F in For input features; Conv 1*1 Represents a 1×1 convolution operation; BN is batch normalization; ReLU is the activation function for residual units; Conv3*3 This represents a 3×3 convolution operation.

[0135] Step S22: The RGB branch enhances the target contour representation through deformable convolution to obtain rich texture information.

[0136] After standardization preprocessing, the RGB image is input into a dual-stream feature extraction backbone network. Multi-scale features are extracted using a ResNet-50 architecture, with its first-layer 7×7 convolutional kernel expanding the spatial receptive field to achieve initial feature modeling and obtain the initial RGB features F. f .

[0137] Step S23: The event branch enhances cross-channel features through AFRM grouped convolution spatial attention to obtain clear contour features.

[0138] First, the event frame sequence is tensor-quantized and then input into the dual-stream feature extraction backbone network. The first layer uses a 5×5 convolution kernel to enhance the modeling capability of spatiotemporal event distribution (64 output channels) and improve the capture efficiency of sparse pulse features.

[0139] Subsequently, hierarchical feature extraction is performed through a four-level residual module, which gradually reduces the spatial resolution of the feature map to 1 / 16 of the original size.

[0140] like Figure 3 As shown, in the event branch, an AFRM module is embedded after each residual unit to enhance the event feature representation capability. The specific process is as follows:

[0141] Step S231: Divide the channel dimension into 8 groups using grouped convolution, and perform spatial attention calculation on the individual features after grouping, as shown below:

[0142] F h =PoolH(F group (9);

[0143] F w =PoolW(F group ) T (10);

[0144] HW = Conv 1*1 ([F h ,F w ]) (11);

[0145] Among them, F h F represents the feature map after height-oriented pooling; w The first part represents the feature map after pooling in the width direction; PoolH represents adaptive average pooling in the height direction; PoolW represents adaptive average pooling in the width direction; F groupThe original grouped features are represented by T; T represents the transpose operation; HW represents the merged feature map; [F h ,F w ] represents connecting F along the channel dimension h ,F w .

[0146] Step S232: The spliced ​​tensor can be divided into F using the Split operation. h ′,F w And the global features are obtained using the Sigmoid activation function, as shown below:

[0147] F h ′,F w ′=Split(HW,[h,w],dim=2) (12);

[0148]

[0149] Among them, F h ′,F w ′ represents the height attention feature and width attention feature after segmentation, respectively; F1 represents the global feature output of the spatial attention branch; h represents the height dimension; w represents the width dimension; dim represents the dimension; GN represents the group normalization operation; σ represents the Sigmoid activation function.

[0150] Step S233: To better utilize global and local features, local features are extracted through 3×3 convolution, and then multiplied by the obtained global feature F1 at cross-pixel intervals to obtain the weights, as shown below:

[0151] F2=Conv3*3(F group (14);

[0152] F 11 =Softmax(AvgPool(F1)) T (15);

[0153] F 21 =Softmax(AvgPool(F2)) T (16);

[0154] F 12 =reshape(F2),F 22 =reshape(F1) (17);

[0155] Weights=(F 11 ×F 12 +F 21 ×F 22 (18);

[0156] Among them, F 11 Represents global features; F 21 Indicates local features; F 12 and F 22 These represent the reshaped global and local features, respectively; reshape indicates changing the shape of the parameters; AvgPool represents global average pooling; Softmax is the normalization function; and Weights represents the weights calculated by cross-attention.

[0157] Step S234: Multiply the original grouped features by the Weights to obtain the initial event features F. e As shown below:

[0158] F e =(F group ×σ(Weights)) (19);

[0159] Where σ represents the Sigmoid activation function.

[0160] Step S3: Establish cross-modal feature associations based on the cross-attention mechanism, and achieve semantic alignment using multi-head attention; dynamically optimize the fused features using the mean-standard deviation method, and adaptively adjust the fusion weights according to the statistical distribution, such as... Figure 4 As shown.

[0161] Step S31: Initial characteristics F of the event e With RGB initial feature F f Convolution processing is performed, and global attention is calculated through element-wise multiplication to fuse bimodal features. A spatial-channel hybrid attention mechanism is then used to enhance feature representation, as shown below:

[0162]

[0163] in, and These represent the RGB image features and event features after the convolution operation, respectively; Conv represents the convolution operation. and These represent the RGB enhancement features and event enhancement features after initial fusion, respectively.

[0164] Step S32: Linearly project the fused and enhanced RGB features and event features onto the query, key, and value respectively, as shown below:

[0165]

[0166] Among them, Q f Q represents the projection of the enhanced RGB features onto the query;e K represents the projection of the enhanced event features onto the query; f K represents the projection of the enhanced RGB features onto the key. e V represents the projection of the enhanced event features onto the key; f V represents the projection of the enhanced RGB features onto the value; e This represents the projection of the enhanced event features onto the value; and These represent the learnable linear projections of RGB and event features in the query, respectively; and These represent the learnable linear projections of RGB and event features in the key, respectively; and These represent the learnable linear projections of RGB and event features in the value, respectively.

[0167] Step S33: By calculating the dual-branch cross-attention mechanism, the multimodal features are fused as follows:

[0168]

[0169] in, and These represent the results of RGB-dominated fusion and event-dominated fusion, respectively; D represents the channel dimension of the feature map; and Softmax represents the normalization function.

[0170] Step S34: Analyze the fusion results. The matching style alignment refines the fusion process, employs the mean-standard deviation method to dynamically optimize the fusion features, and utilizes Concat to connect the alignment features, yielding the output result F. output As shown below:

[0171]

[0172] in, and These represent the RGB features and event features after style alignment refinement and fusion, respectively. and The mean of the content features; and The standard deviation representing content characteristics; and The mean of the style characteristics; and The standard deviation represents the style feature; Concat represents the concatenation alignment feature operation.

[0173]

[0174] Where, μ x This represents the formula for calculating the mean; σ x The formula for calculating the standard deviation is given; H represents the height of the feature map; W represents the width of the feature map; ∈ is a variable between 10 and 10. -8 Up to 10 -3 Positive numbers between these ranges are used to avoid division by zero errors.

[0175] Step S4: Construct a multi-level fusion network, input the optimized features into the FPN pyramid for cross-scale fusion, and strengthen the target localization and classification feature expression through the spatial semantic aggregation module.

[0176] Feature Pyramid (FPN) integrates multi-scale features in a hierarchical manner and employs a top-down semantic propagation and a bottom-up detail enhancement mechanism to uniformly solve the problems of missed detection of small targets at long distances and localization offset of large targets at close distances, thus significantly improving the system robustness of target detection in autonomous driving.

[0177] The spatial semantic aggregation module constructs a pyramid structure (P2-P6) containing rich semantic and spatial information by fusing multi-scale feature maps from the backbone network. First, 1×1 convolutions are used to unify all input feature channels to 256 dimensions; then, 3×3 convolutional feature integration is performed on the base layers (P2-P5), and the P6 layer is specially extended to adapt to the needs of large-scale object detection.

[0178] Among them, the P3-P5 features are spatially adapted using the nearest neighbor interpolation method, ultimately forming a five-level scale optimized feature pyramid (P2 / P3 / P4 / P5 / P6), which provides a feature representation basis for subsequent multi-granularity target detection that adapts to objects of different scales.

[0179] P5 feature generation is shown below:

[0180] P5 temp =Conv 1*1 (C5) (31);

[0181] P5 upsampled =Upsample(P5) temp (32);

[0182] P5 = Conv 3×3 (P5 temp (33);

[0183] Where C5 represents the fusion result of layer 4 in the backbone network; P5 temp P5 represents the unified result of features; Upsample represents nearest neighbor upsampling; P5 upsampled P1 represents the upsampling result; P5 represents the P5 feature representation after convolution enhancement.

[0184] P4 and P3 features are generated as follows:

[0185] Pi temp =Conv 1×1 (Ci) (34);

[0186] Pi fused =Pi temp +Pi+1 upsampled (35);

[0187] Pi upsampled =Upsample(Pi temp (36);

[0188] Pi = Conv 3×3 (Pi temp (37);

[0189] Where i takes values ​​in the range [3,4]; Ci represents the fusion result of the (i-1)th layer in the backbone network; Pi temp Indicates the result of feature unification; Pi+1 upsampled Pi represents the upsampling result from the previous layer. fused Indicates the feature fusion result; Pi upsampled Pi represents the upsampling result; Pi represents the feature representation after convolution enhancement.

[0190] In the feature pyramid structure, the generation mechanism of P2 is similar to that of P3 / P4, but since its underlying feature C2 is located in the shallowest layer of the network, it does not need to be upsampled to connect to the next level of features.

[0191] Layer P6 then directly processes the C5 feature map through 3×3 convolutions to generate a larger-scale feature representation, as shown below:

[0192] P6 = Conv 3×3,stride=2 (C5) (38);

[0193] Step S5: Train the detection network using a spatiotemporally aligned multimodal dataset, and optimize the parameter update process by combining a dynamic hard sample mining strategy.

[0194] Current autonomous driving systems primarily rely on traditional onboard sensors (such as RGB cameras) for environmental perception. This approach offers advantages such as high real-time performance and low deployment costs, enabling continuous coverage of the road environment. However, under complex lighting conditions, such as strong glare at tunnel entrances and exits, low light at night, or driving against the light, the limited dynamic range of the sensors leads to a sharp drop in image quality. This, coupled with motion blur caused by high-speed vehicle movement, results in excessively high recall rates for traditional detection models in extreme scenarios. Actual road tests show that such models systematically miss pedestrians (especially those wearing dark clothing) and highly reflective vehicles in low-light environments, severely hindering the functional safety certification of autonomous driving systems.

[0195] Therefore, this invention introduces a neuromorphic visual sensor (event camera) as a key data source. By constructing a dual-branch feature extraction network for event streams and RGB data, event features are used to capture dynamic target edge information, and RGB features are used to extract texture semantic information, achieving cross-modal feature complementarity. A novel unilateral event feature enhancement scheme is proposed, which strengthens the target edge information of the event features extracted by the dual-branch backbone network through an asymmetric feature enhancement module. Then, a novel dynamic dual-branch feature fusion module is proposed to perform pixel-level pre-fusion of the enhanced event feature map and the RGB feature map to reduce the feature imbalance problem. Furthermore, a dual-branch cross-attention module is used to deeply fuse multimodal features, and the fused features are style-aligned to further reduce the feature ambiguity caused by feature differences. After the fused features are enhanced by the attention module, the input decoupled detection head synchronously outputs the target classification probability, bounding box coordinates, and motion state vector, significantly improving the detection robustness in complex traffic scenarios.

[0196] The training process generates parameters such as accuracy, recall, and mAP to help comprehensively understand the model's performance. When the loss function during training becomes smooth, the training of the autonomous driving object detection network based on multimodal feature fusion is complete.

[0197] Step S6: Input the real-time aligned event stream and RGB image into the trained autonomous driving object detection network, and output the object detection box with confidence and motion parameters to provide environmental perception decision support for the autonomous driving system.

[0198] The autonomous driving target detection network model, based on a pre-trained multimodal feature fusion, first inputs any two corresponding RGB images and event frames that were not involved in training; then, it performs feature extraction, one-sided augmentation, and two-branch fusion on the input data; then, based on the processed fused features, the autonomous driving target detection network performs final target classification and bounding box regression, thereby generating a labeled recognition result map of autonomous driving target detection.

[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A target detection method based on the fusion of frame images and event stream features, characterized in that, Includes the following steps: Step S1: Acquire RGB image and event stream data, and use the field of view (FOV) spatial alignment algorithm to achieve spatial matching between the event stream and the RGB image, and construct a spatiotemporally synchronized multimodal data pair; Step S2: Construct a dual-stream feature extraction backbone network to extract event features and RGB features respectively to obtain semantic features; Among them, the RGB branch enhances the target contour representation through deformable convolution to obtain rich texture information; the event branch enhances cross-channel features through AFRM grouped convolution spatial attention to obtain clear contour features. The dual-stream feature extraction backbone network consists of four layers. Each layer has the same structure except for the channel dimension. Feature extraction is achieved through residual operations, as detailed below: Step S21: A single residual unit consists of a 3×3 convolutional layer, batch normalized normalization (BN), and ReLU activation; therefore, the output feature map of a single residual unit... As shown below: (8); in, Input features; This represents a 1×1 convolution operation; For batch normalization; is the activation function for the residual unit; This represents a 3×3 convolution operation; Step S22: After standardization preprocessing, the RGB image is input into the dual-stream feature extraction backbone network. Multi-scale features are extracted using the ResNet-50 architecture, with its first-layer 7×7 convolutional kernel expanding the spatial receptive field to achieve initial feature modeling and obtain the initial RGB features. ; Step S23: The event branch enhances cross-channel features through AFRM grouped convolution spatial attention to obtain clear contour features; First, the event frame sequence is tensor-quantized and then input into the dual-stream feature extraction backbone network. The first layer uses a 5×5 convolutional kernel to enhance the modeling capability of spatiotemporal event distribution, and its output channels are 64. Subsequently, hierarchical feature extraction is performed through a four-level residual module, which gradually reduces the spatial resolution of the feature map to 1 / 16 of the original size. Step S3: Establish cross-modal feature association based on cross-attention mechanism, and achieve semantic alignment using multi-head attention; dynamically optimize fused features using mean-standard deviation method, and adaptively adjust fusion weights according to statistical distribution; Step S4: Construct a multi-level fusion network, input the optimized features into the FPN pyramid for cross-scale fusion, and through the spatial semantic aggregation module, collaboratively enhance the target localization and classification feature representation; Step S5: Train the detection network using a spatiotemporally aligned multimodal dataset and optimize the parameter update process by combining a dynamic hard sample mining strategy; Step S6: Input the real-time aligned event stream and RGB image into the training network, and output the target detection box with confidence and motion parameters to provide environmental perception decision support for the autonomous driving system.

2. The target detection method based on the fusion of frame images and event stream features according to claim 1, characterized in that, In step S1, RGB images and event stream data are acquired, and spatiotemporally synchronized multimodal data pairs are constructed using a field of view (FOV) spatial alignment algorithm. The specific process is as follows: Step S11: Acquire RGB images from a traditional vision sensor and event stream data from a neuromorphic vision sensor, respectively; Step S12: Eliminate multi-camera field-of-view differences in the RGB image through geometric correction, and calculate the transformation matrix between cameras. As shown below: (1); in, This represents the intrinsic parameter matrix of the left camera; This represents the intrinsic parameter matrix of the right camera; This represents the left camera correction rotation matrix; This indicates the rotation matrix for right camera correction; This represents the transformation matrix from camera 0 to camera 1; Step S13: Implement frequency-adaptive representation for the event stream data and dynamically adjust the event accumulation time, as shown below: (2); in, Indicates the time it takes for an event to accumulate; Step S14: For each coordinate point (u,v) in the t-image, calculate the mathematical expression for its coordinate mapping table: (3); (4); (5); (6); in, Represents the homogeneous coordinates in the source image; Represents the homogeneous coordinates of the reprojected point; This represents the unnormalized x-coordinate; This represents the unnormalized y-coordinate; The scaling factor representing homogeneous coordinates; The x-coordinate mapping table used for remapping and The y-coordinate mapping table used for remapping; Step S15: Calculate the mathematical expression for image remapping using the coordinate mapping table: (7); in, Represents the source image; This represents the remapping process; This represents the remapped image.

3. The target detection method based on the fusion of frame images and event stream features according to claim 1, characterized in that, In the event branch, an AFRM module is embedded after each residual unit to enhance the event feature representation capability. The specific process is as follows: Step S231: Divide the channel dimension into 8 groups using grouped convolution, and perform spatial attention calculation on the individual features after grouping, as shown below: (9); (10); (11); in, This represents the feature map after height-oriented pooling. This represents the feature map after pooling in the width direction. This represents adaptive average pooling in terms of height; This represents adaptive average pooling in the width direction; This represents the original grouping features obtained from grouping. Indicates the transpose operation; This represents the merged feature map; Represents connection along the channel dimension ; Step S232: The spliced ​​tensor can be divided into two parts using the Split operation. The global features are obtained using the Sigmoid activation function, as shown below: (12); (13); in, These represent the height attention features and width attention features after segmentation, respectively; Represents the global feature output of the spatial attention branch; Indicates height dimension; Indicates the width dimension; Indicates dimension; This indicates a group normalization operation; This represents the Sigmoid activation function; Step S233: Extract local features through 3×3 convolution and compare them with the obtained global features. Perform cross-pixel multiplication to obtain the weights. As shown below: (14); (15); (16); (17); (18); in, Represents global features; Indicates local features; and These represent the reshaped global and local features, respectively. This indicates a change in the shape of the parameter; Indicates global average pooling; It is a normalization function; This represents the weights calculated using cross-attention; Step S234, Original grouping features and Multiplying the weights yields the initial features of the event. As shown below: (19); in, This represents the Sigmoid activation function.

4. The target detection method based on the fusion of frame image and event stream features according to claim 1, characterized in that, In step S3, the initial characteristics of the event are... With RGB initial features Convolution processing is performed, and global attention is used to fuse bimodal features through element-wise multiplication. A spatial-channel hybrid attention mechanism is then employed to enhance feature representation, as shown below: (20); (21); in, and These represent the RGB image features and event features after the convolution operation, respectively. Indicates the convolution operation; and These represent the RGB enhancement features and event enhancement features after initial fusion, respectively; Step S32: Linearly project the fused and enhanced RGB features and event features onto the query, key, and value respectively, as shown below: (22); (23); (24); in, This represents the projection of the enhanced RGB features onto the query; This represents the projection of the enhanced event features onto the query; This represents the projection of the enhanced RGB features onto the key; This represents the projection of the enhanced event features onto the key; This represents the projection of the enhanced RGB features onto the value; This represents the projection of the enhanced event features onto the value; and These represent the learnable linear projections of RGB and event features in the query, respectively; and These represent the learnable linear projections of RGB and event features in the key, respectively; and These represent the learnable linear projections of RGB and event features in the value, respectively; Step S33: By calculating the dual-branch cross-attention mechanism, the multimodal features are fused as follows: (25); in, and These represent the results of RGB-dominated fusion and event-dominated fusion, respectively; D represents the channel dimension of the feature map; Softmax represents the normalization function. Step S34: Analyze the fusion results. The application of style alignment refines the fusion process, employs the mean-standard deviation method to dynamically optimize the fusion features, and utilizes... Connect the aligned features to obtain the output result. As shown below: (26); (27); (28); in, and These represent the RGB features and event features after style alignment refinement and fusion, respectively. and The mean of the content features; and The standard deviation representing content characteristics; and The mean of the style characteristics; and The standard deviation representing style characteristics; Indicates the join alignment feature operation; (29); (30); in, This represents the formula for calculating the mean. The formula for calculating the standard deviation is given; H represents the height of the feature map; W represents the width of the feature map; ϵ is a value between 10 and 16. -8 Up to 10 -3 Positive numbers between these ranges are used to avoid division by zero errors.

5. The target detection method based on the fusion of frame image and event stream features according to claim 1, characterized in that, In step S4, a multi-level fusion network is constructed, and the optimized features are input into the FPN pyramid for cross-scale fusion. The spatial semantic aggregation module is used to collaboratively enhance the target localization and classification feature representation. The spatial semantic aggregation module constructs a pyramid structure P2-P6 containing rich semantic and spatial information by fusing multi-scale feature maps of the backbone network. First, 1×1 convolution is used to unify all input feature channels to 256 dimensions. Then, 3×3 convolution feature integration is performed on the base layers P2-P5, and the P6 layer is specially extended to adapt to the needs of large-scale target detection. Among them, features P3-P5 are spatially adapted using the nearest neighbor interpolation method, ultimately forming a five-level scale optimized feature pyramid P2 / P3 / P4 / P5 / P6.

6. The target detection method based on the fusion of frame image and event stream features according to claim 5, characterized in that, P5 feature generation is shown below: (31); (32); (33); in, This represents the fusion result of layer 4 in the backbone network; This indicates the unified result of P5 features; Indicates nearest neighbor upsampling; This indicates the upsampling result; This represents the P5 feature representation after convolution enhancement; P4 and P3 features are generated as follows: (34); (35); (36); (37); in, The value range is [3,4]; Indicates the first in the backbone network The result of layer fusion; Indicates the result of feature unification; This indicates the upsampling result from the previous layer; Indicates the result of feature fusion; This indicates the upsampling result; This represents the feature representation after convolution enhancement; In the feature pyramid structure, the generation mechanism of P2 is similar to that of P3 / P4, but its bottom feature C2 is located in the shallowest layer of the network, so there is no need to upsample and connect to the next level of features. Layer P6 directly processes the C5 feature map through 3×3 convolutions, generating the feature representation as shown below: (38)。 7. The target detection method based on the fusion of frame image and event stream features according to claim 1, characterized in that, In step S5, the detection network is trained using a spatiotemporally aligned multimodal dataset, and the parameter update process is optimized by combining a dynamic hard sample mining strategy. The specific process is as follows: First, a neuromorphic visual sensor, namely an event camera, is introduced as the key data source. By constructing a dual-branch feature extraction network for event streams and RGB data, event features are used to capture dynamic target edge information, and RGB features are used to extract texture semantic information, thereby achieving cross-modal feature complementarity. Secondly, a novel unilateral event feature enhancement scheme is designed, which enhances the target edge information by using an asymmetric feature enhancement module to strengthen the event features extracted by the dual-branch backbone network. Then, a novel dynamic dual-branch feature fusion module is designed to perform pixel-level pre-fusion of the enhanced event feature map and the RGB feature map, and to deeply fuse multimodal features through a dual-branch cross-attention module, and to perform style alignment operation on the dual-branch fused features. Then, after the fused features are enhanced by the attention module, they are input into the decoupled detection head and output the target classification probability, bounding box coordinates, and motion state vector simultaneously. Finally, accuracy, recall, and mAP parameters are generated during training to evaluate the model's performance; when the loss function during training becomes smooth, the training of the autonomous driving target detection network based on multimodal feature fusion is complete.