Pulse neural network target detection method fusing rgb image and event data

CN122618418BActive Publication Date: 2026-09-29UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611110765.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-29
Estimated Expiration
2046-07-24

AI Technical Summary

Technical Problem

[0004]本申请的目的在于提供一种融合RGB图像和事件数据的脉冲神经网络目标检测方法,以解决现有技术中存在的现有单一模态脉冲目标检测方法缺少RGB-事件双模态协同处理的技术问题

Benefits of technology

[0014]在一些实施例中,所述基于所述跨模态融合特征表示进行目标分类和边界框回归,输出目标检测结果,包括:利用目标检测头对所述跨模态融合特征表示进行检测特征处理,得到用于分类预测和边界框回归的检测特征;分别利用目标分类分支以及边界框回归分支对所述检测特征的待检测目标进行处理,得到类别预测结果和边界框预测结果;基于所述类别预测结果及所述边界框预测结果,生成所述目标检测结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618418B_ABST
    Figure CN122618418B_ABST
Patent Text Reader

Abstract

The application discloses a pulse neural network target detection method fusing RGB images and event data, comprising: time segmenting event data to obtain a plurality of sub-time periods, performing polarity statistics on the sub-time periods to obtain a polarity-separated event count representation, performing offset correction on the event count representation based on a preset soft baseline threshold, performing normalization processing on the event count representation after offset correction to obtain an event representation tensor; performing double-branch pulse feature extraction on the RGB images and the event representation tensor respectively, performing asymmetric cross-modal feature correction based on multi-scale RGB features and multi-scale event features; performing cross-modal cross-gate fusion on the corrected RGB features and the corrected event features to generate a cross-modal fusion feature representation, performing target classification and boundary box regression based on the cross-modal fusion feature representation, and outputting a target detection result. The application can improve the accuracy, robustness and energy efficiency performance of target detection in complex dynamic scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and artificial intelligence, and in particular to a spiking neural network object detection method that integrates RGB images and event data. Background Technology

[0002] Event cameras are a novel type of neuromorphic visual sensor. Unlike traditional frame-based cameras that output entire frames of images at fixed time intervals, event cameras asynchronously output event information only when pixel brightness changes. They feature high temporal resolution, high dynamic range, and low data redundancy, enabling them to more effectively perceive dynamic scene changes. Since event data has advantages in representing temporal dynamic information, while RGB images are richer in spatial texture and semantic representation, fusing RGB images with event data to achieve complementary advantages between the two modalities has become an important technological direction for object detection in complex scenes.

[0003] While some existing target detection methods based on spiking neural networks can be compatible with RGB static images or event stream inputs, they are usually treated as interchangeable single-modal inputs. They lack dual-branch collaborative processing and cross-modal interaction mechanisms for time-aligned RGB images and event streams, making it difficult to fully leverage the complementary advantages of RGB spatial semantic information and event temporal dynamic information. Summary of the Invention

[0004] The purpose of this application is to provide a spiking neural network target detection method that integrates RGB image and event data, thereby addressing the technical problem of existing single-modal spiking target detection methods lacking RGB-event dual-modal collaborative processing. The various technical effects of the preferred solutions among the many technical solutions provided in this application are detailed below.

[0005] To achieve the above objectives, this application provides the following technical solutions: This application provides a spiking neural network target detection method that fuses RGB image and event data, comprising: acquiring RGB image and event data; dividing the event data into multiple sub-time periods by time segmentation; performing polarity statistics on the sub-time periods to obtain a polarity-separated event count representation; performing offset correction on the event count representation based on a preset soft baseline threshold; and normalizing the offset-corrected event count representation to obtain an event representation tensor; performing bi-branch spiking feature extraction on the RGB image and the event representation tensor respectively to obtain multi-scale RGB features and multi-scale event features; performing asymmetric cross-modal feature correction based on the multi-scale RGB features and the multi-scale event features to obtain corrected RGB features and corrected event features; performing cross-modal cross-gated fusion on the corrected RGB features and the corrected event features to generate a cross-modal fused feature representation; and performing target classification and bounding box regression based on the cross-modal fused feature representation to output target detection results.

[0006] In some embodiments, the normalization process for the offset-corrected event count representation to obtain an event representation tensor includes: performing Poisson-perception normalization on the offset-corrected event count representation to obtain a normalized event count representation; performing a nonlinear mapping on the normalized event count representation based on an inverse hyperbolic sine function to obtain event feature values ​​after dynamic range compression; performing 8-bit uniform quantization on the event feature values ​​after dynamic range compression to obtain 8-bit uniform quantized event feature values; and organizing the 8-bit uniform quantized event representation in chronological order to obtain the event representation tensor.

[0007] In some embodiments, the asymmetric cross-modal feature correction based on the multi-scale RGB features and the multi-scale event features to obtain the corrected RGB features and the corrected event features includes: performing residual pulse convolution preprocessing on the multi-scale RGB features and the multi-scale event features respectively to obtain preprocessed RGB features and preprocessed event features; correcting the preprocessed event features to obtain a time correction gating applied to the RGB features, and using the time correction gating to perform weighted adjustment on the preprocessed RGB features to obtain weighted adjusted RGB features; correcting the preprocessed RGB features to obtain a channel-space correction gating applied to the event features, and using the channel-space correction gating to perform weighted adjustment on the preprocessed event features to obtain weighted adjusted event features; and performing residual refinement and local feature enhancement on the weighted adjusted RGB features and the weighted adjusted event features respectively to obtain the corrected RGB features and the corrected event features.

[0008] In some embodiments, the step of correcting the preprocessed event features to obtain a time-corrected gating applied to the RGB features includes: obtaining the time change response of the event features based on the event feature difference between adjacent time steps of the preprocessed event features; inputting the time change response to the time-gating generation branch to obtain the time-corrected gating, wherein the time-gating generation branch includes at least a deep convolutional layer, a batch normalization layer, a convolutional layer, a non-linear activation layer, and a sigmoid layer.

[0009] In some embodiments, the step of correcting the preprocessed RGB features to obtain a channel-space correction gating applied to event features includes: generating channel weights through global average pooling and fully connected mapping, generating spatial weights through depthwise separable convolution, multiplying the channel weights and the spatial weights element-wise and pruning them to obtain the channel-space correction gating.

[0010] In some embodiments, the cross-modal cross-gated fusion of the modified RGB features and the modified event features to generate a cross-modal fused feature representation includes: generating event gating weights based on the modified event features, and weighting the modified RGB features according to the event gating weights to obtain gated enhanced RGB features; generating RGB gating weights based on the modified RGB features, and weighting the modified event features according to the RGB gating weights to obtain gated enhanced event features; weighting and aggregating the responses of the gated enhanced event features in the time dimension to obtain aggregated event features; concatenating the gated enhanced RGB features with the aggregated event features to obtain bimodal concatenated features; and performing channel recalibration and feature mapping on the bimodal concatenated features to obtain the cross-modal fused feature representation.

[0011] In some embodiments, generating event gating weights based on the modified event features includes: performing impulse activation and spatial mean pooling on the modified event features to obtain an event feature description containing time dimension and channel dimension information; generating event temporal attention weights and event channel attention weights through time dimension convolution and channel dimension convolution; and multiplying the event temporal attention weights and the event channel attention weights element-wise with the event feature description to obtain the event gating weights.

[0012] In some embodiments, the step of weighting the modified RGB features according to the event gating weights to obtain gated enhanced RGB features includes: aligning the event gating tensor to the time dimension of the RGB features, multiplying it element-wise with the modified RGB features, and then obtaining the gated enhanced RGB features through a pulse projection unit, wherein the pulse projection unit is used for pulse activation, convolution, and normalization.

[0013] In some embodiments, the channel recalibration and feature mapping of the bimodal spliced ​​features includes: performing global average pooling on the bimodal spliced ​​features to obtain a global response description for each channel; generating an importance weight for each channel based on the global response description using a learnable nonlinear mapping and a sigmoid function; weighting the bimodal spliced ​​features channel-by-channel using the importance weights to obtain recalibrated features; and performing feature mapping on the recalibrated features to generate a cross-modal fusion feature representation with a unified dimension.

[0014] In some embodiments, the step of performing target classification and bounding box regression based on the cross-modal fusion feature representation and outputting target detection results includes: processing the cross-modal fusion feature representation using a target detection head to obtain detection features for classification prediction and bounding box regression; processing the target to be detected using the detection features using the target classification branch and the bounding box regression branch respectively to obtain category prediction results and bounding box prediction results; and generating the target detection results based on the category prediction results and the bounding box prediction results.

[0015] Implementing one of the above-described technical solutions of this application has the following advantages or beneficial effects: In this application, event data is sequentially processed by time segmentation, polarity separation counting, soft baseline offset correction, and normalization to generate an event representation tensor adapted for spiking neural network processing. Subsequently, bi-branch pulse feature extraction is performed on the RGB image and the event representation tensor respectively. Asymmetric cross-modal correction is performed before fusion, and cross-modal cross-gated fusion is performed during the fusion stage to generate a cross-modal fusion feature representation. Finally, the target detection result is obtained based on the cross-modal fusion feature representation.

[0016] This application differs from multi-scale feature fusion in single-modal pulse target detection. It can simultaneously utilize the spatial semantic advantages of RGB images and the temporal dynamic advantages of event data to improve the accuracy, robustness, and energy efficiency of target detection in complex dynamic scenes. It can be widely applied to scenarios such as embodied intelligence, autonomous driving, and intelligent monitoring. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a schematic flowchart of a spiking neural network target detection method that integrates RGB image and event data according to an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments and depict various exemplary embodiments that may be adopted to implement this application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of this application disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of this application.

[0019] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0020] To illustrate the technical solutions described in this application, specific embodiments are provided below, showing only the parts related to the embodiments of this application.

[0021] like Figure 1As shown, this application provides a spiking neural network target detection method that fuses RGB images and event data, including the following steps (steps S1 to S3): S1. Acquire RGB image and event data, divide the event data into time segments to obtain multiple sub-time periods, perform polarity statistics on the sub-time periods to obtain a polarity-separated event count representation, perform offset correction on the event count representation based on a preset soft baseline threshold, and normalize the offset-corrected event count representation to obtain the event representation tensor.

[0022] Specifically, an event camera can be used to collect event data in the target scene. Each event in the event data can include pixel position, timestamp, and event polarity information.

[0023] In some embodiments, segmenting event data into multiple sub-time periods can include: segmenting continuously input event data into multiple event segments within a corresponding time range according to a preset time window length; and dividing each event segment into multiple sub-time periods. This preserves the change information of the event stream over time and provides a basis for subsequent event counting and statistics.

[0024] In some embodiments, performing polarity statistics on sub-time periods to obtain a polarity-separated event count representation may include: separately counting positive and negative polarity events within each sub-time period to obtain corresponding positive and negative polarity event counts, thus forming a polarity-separated event count representation.

[0025] In some embodiments, offset correction of the event count representation can be expressed as: in, Indicates the first Each sub-time period, p represents polarity, event count graph ,in , Indicates pixel position, To preset the soft baseline threshold, This is the event count value after offset correction.

[0026] By performing offset correction on the event count representation, the influence of background noise and low-intensity perturbations on subsequent feature representation can be reduced, so as to generate a soft baseline offset-corrected event count representation.

[0027] In some embodiments, normalizing the offset-corrected event count representation to obtain an event representation tensor may include: performing Poisson-sensor normalization on the offset-corrected event count representation to obtain a normalized event count representation; performing a nonlinear mapping on the normalized event count representation based on an inverse hyperbolic sine function to obtain event feature values ​​after dynamic range compression; performing 8-bit uniform quantization on the event feature values ​​after dynamic range compression to obtain 8-bit uniform quantized event feature values; and organizing the 8-bit uniform quantized event representation in chronological order to obtain an event representation tensor.

[0028] Among them, the Poisson perception normalization process is a variance stabilization and normalization process performed using the following inverse hyperbolic sine function, which targets the Poisson statistical characteristics of event counting. The inverse hyperbolic sine function also achieves dynamic range compression, and its process can be expressed as follows: in, For the first dynamic range compressed pixel position within a time period The characteristic value of an event with polarity p. Pixel position after soft baseline offset correction The event count value at polarity p. , For preset scale parameters, As a polarity-specific upper limit threshold, this nonlinear mapping can enhance the feature representation ability of low-activity regions and suppress feature saturation in high-activity regions.

[0029] Specifically, eight-bit uniform quantization can include cropping event feature values ​​to the range [0,1], and then linearly mapping the cropped continuous feature values ​​to the integer range of 0 to 255. Eight-bit uniform quantization can be expressed as: in, The pixel position within the τth sub-time period The eigenvalue of an 8-bit uniformly quantized event with polarity p, ranging from 0 to 255; This means restricting the input value to the range [0,1]. This indicates taking the nearest integer. Through the above quantization process, the normalized continuous event features can be converted into a low-bit-width event representation adapted to the input of a spiking neural network.

[0030] S2. Perform bi-branch pulse feature extraction on the RGB image and event representation tensor respectively to obtain multi-scale RGB features and multi-scale event features. Based on the multi-scale RGB features and multi-scale event features, perform asymmetric cross-modal feature correction to obtain the corrected RGB features and corrected event features.

[0031] In some embodiments, performing bi-branch pulse feature extraction on the RGB image and the event representation tensor respectively may include: copying the RGB image along the time dimension and inputting it into the RGB pulse feature extraction branch, and inputting the event representation tensor into the event pulse feature extraction branch, wherein both the RGB pulse feature extraction branch and the event pulse feature extraction branch include three pulse feature extraction units cascaded in sequence.

[0032] Multi-scale feature information can be obtained by performing stepwise feature encoding and hierarchical representation learning on RGB images and event representation tensors using a spiking neural network, specifically the spiking feature extraction branch mentioned above. Simultaneously, independent network parameters and modality-specific temporal unfolding depths can be set according to the modal characteristics of RGB images and event data to preserve the differences in information distribution and temporal response between the two modalities.

[0033] Furthermore, the RGB pulse feature extraction branch may include a first RGB pulse feature extraction unit, a second RGB pulse feature extraction unit, and a third RGB pulse feature extraction unit. The event pulse feature extraction branch may include a first event pulse feature extraction unit, a second event pulse feature extraction unit, and a third event pulse feature extraction unit.

[0034] The three pulse feature extraction units correspond to three multi-scale feature levels from shallow to deep. Specifically, the first pulse feature extraction unit is used to extract shallow detail features such as edges, textures, contours, and local motion responses at a higher spatial resolution; the second pulse feature extraction unit is used to extract the target's local structure, component relationships, and mid-level context features at a medium spatial resolution after downsampling; and the third pulse feature extraction unit is used to extract high-level semantic features related to the target category at a lower spatial resolution and a larger receptive field.

[0035] The RGB pulse feature extraction branch and the event pulse feature extraction branch correspond to each other in terms of hierarchical settings, and their network parameters can be independent and not shared. The RGB pulse feature extraction branch can encode the spatial texture and semantic information in the RGB image, while the event pulse feature extraction branch can encode the temporal dynamics and motion boundary information in the event representation tensor.

[0036] In some embodiments, performing asymmetric cross-modal feature correction based on multi-scale RGB features and multi-scale event features to obtain corrected RGB features and corrected event features may include: Residual pulse convolution preprocessing is performed on the multi-scale RGB features and the multi-scale event features respectively to obtain the preprocessed RGB features and the preprocessed event features. The preprocessed event features are corrected to obtain a time correction gate that applies to the RGB features. The preprocessed RGB features are then weighted and adjusted using the time correction gate to obtain the weighted RGB features. The preprocessed RGB features are corrected to obtain a channel-space correction gate that acts on the event features. The preprocessed event features are then weighted and adjusted using the channel-space correction gate to obtain the weighted event features. Residual refinement and local feature enhancement are performed on the weighted adjusted RGB features and the weighted adjusted event features respectively to obtain the corrected RGB features and the corrected event features.

[0037] In some embodiments, correcting the preprocessed event features to obtain a time correction gating applied to the RGB features may include: obtaining the time change response of the event features based on the event feature difference between adjacent time steps of the preprocessed event features; inputting the time change response to the time gating generation branch to obtain the time correction gating, wherein the time gating generation branch includes at least a deep convolutional layer, a batch normalization layer, a convolutional layer, a nonlinear activation layer, and a sigmoid layer.

[0038] Specifically, for the first Preprocessed event features at each scale We can first calculate the event feature difference between adjacent time steps to obtain the time change response of the event features, that is, the time change information based on the event features, where the time change response can be expressed as: Subsequently, the time change response can be input into the time gating generation branch to obtain the time correction gating. Then, the time correction gating can be multiplied element-wise with the preprocessed RGB features to obtain the RGB features corrected by the time dynamic information, that is, the weighted RGB features.

[0039] In some embodiments, the weighted RGB feature can be represented as: in, This is a time-corrected gating system, used to characterize the weights of event features related to time variations, motion boundaries, and dynamic responses. This indicates element-wise multiplication. Through the above weighting process, the time-dependent responses in the RGB features can be enhanced, while responses with low time-dependent changes can be suppressed.

[0040] In some embodiments, the preprocessed RGB features are corrected to obtain a channel-space correction gating that acts on the event features. This may include: generating channel weights through global average pooling and fully connected mapping, generating spatial weights through depthwise separable convolution, multiplying the channel weights and spatial weights element-wise and pruning them to obtain the channel-space correction gating.

[0041] Specifically, for the preprocessed RGB features at the l-th scale On the one hand, channel weights can be generated through global average pooling and fully connected mapping. On the other hand, spatial weights can be generated through depthwise separable convolutional branches. Channel-space correction gating can be represented as: Subsequently, the channel-space correction gate is multiplied element-by-element by the preprocessed event features to obtain the event features corrected by spatial texture and semantic information, which is the weighted event features.

[0042] In some embodiments, the weighted event characteristics can be represented as: in, Used to characterize the importance of different channels in RGB features Used to characterize the importance of different spatial locations in RGB features Used to jointly weight the channel response and spatial response of an event.

[0043] Through the above weighting process, the spatial texture and semantic information in the RGB features can be used to enhance the target-related response in the event features and suppress background noise or low-correlation motion response.

[0044] S3. Perform cross-modal cross-gated fusion on the corrected RGB features and the corrected event features to generate a cross-modal fusion feature representation. Based on the cross-modal fusion feature representation, perform target classification and bounding box regression, and output the target detection results.

[0045] In some embodiments, cross-modal cross-gated fusion of the corrected RGB features and the corrected event features to generate a cross-modal fused feature representation may include: Event gating weights are generated based on the corrected event features, and the corrected RGB features are weighted and adjusted according to the event gating weights to obtain the gated enhanced RGB features. Unlike the time correction gating of the aforementioned asymmetric cross-modal feature correction, the event gating weights can be used in the cross-modal cross-gating fusion stage to selectively enhance the RGB features based on the time response and channel response in the event features before fusion. RGB gating weights are generated based on the corrected RGB features, and the corrected event features are weighted and adjusted according to the RGB gating weights to obtain the gating-enhanced event features. Unlike the channel-space correction gating of the aforementioned asymmetric cross-modal feature correction, the RGB gating weights can be used in the cross-modal cross-gating fusion stage to selectively enhance the event features based on the time-channel response in the RGB features before fusion. The time-dimensional response of the gated event features is weighted and aggregated to obtain the aggregated event features. The gated and enhanced RGB features are concatenated with the aggregated event features to obtain bimodal concatenated features. Channel recalibration and feature mapping are then performed on the bimodal concatenated features to obtain cross-modal fusion feature representations.

[0046] In some embodiments, generating event gating weights based on the modified event features may include: performing impulse activation and spatial mean pooling on the modified event features to obtain an event feature description containing information in the time dimension and channel dimension; generating event temporal attention weights and event channel attention weights through time dimension convolution and channel dimension convolution; and multiplying the event temporal attention weights and event channel attention weights element-wise with the event feature description to obtain the event gating weights.

[0047] Specifically, for the first Event characteristics corrected at each scale First, the event features are subjected to impulse activation and spatial mean pooling to obtain event feature descriptions containing information in the temporal and channel dimensions. Then, event temporal attention weights and event channel attention weights are generated by convolution in the temporal and channel dimensions, respectively. These weights are then multiplied element-wise with the event impulse activation features to obtain the event gating tensor. Event Gating Tensor It can be represented as: in, The modified event features are the features after pulse activation. Attention weights based on event time. For event channel attention weights, This indicates element-wise multiplication.

[0048] In some embodiments, the modified RGB features are weighted according to the event gating weights to obtain gated enhanced RGB features. This may include: aligning the event gating tensor to the time dimension of the RGB features and multiplying it element-wise with the modified RGB features, and then obtaining the gated enhanced RGB features through a pulse projection unit, wherein the pulse projection unit is used for pulse activation, convolution and normalization.

[0049] In some embodiments, the gated enhanced RGB features can be represented as: in, This is the corrected RGB feature. This represents a time alignment operation that aligns the event gating tensor to the time dimension of the RGB features. This represents a pulse projection unit composed of pulse activation, convolution, and normalization. Through the above processing, the temporal dynamics and channel responses in the event features can be used to selectively weight and enhance the RGB features before fusion.

[0050] In some embodiments, generating RGB gating weights based on the modified RGB features may include: performing impulse activation and spatial mean pooling on the modified RGB features to obtain an RGB feature description containing information in the time dimension and channel dimension; generating RGB temporal attention weights and RGB channel attention weights through temporal convolution and channel convolution; and multiplying the RGB temporal attention weights and RGB channel attention weights element-wise with the RGB feature description to obtain the RGB gating weights.

[0051] Specifically, for the first RGB features corrected at each scale First, pulse activation and spatial mean pooling are applied to the RGB feature description, which contains information in both the temporal and channel dimensions. Then, RGB temporal attention weights and RGB channel attention weights are generated by convolution in the temporal and channel dimensions, respectively. These weights are then multiplied element-wise with the RGB pulse activation feature to obtain the RGB gating tensor. .

[0052] In some embodiments, the RGB gate tensor It can be represented as: In some embodiments, the gating enhancement of the event features is obtained by weighting the modified event features according to the RGB gating weights. This may include: aligning the RGB gating tensor with the time dimension of the event features according to the time dimension, multiplying it element-wise with the modified event features, and then obtaining the gating enhancement of the event features through a pulse projection unit.

[0053] In some embodiments, the gating-enhanced event characteristics can be represented as: In some embodiments, weighted aggregation of the response of the gated enhanced event features in the time dimension may include: performing global average pooling on the event features at each time step to obtain a channel description vector; generating an importance score for each time step using a multilayer perceptron based on the channel description vector; performing Softmax normalization on the importance score to obtain a time attention weight; and performing weighted summation on the event features at each time step according to the time attention weight.

[0054] Specifically, for the first Event characteristics after gating enhancement at individual scales It contains Event characteristics at each time step ,in First, the event characteristics at each time step. Global average pooling is performed to obtain the corresponding channel description vector; then, a multilayer perceptron is used to generate the importance score for that time step, and softmax normalization is performed on the importance scores of all time steps to obtain the temporal attention weights. .

[0055] In some embodiments, temporal attention weights It can be represented as: in, Indicates global average pooling. This represents a multilayer perceptron. Indicates the first Each time step has a time attention weight, and the sum of the weights for each time step is 1. Finally, the event features of each time step are weighted and summed according to the time attention weights to obtain the aggregated event features.

[0056] In some embodiments, the aggregated event characteristics can be represented as: By using the above-mentioned temporal attention pooling process, it is possible to compress the temporal dimension of event features, highlight the event time slices that contribute more to object detection, and suppress the responses of low-relevance or noisy time slices.

[0057] In some embodiments, channel recalibration and feature mapping of bimodal concatenated features may include: performing global average pooling on the bimodal concatenated features to obtain a global response description for each channel; generating importance weights for each channel based on the global response description using a learnable nonlinear mapping and a sigmoid function; weighting the bimodal concatenated features channel by channel using the importance weights to obtain recalibrated features; and performing feature mapping on the recalibrated features to generate a cross-modal fusion feature representation with a unified dimension.

[0058] Specifically, for the first At each scale, the gated enhanced RGB features and the aggregated event features are first concatenated along the channel dimension to obtain the bimodal concatenated features. .

[0059] In some embodiments, dual-modal splicing features It can be represented as: Subsequently, the dual-modal splicing features were analyzed. Perform channel recalibration. Specifically, for Global average pooling is performed to obtain the global response description of each channel. Then, a learnable nonlinear mapping and a sigmoid function are used to generate the importance weights of each channel. .

[0060] In some embodiments, importance weighting It can be represented as: in, Indicates global average pooling. and Indicates learnable mapping parameters, This represents the Sigmoid function. This represents the importance weight of each channel. Furthermore, the importance weight can be adaptively learned during network training based on the object detection loss. A larger weight indicates a higher correlation between the corresponding channel and the object classification or bounding box regression task, while a smaller weight indicates a low correlation or redundant response of the corresponding channel.

[0061] In some embodiments, the features after channel recalibration It can be represented as: in, This indicates element-wise multiplication.

[0062] In some embodiments, cross-modal fusion feature representation It can be represented as: in, This represents a pulse projection unit consisting of pulse activation, convolution, and normalization.

[0063] Through the above processing, the channel responses relevant to the target detection task can be enhanced, low-correlation or redundant channel responses can be suppressed, and RGB features and event features can be mapped into a cross-modal fusion feature representation of a unified dimension.

[0064] In some embodiments, target classification and bounding box regression are performed based on cross-modal fusion feature representations, and the output target detection results may include: The target detection head is used to process the cross-modal fusion feature representation to obtain detection features for classification prediction and bounding box regression. The target to be detected is processed by the target classification branch and the bounding box regression branch respectively to obtain the category prediction result and the bounding box prediction result. Based on the category prediction result and the bounding box prediction result, the target detection result is generated.

[0065] For example, in autonomous driving scenarios, synchronously acquired RGB images and event data can be used to detect targets such as vehicles, pedestrians, and non-motorized vehicles in road scenes. The system outputs category predictions and corresponding bounding box predictions for each detected target, providing target information for subsequent environmental perception, risk assessment, and path planning. Especially in nighttime, backlighting, fast-moving, or high dynamic range scenarios, event data can supplement the temporal dynamic information of RGB images, thereby improving the stability of target detection.

[0066] In this application, event data is sequentially processed through time segmentation, polarity separation counting, soft baseline offset correction, and normalization to generate an event representation tensor adapted for spiking neural network processing. Subsequently, bi-branch pulse feature extraction is performed on the RGB image and the event representation tensor respectively. Asymmetric cross-modal correction is performed before fusion, and cross-modal cross-gated fusion is performed during the fusion stage to generate a cross-modal fusion feature representation. Finally, the target detection result is obtained based on the cross-modal fusion feature representation.

[0067] This application differs from multi-scale feature fusion in single-modal pulse target detection. It can simultaneously utilize the spatial semantic advantages of RGB images and the temporal dynamic advantages of event data to improve the accuracy, robustness, and energy efficiency of target detection in complex dynamic scenes. It can be widely applied to scenarios such as embodied intelligence, autonomous driving, and intelligent monitoring.

[0068] Those skilled in the art will understand that all or part of the features / steps of the above-described method embodiments can be implemented by methods, data processing systems, or computer programs. These features may be implemented without hardware, entirely in software, or in a combination of hardware and software. The aforementioned computer program may be stored in one or more computer-readable storage media. When the computer program is executed (e.g., by a processor), it performs the steps of the above-described embodiments of the spiking neural network object detection method that fuses RGB images and event data.

[0069] The aforementioned storage media capable of storing program code include: static hard disks, solid-state hard disks, random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), optical storage devices, magnetic storage devices, flash memory, magnetic disks or optical disks, and / or combinations of the above devices, that is, they can be implemented by any type of volatile or non-volatile storage devices or combinations thereof.

[0070] This application also provides a processing device embodiment, including one or more processors and a memory; wherein the memory is used to store one or more computer programs, and the one or more processors are used to execute the one or more computer programs stored in the memory, so that the processors execute the features / steps of the above-described embodiment of the spiking neural network object detection method that fuses RGB image and event data.

[0071] The above description is merely a preferred embodiment of this application. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of this application. Furthermore, under the teachings of this application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this application. Therefore, this application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of this application.

Claims

1. A spiking neural network target detection method that integrates RGB image and event data, characterized in that, include: RGB image and event data are acquired. The event data is segmented into multiple sub-time periods. Polarity statistics are performed on the sub-time periods to obtain a polarity-separated event count representation. The event count representation is offset-corrected based on a preset soft baseline threshold. The offset-corrected event count representation is normalized to obtain an event representation tensor. Bi-branch impulse feature extraction is performed on the RGB image and the event representation tensor respectively to obtain multi-scale RGB features and multi-scale event features. Based on the multi-scale RGB features and the multi-scale event features, asymmetric cross-modal feature correction is performed to obtain the corrected RGB features and corrected event features. Cross-modal cross-gated fusion is performed on the modified RGB features and the modified event features to generate a cross-modal fusion feature representation. Target classification and bounding box regression are performed based on the cross-modal fusion feature representation to output the target detection result. The normalization process for the offset-corrected event count representation to obtain the event representation tensor includes: performing Poisson-perception normalization on the offset-corrected event count representation to obtain a normalized event count representation; performing a nonlinear mapping on the normalized event count representation based on an inverse hyperbolic sine function to obtain event feature values ​​after dynamic range compression; performing 8-bit uniform quantization on the dynamically range compressed event feature values ​​to obtain 8-bit uniformly quantized event feature values; and organizing the 8-bit uniformly quantized event feature values ​​in chronological order to obtain the event representation tensor. The asymmetric cross-modal feature correction based on the multi-scale RGB features and the multi-scale event features, to obtain the corrected RGB features and the corrected event features, includes: The multi-scale RGB features and the multi-scale event features are respectively subjected to residual pulse convolution preprocessing to obtain preprocessed RGB features and preprocessed event features; The preprocessed event features are corrected to obtain a time correction gate that acts on the RGB features. The preprocessed RGB features are then weighted and adjusted using the time correction gate to obtain the weighted RGB features. The preprocessed RGB features are corrected to obtain a channel-space correction gating that acts on the event features. The preprocessed event features are then weighted and adjusted using the channel-space correction gating to obtain the weighted event features. The weighted RGB features and the weighted event features are subjected to residual refinement and local feature enhancement respectively to obtain the corrected RGB features and the corrected event features.

2. The spiking neural network target detection method fusing RGB image and event data according to claim 1, characterized in that, The step of correcting the preprocessed event features to obtain a time-corrected gating applied to the RGB features includes: Based on the event feature difference between adjacent time steps of the preprocessed event features, the temporal change response of the event features is obtained; the temporal change response is input to the temporal gating generation branch to obtain the temporal correction gating, wherein the temporal gating generation branch includes at least a deep convolutional layer, a batch normalization layer, a convolutional layer, a nonlinear activation layer, and a sigmoid layer.

3. The spiking neural network target detection method fusing RGB image and event data according to claim 1, characterized in that, The step of correcting the preprocessed RGB features to obtain a channel-space correction gating applied to event features includes: Channel weights are generated by global average pooling and fully connected mapping, and spatial weights are generated by depthwise separable convolution. The channel weights and spatial weights are then multiplied element-wise and pruned to obtain the channel-space correction gating.

4. The spiking neural network target detection method according to claim 1, characterized in that, The step of performing cross-modal cross-gated fusion of the corrected RGB features and the corrected event features to generate a cross-modal fused feature representation includes: Based on the modified event features, an event gating weight is generated, and the modified RGB features are weighted and adjusted according to the event gating weight to obtain the gating-enhanced RGB features; RGB gating weights are generated based on the modified RGB features, and the modified event features are weighted and adjusted according to the RGB gating weights to obtain gating-enhanced event features; The time-dimensional response of the gated enhanced event features is weighted and aggregated to obtain the aggregated event features; The gated enhanced RGB features are concatenated with the aggregated event features to obtain bimodal concatenated features. Channel recalibration and feature mapping are then performed on the bimodal concatenated features to obtain the cross-modal fusion feature representation.

5. The spiking neural network target detection method fusing RGB image and event data according to claim 4, characterized in that, The generation of event gating weights based on the modified event features includes: The modified event features are subjected to impulse activation and spatial mean pooling to obtain an event feature description containing information in the time and channel dimensions. Event temporal attention weights and event channel attention weights are generated by convolution in the time and channel dimensions. The event temporal attention weights and the event channel attention weights are multiplied element-wise with the event feature description to obtain the event gating weights.

6. The spiking neural network target detection method according to claim 4, characterized in that, The step of weighting the modified RGB features according to the event gating weights to obtain gated enhanced RGB features includes: aligning the event gating weights to the time dimension of the RGB features, multiplying them element-wise with the modified RGB features, and then obtaining the gated enhanced RGB features through a pulse projection unit, wherein the pulse projection unit is used for pulse activation, convolution, and normalization.

7. The spiking neural network target detection method according to claim 4, characterized in that, The process of channel recalibration and feature mapping for the bimodal spliced ​​features includes: performing global average pooling on the bimodal spliced ​​features to obtain global response descriptions for each channel; generating importance weights for each channel based on the global response descriptions using a learnable nonlinear mapping and a sigmoid function; applying the importance weights to the bimodal spliced ​​features channel-wise to obtain recalibrated features; and performing feature mapping on the recalibrated features to generate a unified-dimensional cross-modal fusion feature representation.

8. The spiking neural network target detection method according to claim 1, characterized in that, The process of classifying and regressing bounding boxes based on the cross-modal fusion feature representation, and outputting target detection results, includes: The cross-modal fusion feature representation is processed using a target detection head to obtain detection features for classification prediction and bounding box regression. The target to be detected in the detection features is processed using the target classification branch and the bounding box regression branch respectively to obtain the category prediction result and the bounding box prediction result. The target detection result is generated based on the category prediction result and the bounding box prediction result.

Citation Information

Patent Citations

  • Multi-scale dynamic fusion target detection method and system based on spiking neural network

    CN120726303A

  • Pulse neural network tracking method and system based on event camera

    CN120997257A