Unmanned aerial vehicle detection and identification method and device based on RGB-event fusion
Patent Information
- Application Number
- CN202611008524.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-08
AI Technical Summary
该方法有效解决了现有融合技术中存在的多尺度语义漂移、互补不充分、融合决策过拟合以及对小目标不敏感等关键问题
[0023] Multimodal deep fusion enhances environmental adaptability: By dynamically adjusting the time window through modal consistency detection and combining time consistency attention with frequency domain fusion, the motion blur caused by high-speed motion and the imaging degradation under extreme lighting conditions are effectively solved, realizing the complementary advantages of RGB and event modality.
Smart Images

Figure CN122574822B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision, specifically relating to a method and apparatus for drone detection and recognition using RGB-Event fusion. Background Technology
[0002] The widespread application of unmanned aerial vehicles (UAVs) has significantly improved efficiency and convenience across various industries, but it has also posed increasingly serious challenges to airspace management and public safety. Therefore, reliable and real-time detection and identification of UAVs has become a crucial and urgent technical requirement in the field of airspace security and intelligent monitoring. However, this task faces many inherent technical challenges in practical applications: First, as a type of low-altitude small aircraft, UAVs are typically very small in physical size, especially under long-distance imaging conditions, often occupying only a few pixels in an image, with extremely limited morphological features, making them easily confused with background noise or other small objects; second, UAVs generally possess high-speed maneuverability, and their trajectories are complex and varied. Coupled with their often complex and diverse operating environments, such as strong light, backlighting, or dim lighting, these factors combined significantly increase the difficulty of achieving stable detection and continuous tracking. These challenges mean that UAV detection technology must not only possess high environmental adaptability but also meet extreme requirements for algorithm robustness, response speed, and detection accuracy.
[0003] At the sensor level, current technologies primarily rely on two approaches: traditional RGB cameras and emerging event cameras. These technologies offer complementary information sources to address the aforementioned challenges from different perspectives. RGB cameras, as mainstream visual sensors, can capture rich color, texture, and semantic information, forming a crucial foundation for target recognition and classification. However, limited by their imaging mechanism, RGB cameras suffer from inherent physical limitations such as limited frame rate and narrow dynamic range. They are prone to severe motion blur when targets move at high speeds, and are susceptible to localized overexposure or underexposure under drastic lighting conditions, resulting in the loss of critical visual information and impacting detection performance. In contrast, event cameras (such as dynamic vision sensors like DVS) employ a novel biomimetic sensing mechanism. They record pixel-level brightness change events asynchronously and sparsely, possessing inherent advantages such as microsecond-level ultra-high temporal resolution, high dynamic range, and low power consumption. This allows them to effectively capture the fast-moving edges and instantaneous contour changes of targets, effectively compensating for the perception shortcomings of RGB modal sensors in dynamic and extreme lighting scenarios.
[0004] While RGB cameras and event cameras can theoretically complement each other well, each still has significant limitations. Event camera outputs frameless, asynchronous streams of events, lacking color information and stable texture semantics. Using them alone makes high-precision target recognition and localization difficult. Therefore, effectively fusing RGB and event modalities to fully leverage the synergistic effect of "1+1>2" has become an inevitable direction for technological development in this field. However, designing a fusion architecture that can deeply integrate, fully exploit, and organically coordinate the advantages of these two heterogeneous data types remains a core scientific problem that urgently needs to be solved.
[0005] While existing research on RGB-Event fusion has made some progress, such as through constructing multi-scale event representations or introducing attention mechanisms for preliminary feature integration, a series of systemic defects and shortcomings still exist. First, existing methods lack effective modeling of the semantic consistency between event frames at different time scales, leading to temporal instability in multimodal fusion results. Second, most fusion strategies remain at the level of unidirectional information guidance or simple feature concatenation, failing to construct deep, bidirectional cross-modal complementarity mechanisms, thus limiting further improvements in fusion performance. More critically, in high-level fusion and decision-making processes (such as multi-expert model selection), existing methods often rely solely on the attention weights learned by the network itself, neglecting the inherent physical and statistical properties of the input data (such as signal-to-noise ratio and information entropy) and the uncertainties in the model's prediction process. This makes the model prone to overfitting on training data and results in insufficient generalization ability and robustness in real-world complex scenarios. Furthermore, during the model optimization phase, the commonly used loss function design is relatively generic and fails to differentiate and target the core challenge of "small target detection," resulting in limited improvement in the model's accuracy and recall for bounding box regression of small drones.
[0006] In summary, to achieve efficient and robust detection of UAVs in complex environments, a novel multi-source data fusion methodology is urgently needed. This method requires systematically deconstructing multimodal complementarity mechanisms, fully integrating underlying physical characteristics with high-level semantic information, and introducing a more scientific fusion and decision-making paradigm. This invention starts with data preprocessing, systematically handles cross-modal feature enhancement and complementarity, further advances to intelligent fusion decision-making, and simultaneously conducts targeted loss optimization to address key issues throughout the entire chain. This comprehensive approach effectively unlocks the potential of RGB image and event data fusion, thereby better meeting the stringent requirements of system accuracy, processing speed, and robustness in practical applications, ensuring the high efficiency and reliability of the technical solution in complex scenarios. Summary of the Invention
[0007] This invention addresses the challenges of small size, high-speed motion, and complex lighting conditions in UAV target detection. It constructs an RGB-Event fusion-based UAV detection and recognition method and device by deeply fusing optical image (RGB) and event stream data and systematically introducing a series of innovative modules. The method first adaptively adjusts the processing granularity of the input data through modal consistency; then, it performs time-frequency dual-domain enhancement on the event stream to extract pure motion features; in the feature extraction stage, it designs a bidirectional mutual guidance mechanism to achieve deep cross-modal complementarity; it employs a sparse fusion strategy guided by distribution and uncertainty to achieve intelligent feature aggregation; and finally, it improves detection performance by optimizing the loss function for small targets. This method effectively solves key problems in existing fusion technologies, such as multi-scale semantic drift, insufficient complementarity, overfitting of fusion decisions, and insensitivity to small targets. In summary, the method of this invention dynamically perceives data quality through modal consistency detection (MCA); constructs robust event representations (TMAF++) through temporal consistent attention (TCA) and frequency domain fusion (FDF); achieves bidirectional complementary enhancement at the feature level through cross-modal guided residual blocks (CGRB); achieves adaptive fusion based on data quality and uncertainty through distributed guided sparse noise-gated attention fusion (DG-SNGAF); and optimizes the model's detection capability for small targets through small target saliency-weighted loss (SOD-loss), forming a complete technical solution from data preprocessing, feature enhancement, complementary fusion to loss optimization. This invention addresses the problems in existing technologies such as loss of RGB image details, sparse event data information, unstable multimodal fusion, lack of statistical guidance in decision-making mechanisms, and insufficient small target detection performance caused by motion blur and exposure issues.
[0008] This invention provides a method for detecting and recognizing drones using RGB-Event fusion, comprising the following steps:
[0009] The first step involves using an event camera to simultaneously acquire optical images (RGB) and event streams. Based on event density and RGB sharpness, modal quality indices are calculated, and time window adjustment factors are generated. ;
[0010] The second step is based on the adjustment factor. Event features are generated and enhanced across multiple time scales, and robust enhanced event features are obtained through temporal consistency attention and frequency domain fusion. ;
[0011] The third step is to combine the image RGB values with... Inputting into a dual-stream backbone network, bidirectional feature complementarity is achieved through cross-modal mutually guided residual blocks to obtain enhanced features. and ;
[0012] The fourth step is to enhance the features. and Multiple cross-modal fusion experts are input, and scores are calculated based on attention, modality statistics, and uncertainty. After noise injection and Top-K selection, the results are weighted and fused to obtain the final fused features. ;
[0013] Fifth, during the training phase, the detection network is trained using the Small Target Saliency Weighted Loss (SOD-loss); during the detection phase, Input the detection network to output the drone detection bounding box and confidence score.
[0014] A drone detection and recognition device based on RGB-Event fusion includes:
[0015] The acquisition module simultaneously acquires optical image RGB and event streams, calculates modal quality indices based on event density and image sharpness, and dynamically generates time window adjustment factors. ;
[0016] The feature acquisition module, based on the adjustment factor Event features are generated and enhanced across multiple time scales, and robust enhanced event features are obtained through temporal consistency attention and frequency domain fusion. ;
[0017] Complementary module, which combines the RGB values of the image with... Inputting a two-stream network, bidirectional feature complementarity is achieved through cross-modal mutually guided residual blocks to obtain enhanced features. and ;
[0018] The fusion module will enhance features. and Multiple cross-modal fusion experts are input, and scores are calculated based on attention, modality statistics, and uncertainty. After noise injection and Top-K selection, the results are weighted and fused to obtain the final fused features. ;
[0019] The result acquisition module will Input the detection network and train it using the Small Object Saliency Weighted Loss (SOD-loss), where the loss contribution is adjusted by saliency weights for each predicted box, thereby outputting the drone detection box and confidence score.
[0020] An electronic device includes: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method.
[0021] A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described thereon.
[0022] The beneficial effects of this invention are as follows:
[0023] Multimodal deep fusion enhances environmental adaptability: By dynamically adjusting the time window through modal consistency detection and combining time consistency attention with frequency domain fusion, the motion blur caused by high-speed motion and the imaging degradation under extreme lighting conditions are effectively solved, realizing the complementary advantages of RGB and event modality.
[0024] Intelligent decision-making enhances generalization ability: cross-modal mutual guidance residual blocks achieve bidirectional feature enhancement, and distributed guidance sparse noise gating integrates physical statistics and uncertainty to perform adaptive expert selection, which reduces the risk of overfitting and significantly improves the robustness and generalization performance of the model in complex scenarios.
[0025] Small target detection accuracy optimization: By using a small target saliency weighted loss function, higher weights are applied to small drone targets, which specifically addresses the problem of weak features and easy missed detection of small targets, effectively improving the accuracy and recall of bounding box regression. Attached Figure Description
[0026] Figure 1 This is a flowchart of the overall process for the RGB-Event fusion-based drone detection and recognition method of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, this invention adopts the following technical solution.
[0028] like Figure 1 As shown, this embodiment of the invention provides a drone detection and recognition method based on RGB-Event fusion, which may include:
[0029] S110: Simultaneously acquires optical image RGB and event stream, calculates modal quality indices based on event density and image sharpness, and dynamically generates time window adjustment factors. ;
[0030] S120: Based on adjustment factor Event features are generated and enhanced across multiple time scales, and robust enhanced event features are obtained through temporal consistency attention and frequency domain fusion. ;
[0031] S130: Combine image RGB with Inputting into a dual-stream backbone network, bidirectional feature complementarity is achieved through cross-modal mutually guided residual blocks to obtain enhanced features. and ;
[0032] S140: Will and Multiple cross-modal fusion experts are input, and scores are calculated based on attention, modality statistics, and uncertainty. After noise injection and Top-K selection, the results are weighted and fused to obtain the final fused features. ;
[0033] S150: During the training phase, the detection network is trained using the Small Target Saliency Weighted Loss (SOD-loss); during the detection phase, Input the detection network to output the drone detection bounding box and confidence score.
[0034] The above-described method first adaptively adjusts the quality of input data through modal consistency detection; then, it performs time-frequency dual-domain enhancement on the event stream to extract clean motion features; in the feature extraction stage, it designs a bidirectional guidance mechanism to achieve deep cross-modal complementarity; it adopts an intelligent fusion strategy guided by distribution and uncertainty to aggregate multi-expert features; and finally, it optimizes the small target detection performance through a targeted loss function, forming a complete UAV detection solution.
[0035] In some embodiments, S110 is implemented by the following method:
[0036] Using a DAVIS or similar event camera, simultaneously outputting RGB image sequences Event Flow ,in Indicates the polarity of brightness change; These represent the x and y coordinates, respectively. Indicates time;
[0037] To adapt to different lighting conditions and motion speeds, event frames employ dynamic time windows. Event density. With time window adjustment factor They are represented as follows:
[0038] (1)
[0039] (2)
[0040] in, Adjustment factor for time window This represents the total number of events within the current time window. and These represent the resolution height and width of the event camera sensor, respectively. It is a reference density. Adjust the scaling constant for the time window. and These are the preset minimum and maximum values for the time window length, used to prevent extreme fluctuations in the time window due to excessively low or high event density. This is the clipping function.
[0041] Using Laplace energy to assess RGB clarity :
[0042] (3)
[0043] in, Indicates RGB resolution. Indicates the image The result after performing Laplacian operator convolution. Indicates variance.
[0044] Modal signal-to-noise ratio estimation:
[0045] (4)
[0046] (5)
[0047] in, and These are the estimated RGB mode signal-to-noise ratio and the event mode signal-to-noise ratio, respectively. The current RGB image frame; This is a two-dimensional event count map within the current time window, where each pixel value represents the cumulative number of events at that location. This represents the mean.
[0048] Modal information entropy estimation:
[0049] (6)
[0050] (7)
[0051] in, and These are the calculated RGB entropy and the event entropy, respectively. This indicates the grayscale or intensity of a pixel in an RGB image. Probability distribution of occurrence; Representing a two-dimensional event counting graph The median pixel value is The probability distribution of occurrence.
[0052] Output dynamic time window adjustment factor RGB clarity Modal signal-to-noise ratio , For use by subsequent modules.
[0053] In some embodiments, in S120, the enhanced event representation (TMAF++) is implemented by the following method:
[0054] Using MCA output Set three incremental time scales , , (usually taken) , , At the current moment, respectively in , , Accumulated event stream over a time span generates corresponding event frames. , , ;
[0055] Each process is handled using a lightweight CNN (3 convolutional layers) with shared weights. ,get For each Perform two-dimensional discrete wavelet transform (DWT) to obtain the low-frequency subband. and high-frequency subbands in three directions ;
[0056] Introduce two learnable channel weight vectors , These two vectors are used to perform channel weighting on the low-frequency and high-frequency sub-bands respectively, and the weighted features are obtained. for: ;in This represents the Hadamard product, which is an element-wise multiplication along the channel dimension. and These are adaptive channel importance weights for low-frequency and high-frequency components, respectively, and their values are automatically updated through backpropagation during network training.
[0057] Calculate the first one respectively The first scale and the first Cosine similarity between scales Based on similarity Calculate the first Consistency score of each scale And using the Softmax function to calculate the consistency score Converted into final fusion weights :
[0058] (8)
[0059] (9)
[0060] (10)
[0061] In formula (8) This represents the total number of feature channels; For channel indexing; and They represent the first The first scale and the first The scale feature at the th in the ... Feature vectors on each channel;
[0062] In formula (9), The total number of time scales. Let be the consistency score for the i-th scale. Reflects the first The average similarity between features at each scale and features at other scales;
[0063] In formula (10), Temperature coefficient used to control the smoothness of the Softmax function.
[0064] Features after weighting at each scale The layers are concatenated along the channel dimension and then dimensionality reduced using a 3×3 convolutional layer. Finally, a weighted summation is performed using TCA weights.
[0065] (11)
[0066] Enhanced event features .
[0067] In some embodiments, in S130, the bimodal cross-guided feature extraction, wherein the cross-modal cross-guided residual block (CGRB) is implemented by the following method:
[0068] After the second residual stage in a two-stream backbone network (such as ResNet), the feature outputs of the RGB branches are obtained. Feature output of event branches , as input to CGRB;
[0069] The process guides the execution of RGB-to-event mapping, utilizing the semantic contextual information of RGB through a channel attention mechanism to calibrate event features:
[0070] Global average pooling is performed on RGB features to extract channel-level global information:
[0071] (12)
[0072] in, This represents the aggregated global channel feature vector; Representation of features In the One channel, spatial coordinates are The feature response value at that location. Channel attention weight vectors are generated using a multilayer perceptron with a "bottleneck" structure. Channel attention weight vectors are generated using a multilayer perceptron with a "bottleneck" structure.
[0073] (13)
[0074] in, Attention weights for the generated RGB channels; Use the Sigmoid activation function; It is a multilayer perceptron.
[0075] Event features are modulated using residual scaling to obtain semantically guided enhanced event features. :
[0076] (14)
[0077] in, For semantically guided enhanced event features, The event characteristics before enhancement.
[0078] The execution event guides the RGB process, enhancing RGB features in residual form through high-frequency edge information:
[0079] From the characteristics of the event High-frequency spatial components were extracted. The Sobel gradient operator was used to calculate the horizontal and vertical gradients, respectively. , And merged into a high-frequency image. :
[0080] (15)
[0081] in, The fused high-frequency map represents the spatial edge and texture intensity of the event flow; and These are gradient maps of the event features in the horizontal and vertical directions, respectively. To prevent small constants from causing instability in numerical calculations.
[0082] Introduce a learnable scaling factor (Initial value 0.1), the high-frequency image is added to the original RGB features as an adaptive edge enhancement residual to obtain the edge-enhanced RGB features. :
[0083] (16)
[0084] in, This is the fused high-frequency graph; This is the original RGB feature. This represents the RGB features after edge enhancement. Finally, CGRB outputs the enhanced bimodal features. and To the subsequent network layers.
[0085] In some embodiments, in S140, the distribution-guided sparse fusion (DG-SNGAF) is implemented by the following method:
[0086] RGB features from the dual-stream backbone network and processed by cross-modal mutual guidance residual blocks Event characteristics Input is fed to N parallel cross-modal fusion experts, each expert outputting fusion feature candidates. and their corresponding attention scores ;
[0087] The RGB clarity output by the modal consistency detection module RGB mode signal-to-noise ratio Event mode signal-to-noise ratio RGB entropy and event entropy We perform weighted fusion to obtain the combined statistics:
[0088] (17)
[0089] (18)
[0090] in, and These are the combined signal-to-noise ratio statistic and the combined information entropy statistic after fusion, respectively, used to characterize the overall quality and information richness of the current input data; and These are the fusion weights of statistics for RGB modality and event modality, respectively. , , , These are the modal quality indices obtained from the first step of the calculation.
[0091] Meanwhile, the prediction uncertainty is estimated using the Monte Carlo Dropout method. Perform T random Dropout forward propagations on the input features and calculate the output variance.
[0092] (19)
[0093] in, The estimated overall prediction uncertainty is used to measure the model's confidence in predicting the current input data; This represents the feature output of the t-th forward propagation. This represents the average of the T outputs; The total number of samples taken in Monte Carlo.
[0094] Calculate a comprehensive score for each expert The score combines attention score, combined statistics, and uncertainty.
[0095] (20)
[0096] in, For the first An initial comprehensive score from a fusion expert; The original attention score output by the expert; , , These are the corresponding balance coefficients.
[0097] To increase the exploratory nature of expert selection and prevent overfitting, adaptive Gaussian noise is injected into the comprehensive score:
[0098] ;
[0099] in, The final score after noise injection; This is a random noise term sampled from a Gaussian distribution; This indicates that the mean is 0 and the variance is 0. The normal distribution; The standard deviation hyperparameter for controlling noise intensity is used to promote diversity in expert selection in the early stages of training.
[0100] Noise added to the score The Top-K selection strategy (usually K=2) is implemented, retaining only the K experts with the highest scores, and finally fusing the features. The weighted sum output by the selected experts.
[0101] In some embodiments, S150, the target detection and optimization are implemented by the following method:
[0102] An FPN+RetinaNet structure is adopted. The input of FPN is the output of layers 2, 3, and 4 (layers 2, 3, and 4) of a dual-stream backbone network (ResNet-50) (after passing through CGRB and DG-SNGAF, the corresponding layers of fused features are taken). High-resolution feature layers such as P2 and P3 are retained for small object detection and to generate target bounding boxes and class confidence scores.
[0103] During the training phase, a small objective significance-weighted loss (SOD-loss) is used:
[0104] (twenty one)
[0105] in, This is the total value of the significance-weighted loss for the small objective; This represents the total number of bounding boxes within the batch; and These represent the predicted class probability and the true class label, respectively. and These are the predicted bounding box coordinates and the actual bounding box coordinates, respectively. The classification focus loss function; For regression loss function; To balance the hyperparameter weights of classification loss and regression loss; for each predicted box... Significance weighting :
[0106] (twenty two)
[0107] in, To control the attenuation coefficient of area sensitivity; This is the normalized area of the predicted bounding box. This weight makes the model pay more attention to small targets during training, thereby effectively improving the detection performance of small drones.
[0108] The methods described above achieve high-precision and robust detection and recognition of UAVs in complex scenarios. Due to the collaborative design based on Modal Consistency Detection (MCA), Cross-Modal Mutual Guidance (CGRB), and Distributed Guidance Fusion (DG-SNGAF), this method ensures spatiotemporal consistency and modal complementarity throughout the entire process of data processing, feature enhancement, and decision fusion, thereby significantly improving the system's stability under conditions of small targets, rapid movement, and variable lighting. This method can achieve reliable real-time UAV perception in practical applications such as urban security and airspace surveillance, providing key technical support for low-altitude safety early warning, illegal target verification, and autonomous countermeasure systems.
[0109] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0110] The above embodiments are provided merely for the purpose of describing the present invention and are not intended to limit the scope of the invention. The scope of the invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the invention should be covered within the scope of the invention.
Claims
1. A method for detecting and recognizing drones using RGB-Event fusion, characterized in that, Includes the following steps: The first step involves simultaneously acquiring optical image RGB data and event streams, calculating modal quality indices based on event density and image sharpness, and dynamically generating time window adjustment factors. ; The second step is based on the adjustment factor. Event features are generated and enhanced across multiple time scales, and robust enhanced event features are obtained through temporal consistency attention and frequency domain fusion. ; The third step is to combine the image RGB values with... Inputting into a dual-stream backbone network, bidirectional feature complementarity is achieved through cross-modal mutually guided residual blocks to obtain enhanced features. and ; The fourth step is to enhance the features. and Multiple cross-modal fusion experts are input, and scores are calculated based on attention, modality statistics, and uncertainty. After noise injection and Top-K selection, the results are weighted and fused to obtain the final fused features. ; Fifth, during the training phase, the detection network is trained using the Small Target Saliency Weighted Loss (SOD-loss). During the testing phase, Input the detection network to output the drone detection bounding box and confidence score; The first step includes: Acquire image RGB sequence Event Flow ,in Indicates the polarity of brightness change. These represent the x and y coordinates, respectively. Indicates time; Event frames use dynamic time windows; event density With time window adjustment factor They are represented as follows: (1) (2) in, This represents the total number of events within the current time window. and These represent the resolution height and width of the event camera sensor, respectively. It is a reference density. Adjust the scaling constant for the time window. and These are the preset minimum and maximum values for the time window length, respectively. Represents the clipping function; Using Laplace energy to assess RGB sharpness: (3) in, Indicates the RGB resolution of the image. This represents the result of performing Laplacian convolution on the RGB values of the image. Variance; Modal signal-to-noise ratio estimation: (4) (5) in, and These are the estimated RGB mode signal-to-noise ratio and the event mode signal-to-noise ratio, respectively. The current RGB frame of the image; This is a two-dimensional event count map within the current time window, where each pixel value represents the cumulative number of events at that location. This represents the mean; Modal information entropy estimation: (6) (7) in, and These are the calculated RGB entropy and the event entropy, respectively. This indicates the grayscale or intensity of a pixel in an RGB image. Probability distribution of occurrence; Representing a two-dimensional event counting graph The median pixel value is The probability distribution of occurrence.
2. The RGB-Event fusion-based drone detection and recognition method according to claim 1, characterized in that, The second step includes: Use adjustment factor Set three incremental time scales , , At the current moment, respectively in , , Accumulated event stream over a time span generates corresponding event frames. , , ; Each event frame is processed using a lightweight convolutional CNN with shared weights. i=1,2,3; preliminary features are obtained. For each Perform a two-dimensional discrete wavelet transform to obtain the low-frequency subband. and high-frequency subbands in three directions ; Introduce two learnable channel weight vectors , These two vectors are used to perform channel weighting on the low-frequency and high-frequency sub-bands respectively, and the weighted features are obtained. for: ;in This represents the Hadamard product, which is an element-wise multiplication along the channel dimension. and These are adaptive channel importance weights for low-frequency and high-frequency components, respectively, and their values are automatically updated through backpropagation during network training. Calculate the first one respectively The first scale and the first Cosine similarity between scales Based on similarity Calculate the first Consistency score of each scale And using the Softmax function to calculate the consistency score Converted into final fusion weights : (8) (9) (10) In formula (8) This represents the total number of feature channels; For channel indexing; and They represent the first The first scale and the first The scale feature at the th in the ... The feature vectors on each channel; in formula (9), The total number of time scales. Let be the consistency score for the i-th scale. Reflects the first The average similarity between features at each scale and features at other scales; in formula (10), Temperature coefficient used to control the smoothness of the Softmax function; Features after weighting at each scale The concatenation is performed along the channel dimension, and then dimensionality is reduced using a 3×3 convolutional layer; finally, a weighted sum is calculated using fusion weights. (11) Enhanced event features .
3. The RGB-Event fusion-based drone detection and recognition method according to claim 2, characterized in that, In the third step, the cross-modal mutual guided residual block CGRB is implemented through the following method: After the second residual stage of the dual-stream backbone network, the feature outputs of the RGB branches are obtained. Feature output of event branches , as input to the cross-modal mutual guided residual block CGRB; The process guides the execution of RGB-to-event mapping, utilizing the semantic contextual information of RGB through a channel attention mechanism to calibrate event features: Global average pooling is performed on RGB features to extract channel-level global information: (12) in, This represents the aggregated global channel feature vector. Indicates the first The global channel feature vector after aggregating all channels; express In the One channel, spatial coordinates are Feature response values at the location; channel attention weight vectors are generated using a multilayer perceptron with a bottleneck structure. : (13) in, Attention weights for the generated RGB channels; Use the Sigmoid activation function; It is a multilayer perceptron; Event features are modulated using residual scaling to obtain semantically guided enhanced event features. : (14) in, For semantically guided enhanced event features, The event characteristics before enhancement; From the characteristics of the event High-frequency spatial components are extracted; the Sobel gradient operator is used to calculate the horizontal and vertical gradients respectively. , And merged into a high-frequency image. : (15) in, The fused high-frequency map represents the spatial edge and texture intensity of the event flow; and These are gradient maps of the event features in the horizontal and vertical directions, respectively. To prevent small constants from causing instability in numerical calculations; The execution event guides the RGB process, enhancing RGB features in residual form through high-frequency edge information: Introduce a learnable scaling factor , high frequency graph As an adaptive edge enhancement residual, it is added to the original RGB features to obtain the edge-enhanced RGB features. ; (16) in, This is the fused high-frequency graph; This is the RGB feature after edge enhancement.
4. The RGB-Event fusion-based drone detection and recognition method according to claim 3, characterized in that, The fourth step includes: RGB features Event characteristics Input is fed to N parallel cross-modal fusion experts, each expert outputting fusion feature candidates. and their corresponding attention scores ; RGB mode signal-to-noise ratio Event mode signal-to-noise ratio RGB entropy and event entropy We perform weighted fusion to obtain the combined statistics: (17) (18) in, and These are the combined signal-to-noise ratio statistic and the combined information entropy statistic after fusion, respectively, used to characterize the overall quality and information richness of the current input data; and These are the fusion weights of statistics for RGB modality and event modality, respectively. Simultaneously, the input features are subjected to T random Dropout forward propagations, and the output variance is calculated: (19) in, The estimated overall prediction uncertainty is used to measure the model's confidence in predicting the current input data; This represents the feature output of the t-th forward propagation. This represents the average of the T outputs; Total number of samples taken for Monte Carlo; Calculate a comprehensive score for each expert The score combines attention score, combined statistics, and uncertainty. (20) in, For the first An initial comprehensive score from a fusion expert; , , These are the corresponding balance coefficients; adaptive Gaussian noise is injected into the overall score: ; in, The final score after noise injection; This is a random noise term sampled from a Gaussian distribution; This indicates that the mean is 0 and the variance is 0. The normal distribution; To control the standard deviation hyperparameter of noise intensity; Noise added to the score The Top-K selection strategy is implemented, retaining only the K experts with the highest scores, and finally fusing the features. The weighted sum output by the selected experts.
5. The RGB-Event fusion-based drone detection and recognition method according to claim 1, characterized in that, In the fifth step, the detection network adopts the structure of Feature Pyramid Network (FPN) and detection head RetinaNet. The input of FPN is the output of layers 2, 3, and 4 of the dual-stream backbone network.
6. The RGB-Event fusion-based drone detection and recognition method according to claim 5, characterized in that, In the fifth step, the fused features are input into a detection network containing a Feature Pyramid Network (FPN) and a RetinaNet detection head to generate target bounding boxes and class confidence scores; during the training phase, a small target saliency-weighted loss is used. (21) in, This is the total value of the significance-weighted loss for the small objective; This represents the total number of bounding boxes within the batch; and These represent the predicted class probability and the true class label, respectively. and These are the predicted bounding box coordinates and the actual bounding box coordinates, respectively. The classification focus loss function; For regression loss function; To balance the hyperparameter weights of classification loss and regression loss; for each predicted box... Significance weighting : (22) in, To control the attenuation coefficient of area sensitivity; This is the normalized area of the prediction box.
7. A drone detection and recognition device using RGB-Event fusion, which implements the method of claim 1, characterized in that, include: The acquisition module simultaneously acquires optical image frames and event streams, calculates modal quality indices based on event density and image sharpness, and dynamically generates time window adjustment factors. ; Feature acquisition module, based on factors Event features are generated and enhanced across multiple time scales, and robust enhanced event features are obtained through temporal consistency attention and frequency domain fusion. ; Complementary module, which combines RGB images with... Inputting into a dual-stream backbone network, bidirectional feature complementarity is achieved through cross-modal mutually guided residual blocks to obtain enhanced features. and ; The fusion module will enhance and Multiple cross-modal fusion experts are input, and scores are calculated based on attention, modality statistics, and uncertainty. After noise injection and Top-K selection, the results are weighted and fused to obtain the final fused features. ; The result acquisition module will Input the detection network and train it using the Small Object Saliency Weighted Loss (SOD-loss), where the loss contribution is adjusted by saliency weights for each predicted box, thereby outputting the drone detection box and confidence score.
8. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the RGB-Event fusion drone detection and recognition method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, cause the processor to implement the RGB-Event fusion drone detection and recognition method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
RGB / Event-based spatial-temporal feature fusion target tracking method
CN122156249A