A low-delay visible light target detection method based on event stream inter-frame compensation
Patent Information
- Application Number
- CN202611040607.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-14
AI Technical Summary
[0006]本发明旨在解决现有可见光目标检测方法在低帧率或快速运动场景下存在的检测更新延迟问题,以及事件流与可见光特征直接融合时容易受到事件稀疏、异常触发和跨模态特征不匹配影响的问题
[0017]第一,本发明在相邻两帧可见光图像之间设置多个事件更新时刻,使系统能够在下一帧可见光图像到达之前连续输出帧间目标检测结果,从而降低由可见光帧率限制引起的检测更新延迟。
Smart Images

Figure CN122574744B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, event vision, and target detection technology, and particularly to a low-latency visible light target detection method based on event stream inter-frame compensation. Background Technology
[0002] Visible light cameras provide rich semantic information on texture, color, contour, and category, making them a commonly used visual sensor in object detection systems. However, visible light cameras typically output images at a fixed frame rate, resulting in unavoidable time intervals between adjacent frames. When the target moves quickly, the scene changes drastically, or the system requires a high detection update frequency, relying solely on visible light image frames for object detection can easily lead to problems such as delayed detection state updates, discontinuous target positions, and missing detections at intermediate moments.
[0003] While increasing the frame rate of visible light cameras can reduce detection latency to some extent, it also leads to increased image data volume, higher computational resource consumption, increased power consumption, and higher hardware costs. For edge computing, automotive applications, mobile platforms, or real-time control systems, simply increasing the visible light frame rate is not always feasible.
[0004] An event camera is an asynchronous visual sensor whose pixels independently output events when they detect brightness changes exceeding a trigger threshold. Event cameras feature high temporal resolution, low redundancy output, and sensitivity to motion changes, providing continuous information on brightness, edge, and motion changes between adjacent visible light frames. Therefore, using the event stream as inter-frame dynamic compensation information for visible light target detection helps increase the target detection update frequency without significantly increasing the visible light frame rate.
[0005] However, event streams themselves do not directly contain complete target texture and category semantic information, and event data is characterized by asynchronicity, sparsity, polarity differences, and noise triggering. Directly converting event streams into images or simply concatenating them with visible light features can easily lead to problems such as insufficient reliability of event information, false detections introduced by abnormal events, difficulties in cross-modal feature alignment, and unstable inter-frame detection states. Therefore, a low-latency target detection method is needed that can use visible light detection results as semantic anchors, event streams as inter-frame variation compensation, and adaptively adjust event contributions based on the effectiveness of the event window. Summary of the Invention
[0006] This invention aims to solve the problem of detection update delay in existing visible light target detection methods in low frame rate or fast motion scenarios, as well as the problem that event streams and visible light features are easily affected by event sparsity, abnormal triggering and cross-modal feature mismatch when directly fused.
[0007] To achieve the above objectives, this invention provides a low-latency visible light target detection method based on event stream inter-frame compensation. This method uses the basic detection results of visible light images and multi-scale image features as detection anchors, and the asynchronous event stream output by the event camera as inter-frame dynamic compensation information. Multiple event update times are set between two adjacent visible light image frames, and a corresponding inter-frame target detection result is generated for each event update time.
[0008] Specifically, the method includes:
[0009] Step S1: Construct a synchronous data unit for the visible light image sequence and the asynchronous event stream, acquire the visible light image sequence and the asynchronous event stream output by the event camera, and perform time synchronization and spatial coordinate alignment between the visible light image sequence and the asynchronous event stream;
[0010] Step S2: Generate and cache visible light detection anchor points for the first... Frame visible light image Perform target detection to obtain basic detection results. and multi-scale visible light fusion features and the basic test results and multi-scale visible light fusion features As the first Buffering visible light detection anchor points for frames;
[0011] Step S3: Construct the inter-frame event representation of the current event window, in the... Frame visible light image With the Frame visible light image Set multiple event update times between them Update time for any event Extract event data from the asynchronous event stream within the current event window and convert the event data into an inter-frame event tensor. Then, the multi-scale event features corresponding to the current event update time are obtained through the event feature extraction network. ;
[0012] Step S4: Evaluate the effectiveness of the event window and generate event compensation weights for the current event update time. The corresponding event window is evaluated for validity, and event compensation weights are generated based on the evaluation results. ;
[0013] Step S5: Fuse visible light detection anchor points and event compensation weights, and cache the multi-scale visible light fusion features from step S2. The multi-scale event features obtained in step S3 Perform same-scale alignment and utilize event-compensated weights. By adjusting the contribution of event features, cross-modal compensation features are obtained. ;
[0014] Step S6: Generate and update the inter-frame target detection state, and apply the cross-modal compensation features obtained in step S5. Input the inter-frame detection header to generate the current event update time. Inter-frame target detection results And cache it as the current inter-frame detection state;
[0015] Step S7: Continuously output inter-frame detection status and refresh visible light detection anchor points. Frame visible light image With the Frame visible light image Between these steps, steps S3 to S6 are repeated for each event update time, continuously outputting multiple inter-frame target detection states; when the first... Frame visible light image Upon arrival, step S2 is executed again to generate a new visible light detection anchor point, which is then used as the reference basis for subsequent inter-frame compensation of the event stream.
[0016] Compared with the prior art, the present invention has at least the following beneficial effects.
[0017] First, the present invention sets multiple event update times between two adjacent visible light images, enabling the system to continuously output inter-frame target detection results before the next visible light image arrives, thereby reducing the detection update delay caused by the visible light frame rate limitation.
[0018] Second, the present invention uses visible light detection results and multi-scale visible light fusion features as detection anchors, enabling the inter-frame detection process to inherit the target appearance, category semantics and basic location information in the visible light image, thus avoiding the problem of insufficient semantic information caused by relying solely on event streams.
[0019] Third, the present invention utilizes the high temporal resolution brightness change and motion change information provided by the event stream to compensate for the target position and state changes between adjacent visible light frames, which is beneficial to improving the detection continuity of fast-moving targets.
[0020] Fourth, this invention evaluates the effectiveness of the event window by the number of events, the proportion of active pixels, and the polarity balance, and generates event compensation weights. This can enhance the contribution of event features when the event information is reliable, and reduce the impact of event features when events are sparse, locally triggered, or polarity is abnormal, thereby improving the stability of inter-frame detection.
[0021] Fifth, the present invention adopts a same-scale alignment and fusion method of multi-scale visible light features and multi-scale event features, which can take into account the semantic expression of the target, spatial positioning information and inter-frame dynamic change information, and improve the detection adaptability of targets of different scales in the inter-frame update process.
[0022] Sixth, the present invention outputs the inter-frame target detection status in a unified image coordinate system, rather than the reconstructed visible light image. Therefore, it can avoid the additional computational overhead caused by image reconstruction and facilitate direct access to target tracking, alarm, control or decision-making modules. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments are briefly described below. Obviously, the following drawings are only used to illustrate some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0024] Figure 1 This is a schematic diagram of the overall process of the low-latency visible light target detection method based on event stream inter-frame compensation of the present invention.
[0025] Figure 2 This is a timing diagram illustrating the event update time setting and inter-frame detection result output between two adjacent visible light images of the present invention, used to show the time relationship between the k-th visible light image, the k+1-th visible light image, the event update time, and the inter-frame detection result.
[0026] Figure 3 This is a schematic diagram of the cross-modal fusion structure of visible light detection anchor point features and event compensation features of the present invention, used to illustrate the generation process of multi-scale visible light fusion features, multi-scale event features of the current event window, event compensation weights, and cross-modal compensation features. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, this invention adopts the following technical solution.
[0028] This embodiment provides a low-latency visible light target detection method based on event stream inter-frame compensation, such as... Figure 1As shown, this method uses the target detection results of visible light images and multi-scale image features as detection anchors, and uses the asynchronous event stream output by the event camera as inter-frame dynamic compensation information. Multiple event update times are set between two adjacent visible light image frames, and corresponding inter-frame target detection results are generated.
[0029] The visible light image provides semantic information about the target's appearance and category, while the event stream provides information on brightness changes, edge changes, and motion changes between two frames. The inter-frame target detection result represents the target's detection state in the unified image coordinate system at the current event update time, including the target category, target bounding box coordinates, and target confidence score. It can serve as an intermediate detection output before the next visible light image arrives, reducing the detection update delay caused by the limited visible light frame rate.
[0030] Step S1: Construct a synchronized data unit for visible light images and event streams;
[0031] S1.1 Obtain a sequence of visible light images;
[0032] A sequence of visible light images is acquired using a visible light camera, denoted as:
[0033] ;
[0034] Its corresponding timestamp is:
[0035] ;
[0036] Where K represents the total number of frames in the visible light image sequence. , This represents the visible light image of the k-th frame. This represents the acquisition timestamp of the k-th visible light image frame, with the superscript I indicating that the timestamp belongs to a visible light image frame.
[0037] The visible light image is mainly used to provide semantic information about the appearance, texture, outline, and category of the target.
[0038] S1.2 Obtain the event stream;
[0039] The asynchronous event stream is obtained through the event camera, and the event stream is denoted as:
[0040] ;
[0041] Where N represents the total number of events in the event stream, and each event can be represented as:
[0042] ;
[0043] in, This represents the spatial coordinates of the i-th event in the event camera pixel plane. This represents the timestamp of the i-th event. This indicates the event polarity of the i-th event, and the superscript E indicates that the timestamp belongs to the event stream data.
[0044] The polarity of the event This is used to characterize the direction of brightness change at the corresponding pixel location. The event stream is used to characterize the asynchronous brightness change information of the target or background between two adjacent visible light images, and provides a data foundation for the construction of subsequent inter-frame event representations and target detection time compensation.
[0045] S1.3 Time synchronization;
[0046] Unify the visible light image timestamp and event timestamp to the same time base to ensure that event data within the corresponding time range can be accurately extracted based on the visible light image acquisition time.
[0047] If there is a time offset between the visible light camera and the event camera Then the event timestamp is corrected:
[0048] ;
[0049] The corrected event is:
[0050] ;
[0051] in, This represents the corrected event timestamp. This represents the time correction amount between the event camera's time reference and the visible light camera's time reference. If the visible light camera and the event camera have already been synchronized via hardware triggering or a unified clock, then let:
[0052] ;
[0053] At this point, the corrected event timestamp is the same as the original event timestamp.
[0054] S1.4 Spatial coordinate alignment;
[0055] If the imaging coordinate systems of the visible light camera and the event camera are different, the spatial mapping relationship between the two is obtained through camera calibration, and the event coordinates are transformed into the visible light image coordinate system.
[0056] Let the mapping function from event coordinates to visible light image coordinates be:
[0057] ;
[0058] in, This represents a coordinate mapping function determined by camera intrinsic parameters, extrinsic parameters, or homography transformation. Indicates the first The pixel coordinates of an event in the event camera coordinate system. Indicates the first Each event is mapped to its corresponding coordinates in the visible light image coordinate system.
[0059] To facilitate unified processing later, the spatial coordinates of the event in a unified image coordinate system are defined as follows:
[0060] ;
[0061] If the event coordinates are already aligned with the visible light image coordinates, then:
[0062] ;
[0063] If the event coordinates need to be mapped to the visible light image coordinate system, then:
[0064] ;
[0065] In subsequent steps, the unified spatial coordinates are used. Construct an inter-frame event representation.
[0066] Step S2: Generate and cache visible light detection anchor points;
[0067] For the Frame visible light image Target detection is performed to generate visible light detection anchor points corresponding to the frame image. The visible light detection anchor points include the first... The basic detection results of the visible light image and the multi-scale image features extracted by the target detection network.
[0068] In this embodiment, the visible light target detection network adopts a YOLO-like single-stage target detection network for detecting the first target. Frame visible light image Basic object detection is performed to obtain object category, object bounding box, and object confidence score. The YOLO-like single-stage object detection network includes an input preprocessing module, a backbone feature extraction module, a neck multi-scale feature fusion module, and a head detection head.
[0069] It should be noted that the YOLO-like single-stage object detection network described is only one specific implementation and does not constitute a limitation on the scope of protection of this invention. In other embodiments, Faster R-CNN, RetinaNet, FCOS, YOLOX, YOLOv5, YOLOv8, DETR, or other object detection networks capable of outputting object category, object bounding box, object confidence score, and intermediate features may also be used.
[0070] S2.1 Image preprocessing;
[0071] For the Frame visible light image After resizing, pixel normalization, and channel normalization, the preprocessed visible light image is obtained:
[0072] ;
[0073] in, and These represent the image height and image width input to the object detection network, respectively, and 3 represents the number of channels in the visible light image.
[0074] In one specific embodiment, the input image is adjusted to and normalize the pixel values to Channel standardization can be performed by using preset mean and variance to standardize the pixel values of each channel, thereby reducing the impact of differences in image brightness and color distribution under different acquisition conditions on the target detection results.
[0075] S2.2 Backbone Feature Extraction;
[0076] Preprocessed visible light image Input the YOLO-like network's backbone network to extract image features at different spatial scales.
[0077] In one specific embodiment, the backbone feature extraction module can employ a CSPDarknet, Darknet, ResNet, or lightweight convolutional network structure. The backbone feature extraction module performs progressive convolution and downsampling processing on the input image. It outputs feature maps at multiple scales, denoted as:
[0078] ;
[0079] in:
[0080] ;
[0081] ;
[0082] ;
[0083] in, , , These represent the number of channels in the corresponding scale feature map; , and These represent the feature map space dimensions at different downsampling factors.
[0084] in, It has high spatial resolution and is mainly used to preserve the target's edges, contours, and positional information; Used to characterize target structural information at medium scale; It has a large receptive field and is mainly used to represent the deep semantic information of the target. The above multi-scale image features are used for multi-scale feature fusion in the subsequent Neck module.
[0085] S2.3 Neck multi-scale feature fusion will be implemented in the following steps;
[0086] The multi-scale image features obtained in S2.2 are input into the Neck multi-scale feature fusion module of the YOLO-type object detection network to fuse image features of different scales and obtain multi-scale visible light fusion features.
[0087] In one specific embodiment, the Neck multi-scale feature fusion module adopts a structure combining FPN and PAN. The FPN structure is used to transfer semantic information from deep features to shallow features, thereby enhancing the target semantic expression capability of the shallow features; the PAN structure is used to transfer spatial location information from shallow features to deep features, thereby enhancing the localization expression capability of the deep features.
[0088] Specifically, the Neck multi-scale feature fusion module processes the feature map output by Backbone. , , Upsampling, downsampling, feature concatenation, and convolutional fusion are performed to obtain multi-scale visible light fusion features:
[0089] ;
[0090] in, , , They represent the first Fusion features of visible light images at different scales, superscript This indicates that the feature originates from a visible light image.
[0091] in, It has high spatial resolution and is mainly used to preserve the target's edges, contours, and positional information; Used to characterize structural information of medium-scale targets; It has strong semantic expression capabilities and is mainly used to represent the deep semantic information of the target.
[0092] The multi-scale visible light fusion feature It is cached as part of the visible light detection anchor point and used for inter-frame compensation fusion of event features corresponding to the current event window.
[0093] S2.4 Head detection head prediction;
[0094] The multi-scale visible light fusion features obtained in step S2.3 are:
[0095] ;
[0096] Input the head of the YOLO class object detection network respectively, and predict the first... Target category, target bounding box, and target confidence in a frame of visible light image.
[0097] In one specific embodiment, the head detection head adopts a decoupled detection head structure, including a classification branch, a bounding box regression branch, and a confidence branch. The classification branch is used to predict the class probability of the candidate target, the bounding box regression branch is used to predict the bounding box position parameters of the candidate target, and the confidence branch is used to predict the probability that a target exists at the candidate location.
[0098] For the Visible light fusion features at various scales ,in Branch output:
[0099] ;
[0100] in, Indicates the number of target categories. and They represent the first The height and width of the feature map at each scale.
[0101] Bounding box regression branch output:
[0102] ;
[0103] Four channels are used to represent the positional parameters of the candidate target bounding box.
[0104] Confidence branch output:
[0105] ;
[0106] in, Used to indicate the first The probability that a target exists at each candidate location at each scale.
[0107] By jointly predicting using the classification branch, bounding box regression branch, and confidence branch, the first... Candidate target detection results for frames of visible light images at different scales.
[0108] S2.5 Decoding and filtering of test results;
[0109] The classification results, bounding box regression results, and confidence results output by the detection heads at each scale are decoded to obtain candidate detection boxes.
[0110] For any candidate detection box, its detection confidence is expressed as:
[0111] ;
[0112] in, This indicates the probability of the target category corresponding to the candidate detection box. This indicates the probability that a target exists at the location of the candidate detection box.
[0113] Subsequently, the candidate detection boxes are filtered by confidence level, removing those with confidence levels below a preset threshold. Then, non-maximum suppression is applied to the remaining candidate detection boxes to remove duplicate detection boxes with high overlap, resulting in the [number of boxes]. Basic detection results of the visible light image frame:
[0114] ;
[0115] in, Indicates the first The number of targets detected in a frame of visible light image, the first The basic test results are represented as follows:
[0116] ;
[0117] in, Indicates the first The target category of each objective. Indicates the first The bounding box of each target. Indicates the first Target confidence level for each objective.
[0118] The basic test results Used as a reference for target detection in subsequent inter-frame compensation processes of the event stream.
[0119] S2.6 Detection information caching;
[0120] The multi-scale visible light fusion features obtained in step S2.3 are cached and denoted as:
[0121] ;
[0122] in, , , They represent the first Fusion features of visible light images at different scales, superscript This indicates that the feature originates from a visible light image.
[0123] At the same time, the basic detection results obtained in step S2.5 are cached as the first... Frame visible light detection status. For the first... For each target, its detection status is represented as follows:
[0124] ;
[0125] in, Indicates the first The basic categories of each objective Indicates the first The basic bounding box of each target Indicates the first The baseline confidence level of each objective Indicates the first The acquisition time of a frame of visible light image.
[0126] Therefore, the multi-scale visible light fusion features are thus... and basic test results Together as the first The visible light detection anchor points of the frames are cached. These visible light detection anchor points are used to provide target semantics, basic location, and confidence references for inter-frame compensation in subsequent event streams.
[0127] Step S3: Construct the inter-frame event representation of the current event window;
[0128] For the Frame visible light image With the Frame visible light image Update time of any event between Extract the current event window from the event stream. The event data within the frame is converted into an inter-frame event tensor or an inter-frame event feature.
[0129] in, Indicates the duration of the event update window. This indicates the current event update time. The inter-frame event representation corresponds only to the current event window and is used to characterize the brightness changes, edge changes, and motion changes that occurred within the most recent event update interval. It also serves as the input for subsequent event compensation weight calculation and visible light-event feature fusion.
[0130] S3.1 Set the event update time;
[0131] like Figure 2 As shown, in the first Frame visible light image With the Frame visible light image Between these, multiple event update times are set based on the collection timestamps of both events. Let the first... The acquisition time of the frame visible light image is , No. The acquisition time of the frame visible light image is The visible light frame interval between the two is:
[0132] ;
[0133] According to the visible light frame interval The event update interval is set based on the target's movement speed, event trigger density, and the system's real-time processing capability. .in, Less than or equal to the visible light frame interval .
[0134] In one specific embodiment, the time interval between two adjacent visible light images can be divided into: If there are multiple event update intervals, then the event update interval is:
[0135] ;
[0136] in, Indicates the first Frame visible light image and the first The number of event updates set between frames of visible light images.
[0137] Subsequently, according to the event update interval... Set the event update time:
[0138] ;
[0139] in, And satisfy:
[0140] ;
[0141] The event update time This is not the moment of acquiring a new visible light image, but rather the point in time where the inter-frame target detection state is updated using an event stream before the next visible light image arrives. By setting multiple event update moments, inter-frame target detection results can be continuously generated between two adjacent visible light image frames.
[0142] S3.2 Extract the event time window;
[0143] For any event update time ,by As the end time of the current event window, and based on the event update interval. The time range is extracted from the time-synchronized event stream as the duration of the current event window. From the event data within, obtain the event set corresponding to the current event window:
[0144] ;
[0145] in, This indicates the first time synchronization step in step S1. One event, This represents the corrected event timestamp. If the visible light camera and the event camera have already been synchronized via hardware triggering or a unified clock, then:
[0146] ;
[0147] The event set Indicates the current event update time. The event data generated by the event camera in the most recent event update window is used to construct the inter-frame event representation of the current event window.
[0148] In other embodiments, the duration of the current event window can also be set differently based on the target's movement speed, event trigger density, or the system's real-time processing requirements. The preset window length.
[0149] S3.3 Divide the event into time segments;
[0150] To preserve the chronological order of events within the current event window, the event time window is:
[0151] ;
[0152] According to the preset number of time segments Divided into There are several consecutive time segments. The length of each time segment is:
[0153] ;
[0154] No. Each time segment is represented as:
[0155] ;
[0156] in:
[0157] ;
[0158] For the last time segment, its right endpoint can be set to a closed interval to include timestamps equal to... The incident.
[0159] For each time segment From event set Extract events whose timestamps fall within that time segment to obtain the first... The subset of events corresponding to each time segment:
[0160] ;
[0161] in, Indicates the current event window A set of events within a time segment Indicates the first The event timestamp after time synchronization.
[0162] S3.4 Determine the channel based on the event polarity;
[0163] For each time segment The events are divided into a positive polarity event set and a negative polarity event set according to their polarity:
[0164] ;
[0165] ;
[0166] in, Indicates the first A set of positive polarity events within a time segment. Indicates the first A set of negative events within a time segment.
[0167] Therefore, each time segment corresponds to two event channels: a positive polarity channel and a negative polarity channel.
[0168] S3.5 Construct the inter-frame event tensor;
[0169] Based on the time segmentation results obtained in step S3.3 and the positive and negative polarity event sets obtained in step S3.4, construct the inter-frame event tensor.
[0170] Initialize the inter-frame event tensor:
[0171] ;
[0172] in, Indicates the number of time segments. and These represent the spatial height and spatial width of the event tensor, respectively. The inter-frame event tensor includes... One channel, of which the front One channel is used to record positive polarity events within different time segments, then... Each channel is used to record negative events within different time segments.
[0173] Specifically, for the first Positive events within a time segment are accumulated into the first frame event tensor. The first channel; for the first Negative events within a given time segment are accumulated into the first frame event tensor. One channel. Among them, .
[0174] For any given event, its spatial position in the corresponding channel is determined based on its uniform spatial coordinates. If multiple events correspond to the same time segment, the same event polarity, and the same spatial position, the number of events at that position is accumulated.
[0175] The resulting inter-frame event tensor contains information on the temporal, spatial, and polarity distributions of events, and can be used as input to subsequent event feature extraction networks.
[0176] S3.6 Extract inter-frame event features;
[0177] The inter-frame event tensor obtained in step S3.5 Input the event feature extraction network to obtain inter-frame event features:
[0178] ;
[0179] in, This represents an event feature extraction network. This indicates the event characteristics corresponding to the current event update time.
[0180] In one specific embodiment, the event feature extraction network adopts a lightweight convolutional network structure, including an input convolutional module and multiple event feature extraction stages. The input convolutional module is used to perform preliminary feature mapping on the inter-frame event tensor; the multiple event feature extraction stages are used to progressively extract edge change information, local motion change information, and temporal change information from the event tensor.
[0181] To correspond with the visible light multi-scale fusion features obtained in step S2, the event feature extraction network outputs event features at multiple scales:
[0182] ;
[0183] in, , , These represent the current event update time. Event features at different scales are used for subsequent fusion with visible light multi-scale features. , , Scale alignment and inter-frame compensation fusion are performed.
[0184] S4: Evaluate the effectiveness of the event window and generate event compensation weights;
[0185] Update time of the current event The corresponding event window is evaluated for validity, and event compensation weights are generated based on the evaluation results. Its value range is:
[0186] ;
[0187] The event compensation weight is used to characterize the reliability of event information within the current event window and to adjust the contribution strength of event features in subsequent inter-frame compensation fusion. When the event information within the event window is reliable, the event compensation weight is increased; when the number of events is insufficient, the event distribution is abnormal, or the polarity distribution is abnormal, the event compensation weight is decreased to reduce the impact of abnormal events on the inter-frame detection results.
[0188] S4.1 Count the number of events;
[0189] Statistics of current event window The total number of events within the range is denoted as: ;in, This indicates the number of events in the current event window.
[0190] The number of events Compare with the preset event quantity range. If If the number of events is less than the preset minimum event threshold, the validity evaluation of the current event window will be reduced; if If the number of events exceeds the preset maximum event threshold, it is determined that the current event window may have abnormal triggering or background disturbance, and its effectiveness evaluation is reduced.
[0191] S4.2 Statistical analysis of active pixel ratio;
[0192] The number of pixels within the current event window that have generated at least one event is recorded as: ;
[0193] Calculate the proportion of active pixels based on the spatial dimensions of the event tensor:
[0194] ;
[0195] in, and These represent the spatial height and width of the event tensor, respectively.
[0196] The active pixel ratio is used to represent the spatial distribution range of events within the current event window. If the number of events is large but the active pixel ratio is low, the effectiveness evaluation of the current event window is reduced to minimize the impact of concentrated triggering of local abnormal pixels on inter-frame detection results.
[0197] S4.3 Statistical polarity balance;
[0198] The number of positive and negative events in the current event window are counted and recorded as follows: and ;
[0199] Calculate the polarity balance based on the number of positive and negative polarity events:
[0200] ;
[0201] in, To prevent constants with a denominator of zero. Used to characterize the balance of positive and negative polarity events within the current event window.
[0202] like A lower value indicates a significant skew in the distribution of positive and negative polarity events within the current event window, thus reducing the effectiveness evaluation of the current event window and minimizing the impact of polarity anomalies on inter-frame compensation results.
[0203] S4.4 Generate event compensation weights;
[0204] Based on the number of events, the proportion of active pixels, and the polarity balance obtained in steps S4.1 to S4.3, the event compensation weight of the current event window is generated.
[0205] In one specific embodiment, the number of events, the proportion of active pixels, and the polarity balance are normalized to... The number of events is evaluated. Active pixel evaluation value and polarity evaluation value .
[0206] Then, the event compensation weight is calculated according to the following formula:
[0207] ;
[0208] in, The preset weighting coefficients satisfy:
[0209] ;
[0210] ;
[0211] Calculated event compensation weights Restricted to Within the range.
[0212] In another specific embodiment, event compensation weights can also be generated through preset threshold rules: when the number of events is within a preset range, the proportion of active pixels is greater than a preset threshold, and the polarity distribution does not show obvious abnormalities, the event compensation weights are increased; when the number of events is too small, the proportion of active pixels is too low, or the polarity distribution is obviously abnormal, the event compensation weights are decreased.
[0213] The event compensation weight is used in subsequent steps to adjust the contribution intensity of event features to cross-modal fusion features.
[0214] S5: Fusion of visible light anchor point features and event compensation features;
[0215] like Figure 3 As shown, for the current event update time The visible light detection anchor point features cached in step S2 are fused with the current event window features obtained in step S3, and the event compensation weights obtained in step S4 are used. By adjusting the contribution of event features, cross-modal compensation features are obtained.
[0216] The cross-modal compensation feature is used to combine the target semantic information of the visible light image and the inter-frame change information of the event stream to provide input for the subsequent generation of inter-frame target detection results.
[0217] S5.1 Read the visible light anchor point features and the current event window features;
[0218] Read the cached step S2 Frame-by-frame visible light multi-scale fusion features:
[0219] ;
[0220] in, , , They represent the first Visible light characteristics of a frame of visible light image at different scales.
[0221] At the same time, read the current event update time obtained in step S3. Corresponding multi-scale features of the event:
[0222] ;
[0223] in, , , These represent the event features of the current event window at different scales. The visible light features and event features mentioned above are used for subsequent same-scale alignment and compensation fusion.
[0224] S5.2 Align features of the same scale;
[0225] For the Several scales, among which Visible light characteristics Event characteristics Adjusting to the same spatial size and number of channels yields aligned visible light features. and aligned event features .
[0226] In one specific embodiment, spatial size alignment is achieved through upsampling, downsampling, or interpolation operations, while channel number alignment is achieved through... Convolution implementation.
[0227] S5.3 Adjusting event characteristics based on compensation weights;
[0228] Using the event compensation weights obtained in step S4 The aligned event features are weighted to obtain the event compensation features:
[0229] ;
[0230] in, Indicates the first Aligned event features at each scale This represents the weighted event compensation characteristics.
[0231] when When the value is large, enhance the contribution of event features in subsequent fusion; when When the value is small, the contribution of event features is reduced to minimize the impact of insufficient event information or anomalous events on inter-frame detection results.
[0232] S5.4 Segment and merge two types of features;
[0233] In the At each scale, the aligned visible light features With weighted event compensation features By concatenating the data along the channel dimension, we obtain the concatenated features:
[0234] ;
[0235] in, Indicates the first splicing features at various scales This indicates a channel dimension splicing operation.
[0236] Subsequently, the spliced features are input into the fusion convolution module to obtain the first... Cross-modal compensation features at various scales:
[0237] ;
[0238] in, This indicates a fusion convolutional module. The fusion convolutional module includes... convolution, Convolutional, normalization, and activation function layers are used to fuse visible light semantic information with event frame-to-frame change information.
[0239] Perform the above processing on all scales to obtain the cross-modal compensation features at the current event update time:
[0240] ;
[0241] in, , , These represent cross-modal compensation features at different scales, with superscripts indicating the specific features. This indicates the characteristics of compensation fusion.
[0242] S6: Generate and update the inter-frame target detection status;
[0243] Input the cross-modal compensation features obtained in step S5 into the inter-frame detection head to generate the current event update time. The inter-frame target detection results are cached as the current inter-frame detection state.
[0244] The inter-frame detection head is used to predict target category, target bounding box, and target confidence based on cross-modal compensation features. The inter-frame target detection result is detection data in a unified image coordinate system, rather than a newly generated visible light image or reconstructed image, and is used to represent the target detection state at the current event update time.
[0245] S6.1 Input cross-modal compensation features;
[0246] Input the cross-modal compensation features obtained in step S5 at the current event update time into the inter-frame detection head.
[0247] In one specific embodiment, the inter-frame detection head adopts the same or similar structure as the YOLO-type detection head in step S2, including a classification branch, a bounding box regression branch, and a confidence branch. The classification branch is used to predict the target category, the bounding box regression branch is used to predict the target bounding box parameters, and the confidence branch is used to predict the probability of the target's presence.
[0248] S6.2 Predict the detection result at the current event moment;
[0249] The inter-frame detection head outputs the current event update time based on the cross-modal compensation features. The target category probability, target bounding box parameters, and target confidence.
[0250] Among them, the target category probability is used to represent the category to which the target belongs, the target bounding box parameter is used to represent the position of the target in the unified image coordinate system, and the target confidence score is used to represent the degree of confidence in the existence of the target at the current event update time.
[0251] S6.3 Filter and output inter-frame detection results;
[0252] The candidate detection results output by the inter-frame detection header are subjected to confidence filtering and non-maximum suppression processing to remove low-confidence detection results and duplicate detection boxes, thus obtaining the current event update time. Inter-frame target detection results:
[0253] ;
[0254] in, This indicates the number of targets detected at the current event update time. The inter-frame target detection results are represented as follows:
[0255] ;
[0256] in, Indicates the target category, This represents the bounding box of the target in a unified image coordinate system. Indicates the confidence level of the target.
[0257] S6.4 Cache the current inter-frame detection state;
[0258] The inter-frame target detection results output at the current event update time. The current inter-frame detection state is cached.
[0259] The current inter-frame detection state is used for result association, continuous output, and stability filtering at subsequent event update times. When the next event update time arrives, steps S3 to S6 are re-executed based on the new event window to generate the inter-frame target detection results for the next event update time.
[0260] S7: Continuously output inter-frame detection results and refresh detection anchor points;
[0261] In the Frame visible light image With the Frame visible light image Between each event update, steps S3 to S6 are repeated to continuously output multiple inter-frame target detection results:
[0262] ;
[0263] The multiple inter-frame target detection results are used to represent the continuous detection status of the target between two adjacent visible light images.
[0264] When the Frame visible light image Upon arrival, step S2 is executed again to generate a new visible light detection anchor point, which is then used as the reference basis for subsequent inter-frame compensation of the event stream.
[0265] The inter-frame target detection results can be directly output to the target tracking, alarm, control, or decision-making modules.
Claims
1. A low-latency visible light target detection method based on event stream inter-frame compensation, characterized in that, Includes the following steps: Step S1: Construct a synchronous data unit for the visible light image sequence and the asynchronous event stream, acquire the visible light image sequence and the asynchronous event stream output by the event camera, and perform time synchronization and spatial coordinate alignment between the visible light image sequence and the asynchronous event stream; Step S2: Generate and cache visible light detection anchor points for the first... Frame visible light image Perform target detection to obtain basic detection results. and multi-scale visible light fusion features and the basic test results and multi-scale visible light fusion features As the first Buffering visible light detection anchor points for frames; Step S3: Construct the inter-frame event representation of the current event window, in the... Frame visible light image With the Frame visible light image Set multiple event update times between them Update time for any event Extract event data from the asynchronous event stream within the current event window and convert the event data into an inter-frame event tensor. Then, the multi-scale event features corresponding to the current event update time are obtained through the event feature extraction network. ; Step S4: Evaluate the effectiveness of the event window and generate event compensation weights for the current event update time. The corresponding event window is evaluated for validity, and event compensation weights are generated based on the evaluation results. ; Step S5: Fuse visible light detection anchor points and event compensation weights, and cache the multi-scale visible light fusion features from step S2. The multi-scale event features obtained in step S3 Perform same-scale alignment and utilize event-compensated weights. By adjusting the contribution of event features, cross-modal compensation features are obtained. ; Step S6: Generate and update the inter-frame target detection state, and apply the cross-modal compensation features obtained in step S5. Input the inter-frame detection header to generate the current event update time. Inter-frame target detection results And cache it as the current inter-frame detection state; Step S7: Continuously output inter-frame detection status and refresh visible light detection anchor points. Frame visible light image With the Frame visible light image Between these steps, steps S3 to S6 are repeated for each event update time, continuously outputting multiple inter-frame target detection states; when the first... Frame visible light image Upon arrival, step S2 is executed again to generate a new visible light detection anchor point, which is then used as the reference basis for subsequent inter-frame compensation of the event stream.
2. The low-latency visible light target detection method according to claim 1, characterized in that, In step S1, the visible light image sequence is represented as follows: ; Its corresponding timestamp is: ; in, This represents the total number of frames in a visible light image sequence. , Indicates the first The acquisition timestamp of a frame of visible light image, superscript This indicates that the timestamp belongs to a visible light image frame; The asynchronous event stream is represented as follows: ; in, This represents the total number of events in the event stream, with each event... Represented as: ; in, Indicates the first The spatial coordinates of an event in the event camera pixel plane Indicates the first The trigger timestamp of each event. Indicates the first The event polarity of an event, indicated by the superscript. This indicates that the timestamp belongs to the event stream data.
3. The low-latency visible light target detection method according to claim 2, characterized in that, In step S1, the time synchronization and spatial coordinate alignment specifically include: If there is a time offset between the visible light camera and the event camera Then for the first The trigger timestamp of each event Perform correction: ; The corrected event is: ; in, This indicates the trigger timestamp of the corrected event. This represents the time correction amount between the event camera's time reference and the visible light camera's time reference; if the visible light camera and the event camera have already been synchronized via hardware triggering or a unified clock, then let: ; If the imaging coordinate systems of the visible light camera and the event camera are different, the spatial mapping relationship between them is obtained through camera calibration, and the event coordinates are transformed to the visible light image coordinate system. Let the mapping function from event coordinates to visible light image coordinates be: ; in, This represents a coordinate mapping function determined by camera intrinsic parameters, extrinsic parameters, or homography transformation. Indicates the first Each event is mapped to its corresponding coordinates in the visible light image coordinate system; To facilitate unified processing later, the spatial coordinates of the event in a unified image coordinate system are defined as follows: ; If the event coordinates are already aligned with the visible light image coordinates, then: ; If the event coordinates need to be mapped to the visible light image coordinate system, then: ; In subsequent steps, coordinates are used. Construct an inter-frame event representation.
4. The low-latency visible light target detection method according to claim 1, characterized in that, In step S2, generating and caching visible light detection anchor points specifically includes: For the Frame visible light image The image is preprocessed by performing resizing, pixel normalization, and channel normalization. : ; in, and These represent the image height and image width input to the object detection network, respectively, and 3 represents the number of channels in the visible light image; Preprocessed visible light image Input the backbone of the visible light target detection network, extract image features at different spatial scales, and output feature maps at multiple scales, denoted as: ; in, Used to preserve target edge, contour, and location information; Used to characterize target structural information at medium scale; Deep semantic information used to characterize the target; Feature maps at multiple scales The input multi-scale feature fusion module performs upsampling, downsampling, feature concatenation, and convolutional fusion processing to obtain multi-scale visible light fusion features. : ; in, They represent the first Fusion features of visible light images at different scales, superscript This indicates that the feature originates from a visible light image; Inputting multi-scale visible light fusion features into the detection head, predicting the [missing information]. Target category, target bounding box, and target confidence in a frame of visible light image; for the 1st frame... Visible light fusion features at various scales ,in The classification branch outputs the following features: ; Bounding box regression branch output features: ; Confidence branch output features: ; in, Indicates the number of target categories. and They represent the first The height and width of the feature map at each scale; The classification results, bounding box regression results, and confidence results output by the detection heads at each scale are decoded to obtain candidate detection boxes; for any candidate detection box, its detection confidence is expressed as: ; in, This indicates the probability of the target category corresponding to the candidate detection box. This indicates the probability that a target exists at the location of the candidate detection box; After performing confidence filtering and non-maximum suppression on the candidate detection boxes, the first one is obtained. Basic detection results of visible light images : ; in, Indicates the first The number of targets detected in a frame of visible light image, the first Basic test results Represented as: ; in, Indicates the first The target category of each objective. Indicates the first The bounding box of each target. Indicates the first Target confidence level for each objective; The multi-scale visible light fusion features and basic test results Together as the first The visible light detection anchor points of the frames are cached.
5. The low-latency visible light target detection method according to claim 1, characterized in that, In step S3, the event update time is set as follows: In the Frame visible light image With the Frame visible light image Between them, multiple event update times are set based on the collection timestamps of the two events, let the first one be... The acquisition time of the frame visible light image is , No. The acquisition time of the frame visible light image is The visible light frame interval between the two is: ; According to the visible light frame interval The event update interval is set based on the target's movement speed, event trigger density, and the system's real-time processing capability. ,in Less than or equal to the visible light frame interval ; The time interval between two adjacent visible light images is divided into... If there are multiple event update intervals, then the event update interval is: ; in, Indicates the first Frame visible light image and the first The number of event updates set between frames of visible light images; Subsequently, according to the event update interval... Set the event update time: ; in, And satisfy: ; The event update time This refers to the time point at which the inter-frame target detection state is updated using the event stream before the next visible light image arrives.
6. The low-latency visible light target detection method according to claim 5, characterized in that, In step S3, constructing the inter-frame event tensor of the current event window specifically includes: For any event update time ,by As the end time of the current event window, and based on the event update interval. The time range is extracted from the time-synchronized event stream as the duration of the current event window. From the event data within, obtain the event set corresponding to the current event window: ; in, This indicates the first time synchronization step in step S1. One event, Indicates the corrected event timestamp; Event time window According to the preset number of time segments Divided into There are 3 consecutive time segments, each with a length of: ; No. Each time segment is represented as: ; in: ; For each time segment From the event set Extract events whose timestamps fall within that time segment to obtain the first... The subset of events corresponding to each time segment: ; For each time segment The events are divided into a positive polarity event set and a negative polarity event set according to their polarity: ; ; Based on the time segmentation results and the sets of positive and negative polarity events, initialize and construct the inter-frame event tensor: ; in, Indicates the number of time segments. and These represent the spatial height and spatial width of the event tensor, respectively; the inter-frame event tensor includes... One channel, of which the front One channel is used to record positive polarity events within different time segments, then... Each channel is used to record negative polarity events within different time segments; For the Positive events within a time segment are accumulated into the first frame event tensor. The first channel; for the first Negative events within a given time segment are accumulated into the first frame event tensor. One channel; among them ; For any event, its spatial position in the corresponding channel is determined based on its unified spatial coordinates. If multiple events correspond to the same time segment, the same event polarity, and the same spatial position, the number of events at that position is accumulated.
7. The low-latency visible light target detection method according to claim 6, characterized in that, In step S3, extracting inter-frame event features specifically includes: The inter-frame event tensor obtained in step S3 Input the event feature extraction network to obtain inter-frame event features: ; in, This represents an event feature extraction network. This indicates the event characteristics corresponding to the current event update time; The event feature extraction network adopts a lightweight convolutional network structure, including an input convolutional module and multiple event feature extraction stages; wherein, the input convolutional module is used to perform preliminary feature mapping on the inter-frame event tensor, and the multiple event feature extraction stages are used to extract edge change information, local motion change information and temporal change information in the event tensor step by step; The event feature extraction network outputs event features at multiple scales: ; in, , , These represent the current event update time. Event features at different scales are used for subsequent fusion with visible light multi-scale features. , , Scale alignment and inter-frame compensation fusion are performed.
8. The low-latency visible light target detection method according to claim 7, characterized in that, In step S4, the event compensation weights are generated as follows: Statistics of current event window The total number of events within the range is denoted as: This indicates the number of events in the current event window; The number of pixels within the current event window that have generated at least one event is recorded as: ; Calculate the proportion of active pixels based on the spatial dimensions of the event tensor: ; in, and These represent the spatial height and width of the event tensor, respectively; Count the number of positive and negative events in the current event window, and record them as follows: and The polarity balance is calculated based on the number of positive and negative polarity events: ; in, To prevent constants with a denominator of zero, Used to characterize the balance of positive and negative polarity events within the current event window; The number of events, the proportion of active pixels, and the polarity balance were normalized to [values to be filled in]. The number of events is evaluated. Active pixel evaluation value and polarity evaluation value Then, the event compensation weight is calculated according to the following formula: ; in, The preset weighting coefficients satisfy: ; ; Calculated event compensation weights Restricted to Within the range.
9. The low-latency visible light target detection method according to claim 7 or 8, characterized in that, In step S5, the fusion of visible light detection anchor points and event compensation weights specifically includes: Read the cached step S2 Frame-by-frame visible light multi-scale fusion features: ; in, , , They represent the first Visible light characteristics of a frame of visible light image at different scales; Simultaneously read the current event update time obtained in step S3. Corresponding multi-scale features of the event: ; in, , , These represent the event characteristics of the current event window at different scales; For the There are several scales, among which Visible light characteristics Event characteristics Adjusting to the same spatial size and number of channels yields aligned visible light features. and aligned event features ; Using the event compensation weights obtained in step S4 The aligned event features are weighted to obtain the event compensation features: ; in, Indicates the first Aligned event features at each scale This represents the weighted event compensation characteristics; In the At each scale, the aligned visible light features With weighted event compensation features By concatenating the data along the channel dimension, we obtain the concatenated features: ; in, Indicates the first splicing features at various scales This indicates a channel-level concatenation operation; The concatenated features are input into the fusion convolution module to obtain the first... Cross-modal compensation features at various scales: ; in, This indicates a fusion convolution module, which includes... convolution, Convolutional, normalization, and activation function layers are used to fuse visible light semantic information with event frame change information; Perform the above processing on all scales to obtain the cross-modal compensation features at the current event update time: ; in, , , These represent cross-modal compensation features at different scales, with superscripts indicating the specific features. This indicates the characteristics of compensation fusion.
10. The low-latency visible light target detection method according to claim 9, characterized in that, In steps S6 and S7, generating and outputting the inter-frame target detection results specifically includes: The cross-modal compensation features obtained in step S5 Input the inter-frame detection header to generate the current event update time. Inter-frame target detection results: ; in, This represents the number of targets detected at the current event update time, the th... The inter-frame target detection results are represented as follows: ; in, Indicates the target category. This represents the bounding box of the target in a unified image coordinate system. Indicates the confidence level of the target; The inter-frame target detection results output at the current event update time. Cache the current inter-frame detection state; In the Frame visible light image With the Frame visible light image Between each event update, steps S3 to S6 are repeated to continuously output multiple inter-frame target detection results: ; The multiple inter-frame target detection results are used to represent the continuous detection status of the target between two adjacent visible light images; When the Frame visible light image Upon arrival, step S2 is executed again to generate a new visible light detection anchor point, which is then used as the reference basis for subsequent inter-frame compensation of the event stream. The inter-frame target detection result is the detection state data under a unified image coordinate system, rather than a newly generated visible light image or a visible light image reconstructed from an event stream. The inter-frame target detection result can be directly output to the target tracking, alarm, control, or decision-making module.
Citation Information
Patent Citations
Adaptive target detection method, system and equipment based on event camera
CN115496920A
Target detection method based on computer vision
CN121033508A