Triggered vehicle recognition and tracking method and device based on radar and vision fusion and readable storage medium thereof

By employing a transient asynchronous fusion mechanism and adaptive weight allocation, the problems of time mismatch and wasted computing power in radar and video fusion are solved, enabling efficient vehicle recognition and tracking, and adapting to complex scenarios and edge deployments.

CN122176662AActive Publication Date: 2026-06-09HANGZHOU FEELING TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU FEELING TECH CO LTD
Filing Date
2026-05-11
Publication Date
2026-06-09

Smart Images

  • Figure CN122176662A_ABST
    Figure CN122176662A_ABST
Patent Text Reader

Abstract

The application provides a trigger type vehicle identification and tracking method and device based on radar and vision fusion and a readable storage medium thereof, and belongs to the field of intelligent traffic perception. In view of the problems of time mismatch of existing radar and vision fusion and high consumption of full-time calculation power, the application acquires bimodal data; video frames are used as time anchor points, and the extracted video and radar features are fused to initialize radar and vision fusion transient features; within the video frame interval, the transient features are updated by using the high-frequency features of the radar; in the fusion update, the bimodal weight is adaptively distributed based on the local spatiotemporal feature change; the transient features are monitored, and the vehicle trajectory optimization is started as needed when the trigger condition is met. The application solves the problem of fusion of time-series heterogeneous features, greatly reduces the power redundancy, and is mainly used for accurate vehicle identification and tracking in the automatic driving and roadside monitoring scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation vehicle recognition technology, and in particular to a trigger-based vehicle recognition and tracking method, device and its readable storage medium based on radar-visual fusion, applicable to scenarios such as roadside traffic monitoring and autonomous driving environmental perception. Background Technology

[0002] Vehicle recognition and tracking are core perception tasks in intelligent transportation systems and autonomous driving. Currently, mainstream perception technologies fall into two main categories: radar perception and video perception, which are complementary. Radar perception relies on the principle of electromagnetic wave reflection and can operate stably in adverse environments such as nighttime, rain, and fog. It can accurately acquire motion features such as the relative distance, speed, and azimuth of a target. However, radar point cloud data is sparse and lacks semantic details such as vehicle appearance, texture, license plate, and model, making it difficult to achieve refined recognition. Video perception relies on optical cameras to collect image information and can capture rich semantic features of appearance. Combined with deep learning algorithms, it can achieve refined vehicle classification and feature matching. However, video perception is significantly affected by environmental factors such as lighting and weather. Image quality degrades severely in low-light, strong light, backlight, rain, and fog scenarios, and the limited frame rate of conventional videos makes it difficult to capture the fine dynamics of high-speed vehicles.

[0003] To fully leverage the complementary advantages of the two modalities, existing technologies generally employ radar-video fusion solutions. Among these, synchronous fusion is the most widely used approach. Its core principle is to unify the feature data from radar and video to the same time scale before fusing them. This primarily includes two implementation methods:

[0004] The first step is to downsample the high-frequency features of the radar to match the video frame rate, and the second step is to upsample the low-frequency features of the video to the radar frame rate through interpolation.

[0005] However, the aforementioned synchronous fusion strategies have inherent flaws: radar feature downsampling loses high-frequency temporal dynamic information between frames, failing to capture fine motion changes such as vehicle acceleration / deceleration and rapid lane changes; video feature upsampling results in spatial semantic feature distortion and damages spatial consistency because the interpolated video features are not based on real-world data. Furthermore, existing radar-visual fusion methods often employ globally fixed weights or simple rule-based fusion strategies, failing to adaptively adjust fusion weights according to changes in spatial regions and temporal scenarios. This leads to unreliable modal features interfering with the fusion results, resulting in insufficient fusion robustness. Simultaneously, existing solutions generally employ a full-time, indiscriminate computation mode, continuously performing feature extraction and fusion calculations regardless of whether critical states requiring high-precision perception exist within the monitoring scene, resulting in wasted computing resources and failing to meet the lightweight deployment requirements at the edge.

[0006] Therefore, there is an urgent need for a trigger-based vehicle identification and tracking method, device, and readable storage medium based on radar-visual fusion to solve the problems existing in the prior art. Summary of the Invention

[0007] This invention provides a trigger-based vehicle recognition and tracking method, device, and readable storage medium based on radar-visual fusion. It addresses the time mismatch issues of existing synchronous fusion strategies, such as the loss of high-frequency dynamic information from radar or distortion of video spatial semantic features due to forced alignment of time scales, the inability to adaptively adjust fusion weights according to spatial location and temporal scene to cope with dynamic changes in modal reliability, and the waste of computing power caused by the all-time indiscriminate calculation mode.

[0008] The core technology of this invention mainly adopts a radar-visual transient asynchronous fusion mechanism, which initializes the fusion features with video frames as time anchors, performs pure time-series recursive updates driven by radar data within the video frame interval, and adaptively allocates fusion weights based on changes in local spatial and temporal domain features during the fusion and update process. At the same time, a trigger-based computing mechanism is introduced to achieve on-demand recognition and trajectory optimization.

[0009] In a first aspect, the present invention provides a trigger-based vehicle identification and tracking method based on radar-visual fusion, the method comprising the following steps: Acquire synchronized radar point cloud data and video image data, and preprocess the radar point cloud data and video image data to obtain radar motion feature sequence and video semantic feature sequence; Using the arrival time of the video frame as the time anchor, cross-modal fusion is performed based on the current video frame features and the radar motion features within the corresponding time window to generate the initial radar-visual fusion transient features. Within the time interval between adjacent video frames, with the radar data acquisition cycle as the step size, a purely radar-driven temporal recursive update is performed based on the radar-visual fusion transient features of the previous moment and the radar motion features acquired at the current moment, to obtain the updated radar-visual fusion transient features on the continuous time axis. In the process of cross-modal fusion and temporal recursive update, the fusion weights are adaptively allocated based on the feature changes of radar mode and video mode in the local spatial domain and temporal domain. Based on the change state triggering trajectory optimization of the radar-visual fusion transient features, when the preset triggering conditions are met, the vehicle target is identified and the tracking trajectory is optimized.

[0010] Furthermore, using the arrival time of the video frame as the time anchor, cross-modal fusion is performed based on the features of the current video frame and the radar motion features within the corresponding time window to generate initialized radar-visual fusion transient features, including: The radar motion features within the time window are densified, and the densified radar motion features and the current video frame features are encoded into a unified latent state space respectively. Local adaptive weighted fusion is performed on the encoded radar features and video features to obtain the initialized radar-video fused transient features.

[0011] Furthermore, performing a purely radar-driven timing recursive update includes: The transient features of the radar-visual fusion from the previous moment are used as the query quantity, and the radar motion features acquired at the current moment are encoded and used as the key and value quantities. By using a temporal feature updater to perform cross-attention calculations on query volume, key value volume, and numerical value volume, and injecting the motion dynamic features of the current moment while preserving the consistency of historical semantic features, the updated radar-visual fusion transient features are obtained.

[0012] Furthermore, based on the feature changes of radar modes and video modes in the local spatial and temporal domains, fusion weights are adaptively assigned, including: Based on the spatial geometric distribution relationship between the target node and its local spatial neighboring nodes, determine the spatial domain attention weight; The temporal attention weights are determined based on the rate of change of features between adjacent frames within a local time window. By combining spatial domain attention weights and temporal domain attention weights, dynamic weighting coefficients for cross-modal fusion are generated.

[0013] Furthermore, based on the change state triggering trajectory optimization of the transient features of radar-visual fusion, when the preset triggering conditions are met, the vehicle target is identified and the tracking trajectory is optimized, including at least one of the following: When the instantaneous velocity of a target detected by radar point cloud data reaches a preset velocity threshold and the duration reaches a first preset duration, the triggering condition is determined to be met. When a new target is detected entering the preset monitoring area, and there is no matching identifier within the preset number of historical frames and the confidence level is greater than the preset confidence threshold, the triggering condition is determined to be met. When the intersection-union ratio of the target bounding box is less than the first preset value and the similarity of appearance features is less than the second preset value and continues for a preset number of frames, it is determined that the target is occluded or deoccluded and the triggering condition is met. When the error between the predicted trajectory position and the actual tracking position is greater than or equal to a preset error threshold, and the duration reaches a second preset duration, the trigger condition is determined to be met.

[0014] Furthermore, vehicle target identification and tracking trajectory optimization are performed, including: Constructing a multi-scale feature pyramid based on transient features fused from radar and vision; The multi-scale feature pyramid is input into the trajectory optimization network, and the position, motion state and appearance features of the vehicle target are corrected through an iterative cross-attention mechanism, and the optimized vehicle recognition result and tracking trajectory are output.

[0015] Furthermore, the method also includes a step of training a neural network model for achieving cross-modal fusion and temporal recursive updates, the training step including: Construct a total loss function that includes a detection loss term, a trajectory regression loss term, and a gradient constraint loss term; The gradients of the video branch and the radar branch are calculated based on the total loss function, and the network parameters are updated by constraining the directional consistency between the gradients of the video branch and the radar branch.

[0016] Secondly, the present invention provides a trigger-based vehicle recognition and tracking device based on radar-visual fusion, comprising: The data preprocessing module is used to acquire synchronized radar point cloud data and video image data, and extract radar motion feature sequences and video semantic feature sequences. The transient asynchronous fusion module is used to initialize the transient features of radar-visual fusion with video frames as time anchors, and to perform pure radar-driven time-series recursive updates with radar cycles as steps during the interval between adjacent video frames. The dynamic weight allocation module is used to adaptively allocate fusion weights based on the feature changes of radar mode and video mode in the local spatial domain and temporal domain. The trigger-based optimization module is used to make trigger decisions based on the changing state of the transient features of the radar-visual fusion, and to start vehicle recognition and trajectory tracking optimization when the preset trigger conditions are met.

[0017] Thirdly, the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to execute the above-described trigger-based vehicle recognition and tracking method based on radar-visual fusion.

[0018] Fourthly, the present invention provides a readable storage medium storing a computer program, the computer program including program code for controlling a process to execute the process, the process including the above-described trigger-based vehicle identification and tracking method based on radar-visual fusion.

[0019] The main contributions and innovations of this invention are as follows: 1. This invention uses a radar-visual transient asynchronous fusion mechanism, taking video frames as the initialization reference, and performs a pure radar-driven time-series recursive update with the radar cycle as the step size within the interval between adjacent video frames. It does not interpolate or reuse the video, nor does it downsample the radar, thus fundamentally solving the problem of time-frequency mismatch between radar and video. It realizes a continuously evolving super-video frame rate vehicle feature representation on the physical time axis, significantly improving the continuity and accuracy of vehicle tracking in scenarios such as high-speed driving and rapid lane changes.

[0020] 2. This invention utilizes an adaptive weight allocation mechanism based on changes in local spatial and temporal features to dynamically allocate fusion weights for radar and video in different spatial regions and time scenarios. This enables the fusion process to perceive real-time changes in modal reliability, automatically enhances the contribution of radar features in scenarios where video features degrade due to vehicle occlusion, low light, rain, or fog, and fully leverages the advantages of video semantic features in static or low-speed scenarios. This significantly improves the robustness of fused features and the accuracy of vehicle recognition in complex road scenarios.

[0021] 3. This invention utilizes multi-dimensional triggering conditions based on the transient feature changes of radar-visual fusion to initiate high-precision identification and trajectory optimization processes only when critical events such as the detection of high-speed targets, the entry of new targets, occlusion de-occlusion, or excessive trajectory prediction errors are detected. During non-critical periods, it maintains a low-power standby state, realizing an on-demand computing power scheduling mechanism. This effectively reduces the average computing power consumption and enables the solution to adapt to deployment environments with limited computing power at the edge.

[0022] 4. This invention introduces a gradient constraint loss term during the model training phase to constrain the direction consistency of the gradients of the video branch and the radar branch, effectively alleviating the gradient conflict and negative transfer problem in the joint training of heterogeneous modalities. This ensures the fusion effect from the training mechanism and further improves the convergence stability and generalization ability of the model.

[0023] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description

[0024] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart of a trigger-based vehicle identification and tracking method based on radar-visual fusion according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0026] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.

[0027] This invention provides a trigger-based vehicle recognition and tracking method and apparatus based on radar-eye fusion, aiming to solve the technical problems of time mismatch, inability to adaptively adjust modal fusion weights, and computational redundancy in existing radar-eye fusion schemes. The following refers to... Figure 1 The technical solution of the present invention will be described in detail below. Figure 1 This is a schematic diagram of the algorithm logic of the radar-visual fusion triggered high-efficiency vehicle recognition technology according to an embodiment of the present invention.

[0028] Example 1 This embodiment provides a trigger-based vehicle recognition and tracking method based on radar-visual fusion. For example... Figure 1 As shown, the method mainly includes the following steps: Step 1: Multi-source data acquisition and preprocessing This step is used to acquire synchronized radar point cloud data and video image data, and to preprocess the above data to obtain radar motion feature sequences and video semantic feature sequences.

[0029] Specifically, millimeter-wave radar and high-definition cameras are deployed on the roadside or in vehicles. The millimeter-wave radar acquires 3D point cloud data of the target at a first frequency (e.g., 60Hz), with a single frame acquisition period of approximately 16.7ms; the camera acquires RGB image data at a second frequency (e.g., 25Hz), with a single frame acquisition period of 40ms. The first frequency is higher than the second frequency. In another implementation scenario, the first frequency can also be 30Hz or 77Hz, and the second frequency can also be 30Hz or 50Hz, as long as the first frequency is higher than the second frequency.

[0030] To achieve high-precision synchronization of dual-modal data, a dual synchronization mechanism combining hardware external triggering and software timestamp calibration is adopted. Specifically, a sliding window of length W frames (e.g., W=20) is constructed, and hardware trigger timestamps and software timestamps of multiple consecutive frames are collected. A clock drift model is fitted using the least squares method to eliminate system errors. This model can be expressed as:

[0031] in, For the corrected software timestamp (the term "software timestamp" has been defined above); 'Hardware trigger timestamp' (as defined above); 'a' is the drift coefficient (slope), and 'b' is the fixed delay (intercept). Linear interpolation is used for subframe alignment of high-frequency radar frames to ensure synchronization error is less than 5ms.

[0032] For radar point cloud data, denoising, dimensionality reduction, clustering fitting, and feature encoding are performed sequentially. The denoising operation uses signal-to-noise ratio (SNR) threshold filtering combined with a density-based spatial clustering algorithm (DBSCAN) to remove ground reflection clutter, guardrail interference points, and isolated discrete points. The dimensionality reduction operation projects the 3D point cloud onto a bird's-eye view (BEV) perspective, converting it into a 2D point cloud and eliminating height dimensional redundancy. The clustering fitting operation uses the K-means clustering algorithm to segment the effective point cloud into vehicle targets and combines it with least squares to fit vehicle contours, extracting point cloud clusters for individual vehicles. The feature encoding operation extracts motion features such as contour size, relative distance, driving speed, azimuth angle, and rate of change of speed for each vehicle point cloud cluster, and serializes them according to timestamps to generate a high-frequency radar motion feature sequence with the first frequency.

[0033] For video image data, a target detection algorithm (such as YOLOv8 or other convolutional neural network-based target detectors) is used to locate vehicles in each frame of the image, outputting the bounding box coordinates, confidence score, and category label (0-small car, 1-truck, 2-trailer, 3-bus) of the vehicle target, filtering out invalid targets with a confidence score below 0.5. The local region of the vehicle target is cropped based on the bounding box and scaled to a uniform scale (e.g., 512×512 pixels). Deep visual features, including spatial semantic features such as appearance texture, vehicle type semantics, body color, and local contours, are extracted from the backbone network of the target detection network and serialized according to the timestamp to generate a second-frequency low-frequency video semantic feature sequence V.

[0034] Finally, using the timestamp of the video frame as the reference anchor point, the radar motion feature sequence and the video semantic feature sequence V are precisely aligned through time window matching to form a feature pairing relationship between a single video frame and multiple radar frames, providing a data foundation for subsequent asynchronous fusion.

[0035] Step 2: Initialization of transient features of radar-visual fusion When each video frame arrives, the system uses the arrival time of that video frame as the time anchor point, and performs cross-modal fusion based on the features of the current video frame and the radar motion features within the corresponding time window to generate the initialized radar-visual fusion transient features.

[0036] Specifically, a radar feature matching window centered on the current video frame timestamp is determined. The window range is, for example, Extract all radar feature frames that fall within this window from the radar motion feature sequence.

[0037] The radar motion features within the window are densified to obtain dense radar features that match the spatial dimension of the video semantic features. The densification process is implemented as follows: First, the sparse point cloud features within the window are mapped to a predefined empty feature map with the same spatial resolution as the video feature map, based on their two-dimensional BEV coordinates. Each grid cell hit by the point cloud is assigned a corresponding kinematic feature vector (such as radial velocity, RCS value, etc.), while unhitted grid cells are assigned zero values. Then, a small-scale Gaussian smoothing operation (e.g., a Gaussian kernel size of 3×3) is performed on the feature map, allowing the kinematic information of a single radar point to gently diffuse to its neighboring grids, thus obtaining a denser radar motion feature map.

[0038] The dense radar features are encoded using a first feature encoder (e.g., a lightweight convolutional neural network) to obtain high-dimensional radar feature tokens. The current video frame features are then encoded using a second feature encoder (e.g., a network with a structure symmetrical to the first feature encoder) to obtain high-dimensional video feature tokens. The first and second feature encoders respectively map the features of heterogeneous modalities to a unified latent state space to facilitate subsequent fusion computation.

[0039] Subsequently, based on a local spatial-temporal dual-domain attention mechanism, local adaptive weighted fusion is performed on the radar high-dimensional features and video high-dimensional features to obtain the initialized radar-video fusion transient features. This fusion process can be completed by a cross-modal local weighted fusion operator, which calculates attention weights only within the local spatial region without performing global fusion, thereby reducing computational overhead and maintaining spatial specificity. The initialized radar-video fusion transient features can be expressed as:

[0040] in, The transient fusion feature corresponding to the t-th video frame; For video feature encoders; Let t be the set of target features in the t-th video frame; For radar feature encoders; Denseization operator; R is the radar point cloud feature sequence (denoted as...) ,in (The set of vehicle features at the nth radar time step). Let be the radar feature matching window for the t-th video frame. This is the fusion operator for the RV-CLWF module, which performs adaptive weighting only in the local space and does not perform global fusion.

[0041] Step 3: Timing-based recursive update of inter-frame radar-driven systems Within the time interval between adjacent video frames (e.g., 40ms), the system performs a pure radar-driven time-series recursive update based on the radar data acquisition cycle (e.g., 16.7ms) and the radar motion features acquired at the current time, with a step size of radar data acquisition cycle (e.g., 16.7ms), to obtain the updated radar-visual fusion transient feature sequence on the continuous time axis.

[0042] Specifically, when the When a radar feature frame arrives (where t is the video frame index and Δ is the radar time step), Let Δ represent the radar sampling time after the t-th video frame. First, the radar motion features of this frame are processed by densification and feature encoding as described in step two to obtain the radar incremental features for the current step size. Then, the radar features of the previous time (i.e., the Δ-th time) are processed by Δ. Transient characteristics of radar-visual fusion (each radar time step) As a query, the aforementioned radar incremental features are used as key and value values, and input into the temporal feature updater U (based on cross-attention (Transformer)). The temporal feature updater U is calculated based on the cross-attention mechanism, and its core operation can be represented as:

[0043] The updater internally fuses feature tokens for each spatial location. The update calculation can be further expressed as:

[0044] in, This refers to the transient fusion feature token for the j-th spatial location after the update. The transient fusion feature token from the previous radar moment (as a query) is used. Let N(j) be the value (as key and value) of the feature vector at the i-th position after densification and encoding of the current radar frame, and let N(j) represent the local spatial neighborhood at the j-th position. The local spatial-temporal dual-domain attention weights described in step four.

[0045] The above update process, while preserving the consistency of historical vehicle semantic features and spatial contours, injects the dynamic motion features of the current radar moment (such as instantaneous speed changes and distance displacement) into the transient features, outputting the updated radar-visual fusion transient features. Within a 40ms video frame interval, the 60Hz radar can complete two updates (e.g., corresponding to...). and (Two time points), thus achieving continuous vehicle feature evolution at 60Hz on the physical timeline, solving the problem of motion feature breakage caused by low video frame rate.

[0046] Preferably, this step further includes a recursive error correction mechanism: within the time interval between adjacent video frames, using the radar data acquisition cycle as the step size, based on the radar-visual fusion transient features of the previous moment and the radar motion features acquired at the current moment, a purely radar-driven temporal recursive update is performed, and recursive error detection and hierarchical correction are performed in real time. The specific steps are as follows: 1. The temporal recursive update uses the transient features of the radar-visual fusion from the previous moment as the query value, and encodes the radar motion features of the current moment as the key and value values. These are then input into the temporal feature updater for cross-attention calculation. While preserving the consistency of historical semantic features, the current motion dynamic features are injected to obtain the preliminary update result.

[0047] 2. Recursive error detection: Calculate in real time after each recursion. Position recursion error: the deviation of the recursive position from the reliable anchor point position using Euclidean distance; Velocity recursion error: the difference between the recursive velocity and the radar-measured velocity; Cumulative recursion error: the sum of squared errors from multiple consecutive recursions.

[0048] 3. Graded correction of recursive error: Minor errors (less than 2 pixels): Lightweight smoothing correction using Kalman filtering without interrupting the recursive process; Moderate error (2-6 pixels): A local feature rollback strategy is adopted, and the reliable fused features of the first 3 frames are reused for re-inference; Severe error (greater than 6 pixels): Immediately stop recursion, forcibly trigger video anchor point reinitialization, reconstruct fusion features using the video frame at the current moment, and clear the accumulated error.

[0049] Step 4: Adaptive allocation of cross-modal fusion weights Both the cross-modal fusion initialization in step two and the temporal recursive update in step three involve the allocation of fusion weights between radar and video modes. This invention adaptively allocates fusion weights based on the characteristic changes of radar and video modes in the local spatial and temporal domains.

[0050] Specifically, for any spatial location node j and its neighboring nodes i, the cross-modal attention weights Spatial domain attention weights and temporal attention weights Joint decision:

[0051] 1) Determination of spatial domain attention weights: Based on the spatial geometric distribution relationship between the target node i and its local spatial neighbor nodes j, the spatial location feature mapping function is used. The spatial correlation between nodes is calculated, and the spatial domain local attention weights are obtained through exponential normalization, where... , Let be the two-dimensional coordinates of the node in the image space. This weight reflects the strength of the association between features at different spatial locations.

[0052] 2) Determination of temporal attention weights: based on the feature changes between adjacent frames within a local time window T (i.e., ... ), through time-series feature change mapping function Temporal correlations are calculated and then exponentially normalized to obtain temporal attention weights. These weights reflect the degree of change in features over time.

[0053] Finally, the spatial domain attention weights are multiplied by the temporal domain attention weights to generate dynamic weighting coefficients for cross-modal fusion. The dynamic weighting coefficients are updated in real time as spatial location and temporal scene changes, enabling the fusion process to automatically perceive the reliability differences of different modalities in different regions and at different times. For example, in areas with dense vehicle occlusion, video features become ineffective due to occlusion. In this case, the temporal attention weights will detect the anomaly in the changes of video features, thereby reducing the contribution of the video modality and automatically increasing the fusion weight of radar features. On the other hand, in static scenes with good lighting and sparse vehicles, the reliability of video semantic features is high, while the radar point cloud is sparse. In this case, the weight allocation will be tilted towards the video modality.

[0054] Step 5: Determining the Triggering Condition The system makes trigger decisions based on the changing state of transient features of radar-visual fusion. When the preset trigger conditions are met, it initiates high-precision vehicle recognition and trajectory tracking optimization.

[0055] The preset trigger conditions include at least one of the following: 1) High-speed target triggering: When the instantaneous velocity of a target detected by radar point cloud data reaches a preset velocity threshold and the duration reaches a first preset duration, the triggering condition is determined to be met. In this embodiment, the preset velocity threshold... The speed limit is set to 60 km / h (corresponding to a typical highway speed limit), but this threshold can be adjusted to 40 km / h in urban road scenarios. The instantaneous velocity of the target is calculated from the radar Doppler velocity through coordinate projection. Let the velocity projection components of the target on the x-axis and y-axis in the Cartesian coordinate system be V... x and V y When the target's instantaneous velocity amplitude And two consecutive radar time steps (approximately 16.7) This condition is triggered when the overspeed is maintained (2=33.4ms).

[0056] 2) New Target Trigger: When a new target is detected entering the preset monitoring area, and there is no matching identifier within a preset number of historical frames (e.g., 5 frames), and the target confidence level is greater than a preset confidence threshold (e.g., 0.5), the trigger condition is met. The criteria for determining a new target are that the target's center point coordinates appear in the monitoring area for the first time and cannot be associated with any target in the existing tracking list.

[0057] 3) Occlusion and De-occlusion Triggering: Joint decision is made based on the intersection-over-union (IoU) ratio of the target bounding box and the cosine similarity (Sim) of the appearance features. When the IoU ratio is less than 0.3 and the appearance feature similarity Sim is less than 0.65, and this continues for a preset number of frames (e.g., 3 frames), the target is determined to be occluded and the triggering condition is met. When the IoU ratio recovers to above 0.5 or the similarity recovers to above 0.75, the target is determined to be de-occluded and the triggering condition is met.

[0058] 4) Trajectory Prediction Error Exceeding Threat: When the Euclidean distance error between the predicted trajectory position and the actual tracking position is greater than or equal to a preset error threshold, and the duration reaches a second preset duration, the trigger condition is determined to be met. In this embodiment, the preset error threshold is set to 6 pixels (corresponding to a physical deviation of approximately 0.3 meters in image space), and the second preset duration is 2 consecutive frames. The Euclidean position error calculation formula is:

[0059] in, The x-coordinate of the vehicle center predicted by the model using the Transformer; The model predicts the vehicle center ordinate using the Transformer; The x-coordinate of the vehicle tracking center is the actual output of the radar-view fusion (RTAF). This represents the vertical coordinate of the vehicle tracking center, as actually output by the radar-visual fusion (RTAF).

[0060] Preferably, the concurrent processing rules for multiple triggering conditions are as follows: When two or more triggering conditions are met simultaneously, the following priority is applied: occlusion trigger > trajectory prediction error exceeding the standard trigger > high-speed target trigger > new target entry trigger. Only the highest priority triggering task is executed at any given time, and lower priority conditions are automatically incorporated into this optimization and are not restarted repeatedly.

[0061] Preferably, the upper limit constraint on the triggering frequency is as follows: 1. Single-target cooldown period: After the same target completes one optimization, it enters a cooldown period (5 seconds), during which all triggers are ignored; 2. Global frequency limit: The system can trigger a maximum of 10 times per second. Any triggers exceeding this limit will be queued in a priority order. 3. Continuous Trigger Suppression: If the same target is triggered more than twice within 1 second, the trigger threshold will be automatically and temporarily increased.

[0062] Step Six: Triggered Vehicle Recognition and Trajectory Tracking Optimization When any of the triggering conditions in step five is met, the system initiates the high-precision recognition and trajectory optimization process to identify and optimize the tracking trajectory of the vehicle target.

[0063] First, based on the transient features of the radar-visual fusion at the current trigger time and several preceding times (e.g., the previous 5 time steps), a multi-scale feature pyramid is constructed through a temporal self-attention mechanism and multi-scale upsampling operations. This feature pyramid contains multiple scale layers; for example, three scale layers can be constructed, corresponding to small-sized targets (e.g., distant vehicles), medium-sized targets (e.g., regular vehicles), and large-sized targets (e.g., large trucks at close range), respectively. The feature map sizes of each scale layer are, for example, 1 / 8, 1 / 16, and 1 / 32 of the original size, respectively.

[0064] Subsequently, a multi-scale feature pyramid is input into the trajectory optimization network. This network consists of cascaded Transformer attention layers, which refine the embedding of the vehicle's position, motion state, and appearance features through an iterative cross-attention mechanism. The trajectory optimization network performs a preset number of residual iterative calibrations (e.g., 3 times), using the previous output as a priori in each iteration to successively approximate the optimal trajectory solution. The final output is the optimized vehicle recognition result, including vehicle category, location bounding box, driving speed, and continuous motion trajectory.

[0065] During non-triggered periods, the system maintains a low-power standby state, executing only the basic processes from steps one to three. In this state, the system uses a lightweight filter (e.g., a Kalman filter) to coarsely track and predict the vehicle trajectory based on the transient characteristics of the RTAF output, without activating the computationally intensive Transformer trajectory optimization network. Once optimization is complete, the high-precision trajectory segment output by the Transformer overwrites and replaces the corresponding historical coarse-tracking trajectory segment. The updated trajectory state is then used as a new reference input to the lightweight filter, and the system exits the triggered state, resuming low-power standby mode. Through this trigger-based computation mechanism, on-demand computation scheduling is achieved, effectively reducing average computational power consumption.

[0066] Step 7: Model Training Methods The method in this embodiment also includes the step of training the neural network model used to achieve the above-described cross-modal fusion and temporal recursive updates.

[0067] The training phase first requires constructing a training sample set. The training samples cover various road scenarios (e.g., highway mainlines, tunnels, interchanges, toll stations), various traffic states (e.g., off-peak, peak, congestion, free flow), various vehicle types (e.g., cars, trucks, trailers, buses), and various vehicle motion states (e.g., normal driving, acceleration / deceleration, lane changing, obstruction, high-speed driving). In one specific implementation, at least 10 sets of synchronized data are collected for each vehicle type and motion state combination, and at least 200 valid paired frames are extracted from each set of synchronized data. Data augmentation operations are performed on the training data, including random horizontal flipping of video images, brightness adjustment, contrast adjustment, Gaussian blurring, and adding random position and velocity noise to the radar point cloud to simulate equipment noise and environmental interference in actual perception.

[0068] The training phase constructs a model containing the detection loss term. Trajectory regression loss term and gradient constraint loss term Total loss function:

[0069] in, and This is a balancing coefficient used to adjust the contribution of each individual loss to the total loss. It detects loss items. It includes classification loss and bounding box regression loss, used to supervise the detection accuracy of vehicle targets; gradient constraint loss term. Used to constrain the directional consistency between the video branch gradient and the radar branch gradient; trajectory regression loss term The formula used to monitor the deviation between the predicted trajectory and the true trajectory in the dimensions of position and velocity can be expressed as:

[0070] in, , , To show the predicted position and velocity at time t; , , Let be the actual position and velocity at time t; Specifically, during backpropagation, the video branch gradients are calculated based on the total loss function. and radar branch gradient :

[0071] in, ; These are the parameters for the video branch network (obtained by the video feature encoder). For example, the sigmoid function or the softmax function; This represents the number of effective features in the video branch. The response strength of the i-th feature (reflected by the confidence level or feature map activation value output by the target detection algorithm).

[0072]

[0073] in, ; For radar branch network parameters, The effective point cloud count for the radar branch; This represents the i-th original radar point cloud; This represents the i-th densed radar point cloud; This is the Euclidean distance function, used to calculate the distance deviation between the original point cloud feature vector and the densed point cloud feature vector.

[0074] Gradient constraint loss term By calculating the video branch gradient With radar branch gradient The cosine of the angle between them, i.e. This constrains the gradient directions of the two branches to tend to be orthogonal or in the same direction. During parameter updates, a gradient orthogonality constraint and projection normalization strategy are adopted. When a conflict between two gradient directions is detected (e.g., the dot product is negative), the gradients are adjusted to be in the same or orthogonal direction through projection operations. This prevents gradient conflicts and negative transfer problems in heterogeneous modality joint training, and improves the convergence stability and generalization ability of the model.

[0075] Example 2 Based on Embodiment 1, this embodiment provides another specific implementation of parameter configuration to demonstrate the reasonable generalization of the parameter range in the independent claim.

[0076] In this embodiment, the first frequency of the millimeter-wave radar is 77Hz, and the second frequency of the video camera is 30Hz. At this time, the radar data acquisition period is approximately 13ms, and the video frame interval is approximately 33.3ms. Within the interval between adjacent video frames, the radar-driven timing recursive update can be executed approximately 2 to 3 times (corresponding to...). and (Time). The high-speed target speed threshold in the preset trigger conditions is set to 40 km / h (suitable for urban expressway scenarios), and the first preset duration remains two consecutive radar time steps. The preset error threshold for trajectory prediction error is set to 4 pixels (suitable for high-resolution image scenarios). The loss function balance coefficient during the training phase... and They were set to 0.5 and 0.1 respectively.

[0077] The remaining steps in this embodiment are the same as in Embodiment 1, and will not be repeated here. As can be seen from this embodiment, the method of the present invention has good adaptability and scalability to different sensor configurations and application scenarios.

[0078] Example 3 This embodiment provides a trigger-based vehicle recognition and tracking device based on radar-visual fusion, which is used to implement the method described in Embodiment 1 or Embodiment 2. The device includes the following modules: The data preprocessing module is designed to acquire synchronized radar point cloud data and video image data, and extract radar motion feature sequences and video semantic feature sequences. Specifically, this module performs operations such as hardware-triggered synchronization, timestamp calibration, radar point cloud denoising and dimensionality reduction and feature encoding, and video target detection and feature extraction.

[0079] The transient asynchronous fusion module is configured to initialize the radar-visual fusion transient features using video frames as time anchors, and perform pure radar-driven temporal recursive updates within the interval between adjacent video frames, using the radar cycle as the step size. This module internally includes sub-units such as a feature encoder, a densification operator, and a temporal feature updater.

[0080] The dynamic weight allocation module is designed to adaptively allocate fusion weights based on the feature changes of radar and video modalities in the local spatial and temporal domains. This module calculates and fuses spatial and temporal attention weights to generate dynamic weighting coefficients, which guide the cross-modal fusion process.

[0081] The trigger-based optimization module is configured to make trigger decisions based on changes in the transient features of radar-visual fusion, and to initiate vehicle recognition and trajectory tracking optimization when preset trigger conditions are met. This module includes a trigger condition decision subunit, a multi-scale feature pyramid construction subunit, and a Transformer trajectory optimization network.

[0082] The specific implementation principles and functions of the above modules are consistent with the descriptions of the corresponding steps in Embodiments 1 and 2, and will not be repeated here. The data flow and control flow relationships between the modules are as follows: Figure 1 As shown.

[0083] Example 4 This embodiment also provides an electronic device, see reference. Figure 2 It includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.

[0084] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.

[0085] Memory 404 may include a mass storage device for data or instructions. For example, and not limitingly, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to a data processing device. In a particular embodiment, memory 404 is non-volatile memory. In a particular embodiment, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.

[0086] The memory 404 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 402.

[0087] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any of the trigger-based vehicle recognition and tracking methods based on radar-visual fusion in the above embodiments.

[0088] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the processor 402, and the input / output device 408 is connected to the processor 402.

[0089] The transmission device 406 can be used to receive or send data via a network. Specific examples of the network described above may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 406 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0090] Input / output device 408 is used to input or output information.

[0091] Example 5 This embodiment also provides a readable storage medium storing a computer program, the computer program including program code for controlling a process to execute the process, the process including the trigger-based vehicle recognition and tracking method based on radar-visual fusion according to Embodiment 1.

[0092] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0093] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention can be implemented in hardware, while others can be implemented by firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, these blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0094] Embodiments of the present invention can be implemented by computer software, which may be executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets, and / or macros can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product may include one or more computer-executable components configured to perform the embodiments when the program is run. The one or more computer-executable components may be at least one piece of software code or a portion thereof. Additionally, it should be noted in this respect that, as Figure 1 Any box in the logical flow can represent a program step, or interconnected logic circuits, boxes and functions, or a combination of program steps and logic circuits, boxes and functions. Software can be stored on physical media such as memory chips or blocks of storage implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc. The physical medium is a non-transient medium.

[0095] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0096] The above embodiments are merely illustrative of several implementations of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the appended claims.

Claims

1. A trigger-based vehicle recognition and tracking method based on radar-visual fusion, characterized in that, Includes the following steps: Acquire synchronized radar point cloud data and video image data, and preprocess the radar point cloud data and video image data to obtain radar motion feature sequence and video semantic feature sequence; Using the arrival time of the video frame as the time anchor, cross-modal fusion is performed based on the current video frame features and the radar motion features within the corresponding time window to generate the initial radar-visual fusion transient features. Within the time interval between adjacent video frames, with the radar data acquisition cycle as the step size, a purely radar-driven temporal recursive update is performed based on the radar-visual fusion transient features of the previous moment and the radar motion features acquired at the current moment, to obtain the updated radar-visual fusion transient features on the continuous time axis. In the process of cross-modal fusion and temporal recursive update, fusion weights are adaptively allocated based on the feature changes of radar mode and video mode in the local spatial domain and temporal domain. Based on the change state triggering trajectory optimization of the transient features of the radar-visual fusion, when the preset triggering conditions are met, the vehicle target is identified and the tracking trajectory is optimized.

2. The trigger-based vehicle identification and tracking method as described in claim 1, characterized in that, Using the arrival time of the video frame as the time anchor, cross-modal fusion is performed based on the features of the current video frame and the radar motion features within the corresponding time window to generate initialized radar-visual fusion transient features, including: The radar motion features within the time window are densified, and the densified radar motion features and the current video frame features are encoded into a unified latent state space respectively. Local adaptive weighted fusion is performed on the encoded radar features and video features to obtain the initialized radar-video fusion transient features.

3. The trigger-based vehicle identification and tracking method as described in claim 1, characterized in that, Perform a purely radar-driven timing recursive update, including: The radar-visual fusion transient features of the previous moment are used as the query quantity, and the radar motion features acquired at the current moment are encoded and used as the key value and the numerical value. By performing cross-attention calculation on the query volume, key value volume and numerical value volume through a temporal feature updater, and injecting the motion dynamic features of the current moment while preserving the consistency of historical semantic features, the updated radar-visual fusion transient features are obtained.

4. The trigger-based vehicle identification and tracking method as described in claim 1, characterized in that, Based on the feature changes of radar and video modes in the local spatial and temporal domains, fusion weights are adaptively assigned, including: Based on the spatial geometric distribution relationship between the target node and its local spatial neighboring nodes, determine the spatial domain attention weight; The temporal attention weights are determined based on the rate of change of features between adjacent frames within a local time window. By combining the spatial domain attention weights and the temporal domain attention weights, dynamic weighting coefficients for cross-modal fusion are generated.

5. The trigger-based vehicle identification and tracking method as described in claim 1, characterized in that, Based on the change state triggering trajectory optimization of the transient features of the radar-visual fusion, when a preset triggering condition is met, vehicle target identification and tracking trajectory optimization are performed, including at least one of the following: When the instantaneous velocity of the target detected by the radar point cloud data reaches a preset velocity threshold and the duration reaches a first preset duration, it is determined that the triggering condition is met. When a new target is detected entering the preset monitoring area, and there is no matching identifier within the preset number of historical frames and the confidence level is greater than the preset confidence threshold, the triggering condition is determined to be met. When the intersection-union ratio of the target bounding box is less than the first preset value and the similarity of appearance features is less than the second preset value and continues for a preset number of frames, it is determined that the target is occluded or deoccluded and the triggering condition is met. When the error between the predicted trajectory position and the actual tracking position is greater than or equal to a preset error threshold, and the duration reaches a second preset duration, the trigger condition is determined to be met.

6. The trigger-based vehicle identification and tracking method as described in claim 1, characterized in that, Vehicle target identification and tracking trajectory optimization include: A multi-scale feature pyramid is constructed based on the aforementioned transient features of radar-visual fusion. The multi-scale feature pyramid is input into the trajectory optimization network, and the position, motion state and appearance features of the vehicle target are corrected through an iterative cross-attention mechanism, and the optimized vehicle recognition result and tracking trajectory are output.

7. The trigger-based vehicle identification and tracking method as described in claim 6, characterized in that, The method further includes a step of training a neural network model for implementing the cross-modal fusion and the temporal recursive update, the training step including: Construct a total loss function that includes a detection loss term, a trajectory regression loss term, and a gradient constraint loss term; The video branch gradient and the radar branch gradient are calculated based on the total loss function, and the network parameters are updated by constraining the directional consistency between the video branch gradient and the radar branch gradient.

8. A trigger-based vehicle recognition and tracking device based on radar-visual fusion, characterized in that, include: The data preprocessing module is used to acquire synchronized radar point cloud data and video image data, and extract radar motion feature sequences and video semantic feature sequences. The transient asynchronous fusion module is used to initialize the transient features of radar-visual fusion with video frames as time anchors, and to perform pure radar-driven time-series recursive updates with radar cycles as steps during the interval between adjacent video frames. The dynamic weight allocation module is used to adaptively allocate fusion weights based on the feature changes of radar mode and video mode in the local spatial domain and temporal domain. The trigger-based optimization module is used to make trigger decisions based on the changing state of the transient features of the radar-visual fusion, and to start vehicle recognition and trajectory tracking optimization when the preset trigger conditions are met.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the trigger-based vehicle identification and tracking method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, the computer program including program code for controlling a process to execute the process, the process including the trigger-based vehicle identification and tracking method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Roadside fusion tracking method and system based on Leiye frame fusion

    CN119335527A

  • Traffic state perception and control method based on laser vision fusion

    CN120496331A

  • Vehicle track prediction method and device and vehicle

    CN120747682A

  • Three-dimensional multi-target tracking method fusing radar and vision multiple modes and related equipment

    CN120820938A

  • Target detection tracking method and device based on Leiyu fusion perception and medium

    CN121454510A