A radar visual data association method based on a deep learning algorithm
By using a deep learning algorithm based on the CenterFusion network and leveraging multi-feature fusion and back-projection mechanisms, the problems of anti-interference and false association in data association in the radar-video traffic fusion perception system are solved, achieving higher association accuracy and stability.
Patent Information
- Application Number
- CN202310734115.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-06-20
AI Technical Summary
In existing radar-video traffic fusion perception systems, data association methods are easily affected by the state of individual sensors and have poor anti-interference capabilities. In particular, false associations are serious in dense and complex scenarios, which affects the performance of fusion perception.
A deep learning algorithm based on the CenterFusion network is used to extract multi-class features of radar and visual targets. The positional changes of visual targets are amplified through a back-projection mechanism. Weight coefficients are set based on motion, scale, and appearance similarity to perform multi-source data association, filter out false alarm targets, and improve the association accuracy.
It enhances the anti-interference capability and robustness of radar visual data association, improves the association accuracy in dense scenes, solves the problems of false alarms and misassociations, and improves the stability and adaptability of the fusion perception system.
Smart Images

Figure CN116778290B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of radar visual information fusion technology, specifically to a radar visual data association method based on deep learning algorithms. Background Technology
[0002] As a breakthrough point for improving the perception capabilities of detection systems, radar camera information fusion technology can give full play to the advantages of each sensor, achieve information complementarity, make up for the performance limitations of a single sensor, and obtain more stable and reliable environmentally compatible information, which has broad prospects in military, civilian and other fields.
[0003] In radar-video traffic fusion sensing systems, a crucial issue is determining whether output information from different local nodes points to the same target—the data association problem. Existing technologies address this by searching for the radar frame with the closest temporal sequence to the current video frame as the matching radar frame, and then fusing the current video frame and the matching radar frame. This method does not fully utilize sensor data features, is susceptible to the operating status of individual sensors, and has poor anti-interference capabilities. Another approach involves associating radar echo signals with target attitude recognition from video surveillance to achieve matching and fusion of millimeter-wave radar and video targets. This method is only suitable for scenarios with significant and rapidly changing target motion states; its matching effectiveness drops drastically in situations with high clutter interference and complex scenes. Furthermore, in dense, congested target scenarios, the difficulty of associating data from multiple sensors increases significantly, limiting the performance development of fusion sensing systems.
[0004] Therefore, how to provide a radar visual data association method to achieve better fusion perception performance has become one of the key challenges in the field. Summary of the Invention
[0005] The purpose of this invention is to provide a radar visual data association method based on deep learning algorithms. It extracts multi-class features from radar and visual targets using the CenterFsuion network architecture and designs a back-projection mechanism to project visual detection information onto the radar coordinate system, amplifying changes in the position of visual targets. Multi-source data association is then performed by combining radar and visual continuous frame matching information. The first-stage association filters out a large number of false alarm targets. Then, different feature weight coefficients are set for targets of different sizes, resulting in a more accurate association in the second-stage association, avoiding false associations in dense scenes. This method has stronger scene adaptability and higher robustness in scenarios with false alarms, missed detections, and dense targets.
[0006] To achieve the above objectives, this invention provides a radar visual data association method based on a deep learning algorithm, comprising the following steps:
[0007] S1. Obtain historical fusion trajectory, current radar target trajectory, and visual image frame; the historical fusion trajectory includes the radar visual associated target trajectory, radar target trajectory, and visual target trajectory of the previous moment.
[0008] S2. Input the current visual image frame into the CenterFusion network, output the visual target detection box of visual detection and recognition; and set up a back projection mechanism to back project the visual target detection box to the radar coordinate system to obtain the position of the corresponding visual target in the radar coordinate system.
[0009] S3. Calculate the motion, scale, and appearance similarity between the radar target and the visual target, and preset corresponding weight coefficients for the motion, scale, and appearance similarity to calculate the first association similarity between the radar target and the visual target.
[0010] S4. If the first association similarity is greater than the set first association threshold, the updated position of the corresponding radar target is estimated based on the historical fusion trajectory, and the updated position is matched with the actual position of the corresponding radar target in the historical fusion trajectory. False alarm targets are filtered out according to the matching result.
[0011] S5. Based on the size of the visual target, update the corresponding weight coefficients for the motion, scale, and appearance similarity, and update the first association similarity to the second association similarity; if the second association similarity is greater than the set second association threshold, establish a corresponding radar-visual association pair based on the corresponding radar target and the visual target, and update the historical fusion trajectory based on the corresponding radar target; proceed to the next moment and repeat steps S1 to S5.
[0012] Optionally, in step S1, the echo signal received by the current radar sensor is processed to generate radar target traces; and the current radar target trajectory set R is obtained based on the traces using a joint probabilistic data association algorithm and a Kalman filter algorithm. t , Let m be the trajectory of the current i-th radar target, m be the total number of trajectories in the current radar target trajectory, and t be the number of frames of the echo signal.
[0013] The set of historical integration trajectories is denoted as F. t-1 , For F t-1 The k-th historical fusion trajectory in the set, where p is the total number of trajectories in the historical fusion trajectory set.
[0014] Optionally, step S2 includes:
[0015] S21. Input the current visual image frame into the CenterFusion network to obtain the set of visual targets labeled by the visual target detection boxes. This represents the j-th visual target identified, and n is the total number of visual targets identified.
[0016] S22, Equipment installation height H dev The target space height - H is obtained as prior information. dev +H obj / 2,H obj The target's own height is determined; a position transformation matrix between the visual and radar sensors is obtained through intrinsic parameter calibration and 3D coordinate transformation of the visual sensor; based on the position transformation matrix, the visual target is... The lower edge of the visual target detection bounding box is used as a reference point and projected onto the radar coordinate system to obtain the visual target. Back projection position in radar coordinate system Visual targets The lateral and longitudinal positions of the back projection in the radar coordinate system.
[0017] Optionally, step S3 includes:
[0018] S31. Based on the position of the radar target and the back-projected position of the visual target in the radar coordinate system, obtain the motion similarity Φ between the radar target and the visual target. motion (r t i ,v t j ):
[0019]
[0020] Representing the trajectory r respectively t i The corresponding horizontal and vertical positions of the radar target in the radar coordinate system. Representing visual targets respectively The width and height of the visual target detection bounding box;
[0021] S32, Computer Vision Targets Trajectory in the fusion of history Scale similarity
[0022]
[0023] Representing the trajectory respectively Width and length of the visual target detection box for the associated visual target; Representing visual targets respectively Width and length of the visual target detection bounding box;
[0024] S33. Using the grayscale histogram of the visual target as an appearance feature, calculate the visual target using Bhattacharyya distance. Trajectory in the fusion of history Appearance similarity between
[0025] They represent The grayscale histogram features;
[0026] S34. Set corresponding weight coefficients β1, β2, and β3 for motion, scale, and appearance similarity, and assign preset values to them respectively to obtain the correlation matrix. Among them visual targets With trajectory The first correlation similarity of the corresponding radar target Represented as:
[0027]
[0028] Optionally, step S4 includes:
[0029] S41. If the first association similarity is greater than the set first association threshold, perform a second association matching between the visual target and the historical fusion trajectory; use the Kalman filter algorithm to estimate the trajectory. Updated position of corresponding radar target trajectory The most recent data update time is t′. They represent the trajectories in the radar coordinate system at time t′, respectively. The estimated horizontal and vertical positions;
[0030] S42. If the following relationship is satisfied, the trajectory will be... The corresponding radar and visual targets are placed into the correlation pool; otherwise, the radar and visual targets are treated as false alarm targets and filtered out.
[0031]
[0032] thre_x, thre_y, and thre_t represent the horizontal and vertical positions of the current visual target in the radar coordinate system, respectively, and represent the matching thresholds for the horizontal position, vertical position, and time interval.
[0033] Optionally, step S5 includes:
[0034] S51. Based on the proportion of visual targets in visual image frames, the visual targets in the pool to be associated are divided into large targets, medium targets, and small targets, and the corresponding motion, scale, and appearance similarity weight coefficients β1, β2, and β3 are updated for the large targets, medium targets, and small targets respectively; based on the updated β1, β2, and β3, the second association similarity between the visual targets and radar targets in the pool to be associated is calculated.
[0035] S52. Use the Hungarian algorithm to filter visual targets and radar targets whose second association similarity is greater than the second association threshold to obtain one-to-one radar-visual association pairs.
[0036] S53. Based on the CenterFusion network and radar-visual correlation, the corresponding radar features and visual features are fused, the target information is regressed, and the regression output is stored in the corresponding historical fusion trajectory as the network input information for the next moment; the target information includes any one or more of the target's position, speed, direction of motion, and size.
[0037] Optionally, in step S34, the preset values for β1, β2, and β3 are all 1 / 3.
[0038] Optionally, in step S51, let σ be the pixel value of the visual target, and let Size be the pixel value of the visual image frame; σ1 = Size / 100, σ2 = (3*Size) / 100. When σ < σ1, it is a small target; if σ1 ≤ σ < σ2, it is a medium target; if σ ≥ σ2, it is a large target. Then the updated β1, β2, and β3 are as follows:
[0039]
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] 1) The radar visual data association method based on deep learning algorithms of this invention employs a multi-feature fusion approach. Based on the CenterFsuion network, radar targets are projected onto visual image frames, and visual targets with detection boxes are back-projected onto the radar coordinate system, amplifying changes in the position of the visual targets. In the first-stage association, the motion, scale, and appearance similarity between the visual and radar targets are calculated, and corresponding weight coefficients are assigned to each similarity to obtain the first association similarity. Based on the first association similarity and historical fusion trajectories, matching is performed with the visual targets in a spatiotemporal dimension. False alarm targets are filtered out based on the matching results, improving the anti-interference capability of radar visual data association and solving the problem of decreased association accuracy due to a large number of targets and false detections.
[0042] 2) This invention updates the weight coefficients of motion, scale, and appearance similarity of visual targets based on their different proportions in visual image frames and size differences. This results in a more accurate second-stage association similarity, leading to stable, one-to-one radar-visual association pairs and avoiding erroneous associations in dense scenes. It also solves the problem of chaotic trajectory allocation when radar targets exhibit high similarity or occlusion.
[0043] 3) This invention combines deep learning feature extraction algorithms with historical fusion trajectories. Multi-source data association is performed, making the data association scenarios more adaptable and significantly improving the accuracy of data association. Attached Figure Description
[0044] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the drawings in the following description are one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort:
[0045] Figure 1 , Figure 2 This is a flowchart of a radar visual data association method based on a deep learning algorithm, as described in an embodiment of the present invention.
[0046] Figure 3 This is a schematic diagram illustrating the projection of a radar target onto a visual image frame in an embodiment of the present invention;
[0047] Figure 4 This is a schematic diagram illustrating the back projection of a visual target onto the radar coordinate system in an embodiment of the present invention;
[0048] Figure 5 This is a schematic diagram of the radar-visual correlation pair fused and output in the radar coordinate system in an embodiment of the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0051] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0052] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0053] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0054] Furthermore, in the description of this application, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0055] This invention provides a radar visual data association method based on deep learning algorithms, such as... Figure 1 , Figure 2 As shown, the steps include:
[0056] S1. Acquire the current radar target trajectory, visual image frames, and historical fused trajectories;
[0057] In step S1, the echo signal received by the radar sensor is processed to generate radar target traces; and based on these traces, the current radar target trajectory set R is obtained using a joint probabilistic data association algorithm and a Kalman filter algorithm. t , Let m be the trajectory of the i-th radar target, m be the total number of trajectories in the current radar target trajectory, and t be the frame number.
[0058] The set of historical integration trajectories is denoted as F. t-1 , For F t-1 The k-th historical fusion trajectory in the set, where p is the total number of trajectories in the historical fusion trajectory set. In this embodiment, It includes the trajectory of the corresponding radar target in frame t-1, which has an associated visual target.
[0059] S2. Input the current radar target trajectory and visual image frame into the CenterFusion network, and output the visual target detection box of visual detection and recognition; set up a back projection mechanism to back project the visual target detection box of the visual target to the radar coordinate system, and obtain the position of the corresponding visual target back projection in the radar coordinate system.
[0060] Step S2 includes:
[0061] S21. Input the current visual image frame into the CenterFusion network to obtain the set of visual targets labeled by the visual target detection boxes. Let represent the j-th visual target identified, and n be the total number of visual targets identified. Based on mature multi-sensor spatial synchronization technology, radar detection points are projected onto the pixel space, as illustrated below. Figure 3 As shown.
[0062] S22, Equipment installation height H dev The target space height - H is obtained as prior information. dev +H obj / 2,H obj The target's own height is determined. Based on existing spatial synchronization technology, a position transformation matrix between the visual and radar sensors is obtained through intrinsic parameter calibration and three-dimensional coordinate transformation of the visual sensor; based on the position transformation matrix, the visual target is... The lower edge of the visual target detection bounding box is used as a reference point and projected onto the radar coordinate system, such as... Figure 4 As shown, visual target is obtained Back projection position in radar coordinate system Visual targets The lateral and longitudinal positions of the back projection in the radar coordinate system.
[0063] S3. Calculate the motion, scale, and appearance similarity between the radar target and the visual target, and preset corresponding weight coefficients for the motion, scale, and appearance similarity to calculate the first correlation similarity between the radar target and the visual target.
[0064] Step S3 includes:
[0065] S31. Based on the position of the radar target and the back-projected position of the visual target in the radar coordinate system, obtain the motion similarity Φ between the radar target and the visual target. motion (r t i ,v t j ):
[0066]
[0067] Representing the trajectory r respectively t i The corresponding horizontal and vertical positions of the radar target in the radar coordinate system. Representing visual targets respectively The width and height of the visual target detection bounding box;
[0068] S32, Computer Vision Targets Trajectory in the fusion of history Scale similarity
[0069]
[0070] Representing the trajectory respectively Width and length of the visual target detection box for the associated visual target; Representing visual targets respectively Width and length of the visual target detection bounding box;
[0071] S33. Using the grayscale histogram of the visual target as an appearance feature, calculate the visual target using Bhattacharyya distance. Trajectory in the fusion of history Appearance similarity between
[0072] They represent The grayscale histogram features;
[0073] S34. Set corresponding weight coefficients β1, β2, and β3 for motion, scale, and appearance similarity, and assign preset values to them respectively to obtain the correlation matrix. Among them visual targets With trajectory The first correlation similarity of the corresponding radar target Represented as:
[0074]
[0075] In this embodiment, the preset values of β1, β2, and β3 are all 1 / 3.
[0076] S4. If the first association similarity is greater than the set first association threshold, the updated position of the corresponding radar target is estimated based on the historical fusion trajectory, and the updated position is matched with the actual position of the corresponding radar target in the historical fusion trajectory. False alarm targets are filtered out according to the matching result.
[0077] Step S4 includes:
[0078] S41. If the first association similarity is greater than the set first association threshold, perform a second association matching between the visual target and the historical fused trajectory. Use the Kalman filter algorithm to estimate the trajectory. Updated position of corresponding radar target trajectory The most recent data update time is t′. They represent the trajectories in the radar coordinate system at time t′, respectively. The estimated horizontal and vertical positions;
[0079] S42. If the following relationship is satisfied, the trajectory will be... The corresponding radar and visual targets are placed into the correlation pool; otherwise, the radar and visual targets are treated as false alarm targets and filtered out.
[0080]
[0081] thre_x, thre_y, and thre_t represent the horizontal and vertical positions of the current visual target in the radar coordinate system, respectively, and represent the matching thresholds for the horizontal position, vertical position, and time interval.
[0082] Steps S3 and S4 above complete the first stage of association between radar and visual targets, eliminating interference introduced by false alarm target detection. In step S3, the CenterFusion network is used to extract three types of features: motion, scale, and appearance similarity, and three weighting coefficients are added for radar-visual data association. In step S4, continuous frame trajectory information is obtained to estimate the target position in the current frame, which is then matched with the current detection value in a spatiotemporal dimension. False alarm targets are filtered out based on the association results.
[0083] S5. Based on the size of the visual target, update the corresponding weight coefficients for the motion, scale, and appearance similarity, and update the first association similarity to the second association similarity; if the second association similarity is greater than the set second association threshold, establish a corresponding radar-visual association pair based on the corresponding radar target and visual target, and update the historical fusion trajectory based on the corresponding radar target; proceed to the next moment and repeat steps S1 to S5.
[0084] Step S5 includes:
[0085] S51. Based on the proportion of visual targets in visual image frames, the visual targets in the pool to be associated are divided into large targets, medium targets, and small targets, and the corresponding motion, scale, and appearance similarity weight coefficients β1, β2, and β3 are updated for the large targets, medium targets, and small targets respectively; based on the updated β1, β2, and β3, the second association similarity between the visual targets and radar targets in the pool to be associated is calculated.
[0086] In this embodiment, let σ be the pixel value of the visual target, and let Size be the pixel value of the visual image frame; σ1 = Size / 100, σ2 = (3*Size) / 100. When σ < σ1, it is a small target; if σ1 ≤ σ < σ2, it is a medium target; if σ ≥ σ2, it is a large target. Then the updated β1, β2, and β3 are as follows:
[0087]
[0088] S52. Use the Hungarian algorithm to filter visual targets and radar targets whose second association similarity is greater than the second association threshold to obtain one-to-one radar-visual association pairs.
[0089] S53. Based on the CenterFusion network, the radar and visual features of the radar-visual association pair are fused, and the corresponding target's position, velocity, direction of motion, size and other information are regressed and output. The regression output is then used to generate the historical fused trajectory for the next moment.
[0090] Step S5 is the second stage of association. To avoid the impact of feature mutations on the overall similarity, the targets are divided according to size based on the results of the first stage of association. Different weight coefficients are set for different targets to avoid misassociations in dense scenes and output more accurate association results.
[0091] Ten minutes of traffic data for the collected road segment were processed through data association to obtain relevant evaluation indicators, as shown in Table 1. The fusion output results based on the association are as follows: Figure 5 As shown, compared with traditional data association methods—such as the Hungarian algorithm and CenterFusion network detection point association—the method of this invention incorporates the role of continuous frame trajectories and a two-stage association using multiple features. In occluded and dense scenes, the association results are more stable and accurate, greatly avoiding the phenomenon of chaotic fusion trajectories, and making the fusion perception system more stable and reliable.
[0092] Table 1 Analysis of the Relationship between Leishi and [other companies].
[0093]
[0094]
[0095] The above embodiments achieve accurate association of radar and video target data, enhance the stability of the radar-video fusion system, improve the adaptability of data association, and provide good processing output for situations such as occlusion and jumps in dense scenes. The method has multiple advantages.
[0096] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0097] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A radar-visual data association method based on deep learning algorithm, characterized in that, The method comprises the steps of: S1, obtaining a historical fusion track, a current radar target track and a visual image frame; S2, inputting the current visual image frame into a CenterFusion network to output a visual target detection frame detected and recognized by the CenterFusion network; setting a back projection mechanism to back project the visual target detection frame to a radar coordinate system to obtain a position of the corresponding visual target in the radar coordinate system; S3, calculating a motion similarity, a scale similarity and an appearance similarity between the radar target and the visual target, and presetting corresponding weight coefficients of the motion similarity, the scale similarity and the appearance similarity to obtain a first correlation similarity between the radar target and the visual target; The step S3 comprises: S31, based on the radar coordinate system, the position of the radar target, the back projection position of the visual target, the motion similarity Φ of the radar target and the visual target is obtained motion (r t i ,v t j ) respectively represent the trajectory r t i the corresponding radar target in the radar coordinate system, respectively represent the visual target width, height of the visual target detection box S32, calculate visual target Trajectory in the history fusion trajectory Scale similarity of the dimension respectively represent the width, length of the visual target detection frame of the visual target associated with the trajectory respectively represent the width, length of the visual target detection frame of the visual target associated with the trajectory respectively represent the width, length of the visual target detection frame of the visual target associated with the trajectory respectively represent the width, length of the visual target detection frame of the visual target associated with the trajectory S33. Using the grayscale histogram of the visual target as an appearance feature, calculate the visual target using Bhattacharyya distance. Trajectory in the fusion of history Appearance similarity between respectively represent a gray histogram feature of S34, set corresponding weight coefficients β1, β2, β3 for motion, scale, and appearance similarity, and give them preset values respectively to obtain a correlation matrix where the visual target and the trajectory r t i the first correlation similarity of the corresponding radar target is represented as: S4, if the first correlation similarity is greater than a first correlation threshold, estimating an updated position of the corresponding radar target based on the historical fusion track, matching the updated position with an actual position of the corresponding radar target in the historical fusion track, and filtering out a false alarm target according to a matching result; S5, updating the weight coefficients of the motion similarity, the scale similarity and the appearance similarity based on a size of the visual target, updating the first correlation similarity to a second correlation similarity, and if the second correlation similarity is greater than a second correlation threshold, establishing a corresponding radar-visual correlation pair based on the corresponding radar target and the visual target, and updating the historical fusion track based on the corresponding radar target; entering a next time, and repeating the steps S1 to S5.
2. The radar visual data association method based on deep learning algorithm of claim 1, wherein, In the step S1, a point track of a radar target is generated after a signal processing of a current echo signal received by a radar sensor; and a current radar target track set R is obtained based on the point track through a joint probability data association algorithm and a Kalman filtering algorithm t , r t i is a track of a current i-th radar target, and m is a total number of tracks in the current radar target track. t represents a frame number of the echo signal; The set of historical fusion trajectories is denoted as F t-1 , For the kth historical fusion trajectory in F t-1 , p is the total number of trajectories in the set of historical fusion trajectories.
3. The radar visual data association method based on deep learning algorithm of claim 1, wherein, The step S2 comprises: S21, input the current visual image frame into the CenterFusion network to obtain a visual target set labeled by a visual target detection box denotes the jth visual target identified, and n is the total number of visual targets identified. S22, set up height H of the device dev As the prior information acquisition target space height-H dev +H obj / 2, H obj The target itself height; through the intrinsic calibration of the vision sensor and the three-dimensional coordinate conversion, the position conversion matrix between the vision sensor and the radar sensor is obtained; based on the position conversion matrix, the vision target The lower edge of the vision target detection frame is projected to the radar coordinate system as a reference point, and the vision target The back projection position in the radar coordinate system The lateral position and the longitudinal position of the vision target back projected in the radar coordinate system.
4. The radar vision data association method based on deep learning algorithm of claim 1, wherein, The step S4 comprises: S41, if the first correlation similarity is greater than a set first correlation threshold, performing secondary correlation matching between the radar, visual detection target and the historical fusion trajectory; using a Kalman filtering algorithm to estimate the trajectory updated position of the corresponding radar target trajectory The latest data update time of the trajectory is t', respectively represent the estimated lateral and longitudinal positions of the trajectory in the radar coordinate system at t' S42, if the following relationship is satisfied, the trajectory The corresponding radar target and visual target are put into the pool to be associated, otherwise the radar target and visual target are regarded as false alarm targets and filtered out. thre_x, thre_y, thre_t represent the matching threshold of lateral position, longitudinal position and interval time respectively.
5. The radar vision data association method based on deep learning algorithm of claim 4, wherein, The step S5 comprises: S51, dividing the visual targets in the to-be-associated pool into large targets, medium targets and small targets based on proportions of the visual targets in the visual image frame, and updating the weight coefficients β1, β2 and β3 of the motion similarity, the scale similarity and the appearance similarity of the large targets, the medium targets and the small targets respectively; calculating the second correlation similarity of the visual targets and the radar targets in the to-be-associated pool based on the updated β1, β2 and β3; S52, screening the visual targets and the radar targets with the second correlation similarity greater than the second correlation threshold by using a Hungarian algorithm to obtain one-to-one radar-visual correlation pairs; S53, fusing corresponding radar features and visual features based on the CenterFusion network and the radar-visual correlation pairs, regressing output target information, and storing the regressed output result in the corresponding historical fusion track as network input information of a next time; The target information comprises any one or more of a position, a speed, a motion direction and a size of the target.
6. The radar vision data association method based on deep learning algorithm of claim 1, wherein, In the step S34, the preset values of β1, β2 and β3 are all 1 / 3.
7. The radar vision data association method based on deep learning algorithm of claim 1, wherein, In the step S51, let σ be a pixel value of the visual target, let Size be a pixel value of the visual image frame, σ1=Size / 100 and σ2=(3*Size) / 100; if σ<σ1, the visual target is a small target; if σ1≤σ<σ2, the visual target is a medium target; if σ≥σ2, the visual target is a large target; and the updated β1, β2 and β3 are as follows:
Citation Information
Patent Citations
Target tracking method and device based on correlation filtering
CN113888586A
Radar and video target adaptive association method
CN115932830A