Target tracking method and device based on radar and visual fusion, storage medium and equipment

By using the radar-visual fusion method, video and radar data are acquired and converted to achieve time synchronization and fusion, which solves the problem of low detection accuracy of monocular cameras and millimeter-wave radar in adverse weather conditions and improves the detection accuracy and adaptability of target objects.

CN119151983BActive Publication Date: 2026-03-27GRG BANKING EQUIPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, monocular cameras and millimeter-wave radars suffer from problems such as significant impact from lighting and weather conditions, short sensing distance, and low detection accuracy in cross-scene target tracking algorithms under adverse weather conditions.

Method used

By using the radar-visual fusion method, data from video devices and radar devices are acquired, converted into geodetic coordinates, and then time-synchronized and fused to obtain cross-device radar-visual fusion results, enabling the tracking trajectory of the target object.

Benefits of technology

It improves the detection accuracy of target objects, adapts to complex scenarios such as severe weather, and is suitable for target object tracking in large-scale scenarios such as intersections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151983B_ABST
    Figure CN119151983B_ABST
Patent Text Reader

Abstract

The application discloses a target tracking method and device based on radar and video fusion, a storage medium and equipment. The method comprises the following steps: acquiring video data of a plurality of video devices and radar data of a plurality of radar devices at a traffic intersection, extracting target video frames and target radar frames containing target objects, obtaining target video pixel coordinates and target latitude and longitude information, performing plane coordinate conversion to obtain geodetic plane coordinates, obtaining time synchronization frame data sets according to time stamps of the target video frames and the target radar frames and the geodetic plane video coordinates and the geodetic plane radar coordinates, fusing target object coordinates of the target video frames and target object coordinates of the target radar frames at the same time to obtain cross-device radar and video fusion results, and finally matching cross-device fusion coordinates of each target object in the cross-device radar and video fusion results at different times to corresponding track trackers to obtain tracking tracks of each target object, so that accurate and efficient multi-sensor target fusion is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of image processing, and particularly relates to a target tracking method and device based on radar and vision fusion, a storage medium and equipment. BACKGROUND

[0002] In the intelligent traffic application scenario, monocular cameras and millimeter wave radars are commonly used sensors for target detection. Monocular cameras can obtain target texture information, but are susceptible to light and weather, and have a short sensing distance. Millimeter wave radars can measure the speed and spatial position of moving targets, and have good adaptability to bad weather such as rain and fog, but cannot sense stationary targets. Most current cross-scene target tracking algorithms only use video data, and have the problems of low target detection accuracy and inability to adapt to complex scenes such as bad weather. SUMMARY

[0003] The present application aims to at least solve one of the technical problems in the related art. To this end, the present application provides a target tracking method and device based on radar and vision fusion, a storage medium and equipment, which can improve the detection accuracy of target objects.

[0004] In a first aspect, the present application provides a target tracking method based on radar and vision fusion, which comprises:

[0005] Obtaining video data of a plurality of video devices and radar data of a plurality of radar devices at a traffic intersection, each video device corresponding to a plurality of video frames, and each radar device corresponding to a plurality of radar frames;

[0006] Extracting target video frames containing target objects from the plurality of video frames, and extracting target video pixel coordinates corresponding to each target object from each target video frame; extracting target radar frames containing the target objects from the plurality of radar frames, and extracting target latitude and longitude information corresponding to each target object from each target radar frame;

[0007] Converting the target video pixel coordinates and the target latitude and longitude information into geodetic plane video coordinates and geodetic plane radar coordinates, respectively, and obtaining a time-synchronized frame data set according to the time stamps of the target video frames and the target radar frames, and the geodetic plane video coordinates and the geodetic plane radar coordinates, the time-synchronized frame data set including synchronized target video frames of different video devices having a time-synchronized corresponding relationship, and synchronized target radar frames of different radar devices having a time-synchronized corresponding relationship, the synchronized target video frames and the synchronized target radar frames also having a time-synchronized corresponding relationship;

[0008] fusing target objects in synchronous target video frames of different video devices at the same time to obtain a single-source target video frame result after fusion, fusing target objects in synchronous target radar frames of different radar devices at the same time to obtain a single-source target radar frame result after fusion, and fusing the single-source target video frame result and the single-source target radar frame result at the same time to obtain a cross-device radar-video fusion result;

[0009] obtaining a tracking trajectory of each target object according to cross-device fusion coordinates of each target object in the cross-device radar-video fusion result at different times.

[0010] According to the target tracking method based on radar-video fusion, target video frames containing target objects of different video devices and target radar frames containing target objects of different radar devices are obtained, target video pixel coordinates corresponding to target objects in the target video frames and target longitude and latitude information of the target objects in the target radar frames are respectively converted into geodetic plane video coordinates and geodetic plane radar coordinates in a same plane coordinate system, target video frames of different video devices at the same time are time-synchronized, target radar frames of different radar devices at the same time are time-synchronized, and the target video frames and the target radar frames at the same time are time-synchronized to obtain a time-synchronized frame data set having a time-synchronized corresponding relationship, target objects in synchronous target video frames of different video devices at the same time are fused to obtain a single-source target video frame result after fusion, target objects in synchronous target radar frames of different radar devices at the same time are fused to obtain a single-source target radar frame result after fusion, the single-source target video frame result and the single-source target radar frame result at the same time are fused to obtain a cross-device radar-video fusion result, and finally, a tracking trajectory of each target object is obtained according to cross-device fusion coordinates of each target object in the cross-device radar-video fusion result at different times. The method can realize accurate and efficient fusion of target objects of different sensors, and can not only improve the detection accuracy of target objects, but also adapt to complex scenes such as bad weather, and can simultaneously adapt to tracking of multiple target objects and can be applied to large-scale scenes such as crossroads, and realizes adaptive tracking of target objects.

[0011] In a second aspect, the present application provides a target tracking device based on radar-video fusion, which comprises:

[0012] a radar-video data acquisition module, configured to acquire video data of multiple video devices and radar data of multiple radar devices at a traffic intersection, each video device corresponding to multiple frames of video data, and each radar device corresponding to multiple frames of radar data;

[0013] a target coordinate extraction module, configured to extract target video frames containing target objects from the plurality of video frames, and extract target video pixel coordinates corresponding to each of the target objects from each of the target video frames, extract target radar frames containing the target objects from the plurality of radar frames, and extract target longitude and latitude information corresponding to each of the target objects from each of the target radar frames;

[0014] a conversion synchronization module, configured to convert the target video pixel coordinates and the target longitude and latitude information into geodetic plane video coordinates and geodetic plane radar coordinates respectively, and obtain a time-synchronized frame data set according to time stamps of the target video frames and the target radar frames, and the geodetic plane video coordinates and the geodetic plane radar coordinates, wherein the time-synchronized frame data set includes synchronized target video frames of different video devices having a time-synchronized corresponding relationship, synchronized target radar frames of different radar devices having a time-synchronized corresponding relationship, and the synchronized target video frames and the synchronized target radar frames also have a time-synchronized corresponding relationship;

[0015] a radar and video fusion module, configured to fuse target objects in synchronized target video frames of different video devices at the same time to obtain a single-source target video frame result after fusion, fuse target objects in synchronized target radar frames of different radar devices at the same time to obtain a single-source target radar frame result after fusion, and fuse the single-source target video frame result and the single-source target radar frame result at the same time to obtain a cross-device radar and video fusion result;

[0016] a trajectory tracking module, configured to obtain a tracking trajectory of each of the target objects according to cross-device fusion coordinates of each of the target objects in the cross-device radar and video fusion result at different times.

[0017] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the target tracking method based on radar and video fusion as described in the first aspect when executing the computer program.

[0018] In a fourth aspect, the present application provides a non-transitory computer readable storage medium having a computer program stored thereon, and the computer program is executable by a processor to implement the target tracking method based on radar and video fusion as described in the first aspect.

[0019] In a fifth aspect, the present application provides a chip, including a processor and a communication interface, the communication interface and the processor are coupled, and the processor is configured to run a program or an instruction to implement the target tracking method based on radar and video fusion as described in the first aspect.

[0020] In a sixth aspect, the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the target tracking method based on radar and vision fusion as described in the first aspect above.

[0021] The one or more technical solutions described above in the embodiments of the present application have at least the following technical effects:

[0022] According to the target tracking method, device, storage medium and equipment based on radar and vision fusion of the present application, by acquiring target video frames containing target objects of different video devices and target radar frames containing target objects of different radar devices, and converting target video pixel coordinates of the target objects in the target video frames and target longitude and latitude information of the target objects in the target radar frames into geodetic plane video coordinates and geodetic plane radar coordinates in a same plane coordinate system respectively, then time-synchronizing the target video frames of different video devices at a same time, time-synchronizing the target radar frames of different radar devices at a same time, and time-synchronizing the target video frames and the target radar frames at a same time, time-synchronized frame data sets with time synchronization corresponding relationship are obtained, further, the target objects in the synchronized target video frames of different video devices at a same time are fused to obtain single-source target video frame results after fusion, the target objects in the synchronized radar frames of different radar devices at a same time are fused to obtain single-source target radar frame results after fusion, the single-source target video frame results and the single-source target radar frame results at a same time are fused to obtain cross-device radar and vision fusion results, and finally, the tracking trajectories of the target objects are obtained according to cross-device fusion coordinates of the target objects in the cross-device radar and vision fusion results at different times, the method and device can realize accurate and efficient target object fusion of different sensors, and due to the complementary characteristics of the fusion of video data and radar data, the detection accuracy of the target objects can be improved, the method can also adapt to complex scenes such as bad weather, the method can be applied to tracking of multiple target objects, can be applied to large-scale scenes such as crossroads, has good scene versatility, and realizes adaptive tracking of the target objects.

[0023] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0024] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings.

[0025] Figure 1 is a flowchart of the target tracking method based on radar and vision fusion provided by the embodiments of the present application;

[0026] Figure 2 is a structural schematic diagram of a target tracking device based on radar and visual fusion provided by an embodiment of the present application.

[0027] Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.

[0029] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0030] The target tracking method based on radar and visual fusion, the target tracking device based on radar and visual fusion, the electronic device, and the readable storage medium provided by the embodiments of the present application will be described in detail below with reference to the drawings, through specific embodiments and their application scenarios.

[0031] The target tracking method based on radar and visual fusion can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.

[0032] The target tracking method based on radar and visual fusion provided by the embodiments of the present application, the execution subject of the target tracking method based on radar and visual fusion can be an electronic device or a functional module or functional entity capable of realizing the target tracking method based on radar and visual fusion in the electronic device. The electronic device mentioned in the embodiments of the present application includes but is not limited to a mobile phone, a tablet computer, a computer, a camera, a wearable device, and the like. The target tracking method based on radar and visual fusion provided by the embodiments of the present application will be described below with the electronic device as an execution subject.

[0033] As shown in the figure, the target tracking method based on radar and visual fusion includes the following steps: Figure 1

[0034] ​S100, acquire video data of a plurality of video devices and radar data of a plurality of radar devices of a traffic intersection, each of the video devices corresponding to a plurality of video data frames, and each of the radar devices corresponding to a plurality of radar data frames.

[0035] In order to adapt to the application of a large range of scenes such as traffic intersections, a plurality of video devices and radar devices are arranged at the intersection, each video device captures video data of a certain time period to obtain a plurality of video data frames, and each radar device captures radar data of a certain time period to obtain a plurality of radar data frames.

[0036] S200, extracting target video frames containing target objects from the plurality of video data frames, and extracting target video pixel coordinates corresponding to each of the target objects from each of the target video frames, extracting target radar frames containing the target objects from the plurality of radar data frames, and extracting target latitude and longitude information corresponding to each of the target objects from each of the target radar frames.

[0037] The angle range of the image captured by the plurality of video devices and the plurality of radar devices may be different, and there may be cases where the target objects are not included in part of the video data and radar data. Therefore, only the target video frames containing the target objects and the target radar frames containing the target objects are selected to reduce the amount of calculation in subsequent fusion.

[0038] In some embodiments, the target video pixel coordinates corresponding to each of the target objects are extracted from each of the target video frames, including: extracting each of the target objects from each of the target video frames. The bottom midpoint pixel coordinates of each of the target objects are taken as the target video pixel coordinates.

[0039] In some embodiments, other position pixel coordinates of the target objects can also be selected as the target video pixel coordinates, such as the barycentric pixel coordinates, the top midpoint pixel coordinates, etc. However, the bottom midpoint pixel coordinates are closest to the ground plane and do not produce errors in height, so the coordinate conversion is more accurate. Therefore, in order to make the coordinate conversion more accurate, only the bottom midpoint pixel coordinates of the target objects are selected as the target video pixel coordinates. The target objects include motor vehicles, pedestrians, non-motor vehicles, etc.

[0040] In some embodiments, the set of target video pixel coordinates and ground plane video coordinates can be represented as: wherein, represents the bottom midpoint pixel coordinates of the target objects in the target video frames under each video device, i.e., the target video pixel coordinates, represents the ground plane coordinates of the bottom midpoint of the target objects in the target video frames under each video device, i.e., the ground plane video coordinates, represents the number of video devices, a number of target video frames of each video device, a number of target objects in each target video frame of each video device, a category of target objects in each target video frame.

[0041] In some embodiments, the set of geodetic plane radar coordinates can be represented as: wherein, a geodetic plane coordinate of a bottom midpoint of a target object in a target radar frame of each radar device, i.e., a geodetic plane radar coordinate, a number of radar devices, a number of target radar frames of each radar device, a number of target objects in each target radar frame of each radar device, a category of target objects in each target radar frame.

[0042] It should be noted that the longitude and latitude of the bottom midpoint position do not exist in the target radar frame, so the longitude and latitude information can be directly used.

[0043] S300, converting the target video pixel coordinates and the target longitude and latitude information into geodetic plane video coordinates and geodetic plane radar coordinates respectively, and obtaining a time-synchronized frame data set according to the time stamps of the target video frames and the target radar frames and the geodetic plane video coordinates and the geodetic plane radar coordinates, wherein the time-synchronized frame data set includes synchronized target video frames of different video devices having a time-synchronized corresponding relationship and synchronized target radar frames of different radar devices having a time-synchronized corresponding relationship, and the synchronized target video frames and the synchronized target radar frames also have a time-synchronized corresponding relationship.

[0044] In order to fuse the video pixel coordinates and the radar longitude and latitude information in the subsequent process, it is necessary to convert the video pixel coordinates and the radar longitude and latitude information into the same coordinate system, and the geodetic plane coordinate system can be used for both. Therefore, the target video pixel coordinates are converted into geodetic plane video coordinates, and the target longitude and latitude information is converted into geodetic plane radar coordinates. In the process of converting the target video pixel coordinates into geodetic plane video coordinates, a mapping relationship of the video data to the geodetic plane coordinate system can be obtained through the video data and the high-precision map data, and then the target video pixel coordinates are converted into geodetic plane video coordinates according to the mapping relationship. Specifically, a road surface arrow mark can be selected in the video frame and the high-precision map, a homography matrix of the plane where the road in the video frame is located and the plane where the high-precision map is located can be calculated based on the road surface arrow mark, and a mapping relationship of the video road plane to the geodetic plane coordinate system can be obtained. The matrix calculation method for obtaining the mapping relationship can be implemented by using the existing technology, and will not be described here.

[0045] In some embodiments, the step of obtaining the time-synchronized frame dataset according to the timestamps of the target video frames and the target radar frames and the geodetic plane video coordinates and the geodetic plane radar coordinates in step S300 can be implemented in the following manner:

[0046] determining the initial radar matching frame within the preset range corresponding to the target video frame at the same time according to the timestamps of the target video frame and the target radar frame.

[0047] calculating the Euclidean distance between each target object in the target video frame at the same time and all target objects in the initial radar matching frame according to the geodetic plane video coordinates and the geodetic plane radar coordinates, and taking the initial radar matching frame in which each target object is located when the Euclidean distance is the smallest as the radar matching frame at the corresponding time of the target video frame.

[0048] taking all target video frames at all times and their corresponding radar matching frames as the time-synchronized frame dataset.

[0049] It should be noted that this embodiment is to achieve time synchronization of target video frames and target radar frames at different times. It can be understood that the target video frames can be time-synchronized only by timestamps, and the target radar frames can also be time-synchronized only by their respective times, but the target video frames and the target radar frames cannot be determined to be time-synchronized only by timestamps due to the difference in frame rate.

[0050] Specifically, the target video frame at the same time and the target radar frame at the same time are determined according to the timestamps of the target video frame and the target radar frame, respectively, and then the target radar frame set of different radar devices having the same timestamp as the target video frame at the same time is determined according to the timestamps, denoted as To overcome the problem of different frame rates, the target radar frame within the preset range of the target radar frame set is selected again, and the initial radar matching frame is obtained by taking the union of the target radar frame set and the target radar frame, denoted as (m k -r,m k +r), the Euclidean distance is calculated through the geodetic plane video coordinates of each target object of the target video frame at the same time and the geodetic plane radar coordinates of each target object in the initial radar matching frame, and the initial radar matching frame in which each target object is located when the Euclidean distance is the smallest is taken as the radar matching frame at the corresponding time of the target video frame, that is, through the formula obtained, thereby realizing time synchronization of the target video frame and the target radar frame, obtaining all target video frames at all times and their corresponding radar matching frames, and taking the set of all target video frames at all times as:

[0051]

[0052] The target radar frame set at all times is denoted as:

[0053]

[0054] N m 、 respectively represent the number of time-synchronized frames (including the number of target video frames and the number of target radar frames), the number of target objects corresponding to each target video frame, the number of target objects corresponding to each target radar frame, and the superscript m of the coordinates and target object categories represents the result after time synchronization.

[0055] S400, the target objects in the synchronized target video frames of the same time of different video devices are fused to obtain a single-source target video frame result after fusion, the target objects in the synchronized target radar frames of the same time of different radar devices are fused to obtain a single-source target radar frame result after fusion, and the single-source target video frame result and the single-source target radar frame result of the same time are fused to obtain a cross-device radar-visual fusion result.

[0056] It can be understood that, on the basis of step S300, the time-synchronized synchronized target video frames are obtained through timestamps, and then the target objects in the synchronized target video frames of the same time can be fused to obtain a single-source target video frame result after fusion. Specifically, in some embodiments, the single-source target video frame result can be obtained in the following manner:

[0057] A reference video device is selected from the video devices at a certain time, and the remaining video devices are taken as comparative video devices.

[0058] The target objects in the comparative video devices are aligned with the corresponding target objects in the reference video device by using the terrestrial plane video coordinates of the reference video device and the terrestrial plane video coordinates of the comparative video devices, and the aligned target objects in each video device are taken as aligned video coordinates.

[0059] The mean value of all the aligned video coordinates of each target object is calculated to obtain the fusion video pixel coordinates of each target object.

[0060] The single-source target video frame result at the certain time is obtained according to the fusion video pixel coordinates of each target object.

[0061] It should be noted that since the synchronization target video frames at the same time after time synchronization can be multiple, and one video device can only collect one frame of video data at the same time, each synchronization target video frame corresponds to one video device, and therefore, fusing the target objects in multiple synchronization target video frames is equivalent to fusing the target objects in the target video frames collected by multiple video devices at the same time. In some embodiments, a reference video device is selected from the multiple video devices at a certain time, and in some embodiments, the video device with the largest number of target objects is selected as the reference video device, and if there are multiple video devices with the largest number of target objects, the video device with the smaller device number is selected as the reference video device, and the same applies to the selection of the reference video device. If the number of target objects in all video devices is the same, the video device with device number 1 is selected as the reference video device, and the other video devices are selected as comparison video devices.

[0062] After the reference video device and the comparison video device are selected, the target objects in the comparison video device are aligned with the corresponding target objects in the reference video device by using the geodetic plane video coordinates of the reference video device and the geodetic plane video coordinates of the comparison video device, and the geodetic plane video coordinates of the aligned target objects are used as the aligned video coordinates. In some embodiments, when aligning the target objects in the comparison video device with the corresponding target objects in the reference video device, the formula is used to calculate the Euclidean distance between the target objects in the comparison video device and the corresponding target objects in the reference video device, where represents the geodetic plane video coordinates of the target objects in the reference video device, represents the geodetic plane video coordinates of the target objects in the comparison video device, and a target object matching cost table between the comparison video device and the reference video device is constructed, and the size of the cost table is According to the formula of the above Euclidean distance, the column index j corresponding to the minimum cost value of each row i is calculated, if i≠j, the id of the target object in the comparison video device is exchanged with the id of the corresponding target object in the reference video device, otherwise no processing is performed, thereby aligning the target objects in the comparison video device with the corresponding target objects in the reference video device. For example, aligning the 6 target objects in the second video device (comparison video device) with the 5 target objects in the first video (reference video device), if the first target object in the first video device is closest to the third target object in the second video device, the id of the third target object in the second video device needs to be modified to 1.

[0063] The geodetic plane video coordinates of each target object after alignment are taken as alignment video coordinates, and the mean of all alignment video coordinates of each target object is calculated to obtain the fusion video pixel coordinates of each target object, and then the fusion video pixel coordinates of each target object are fused into the same video frame to obtain a single-source target video frame result at a certain moment, thereby realizing single-source fusion of a single-source video frame. A single-source video frame set at different moments is denoted as: wherein c represents the result after single-source fusion, V is a subscript, s in the coordinate superscript and target object category superscript also represents the result after single-source fusion, and different symbols are used here only to distinguish.

[0064] In some embodiments, the step of fusing target objects in synchronous radar frames of different radar devices at the same moment in step S400 to obtain a fused single-source target radar frame result specifically includes:

[0065] A reference radar device is selected from the radar devices at a certain moment, and the remaining radar devices are taken as comparative radar devices.

[0066] The target objects in the comparative radar devices are aligned with the corresponding target objects in the reference radar device by using the geodetic plane radar coordinates of the reference radar device and the geodetic plane radar coordinates of the comparative radar devices, and the geodetic plane radar coordinates of each target object in each radar device after alignment are taken as alignment radar coordinates.

[0067] The mean of all alignment radar coordinates of each target object is calculated to obtain the fusion radar coordinates of each target object.

[0068] The single-source target radar frame result at the certain moment is obtained according to the fusion radar coordinates of each target object.

[0069] Similarly, in some embodiments, target objects in synchronous radar frames of different radar devices at the same moment are fused to obtain a fused single-source target radar frame result, which can also be obtained in a manner similar to that of a single-source target video frame result.

[0070] Each radar device can only collect one frame of radar data at the same time, each synchronous target radar frame corresponds to one radar device, so the fusion of target objects in multiple synchronous target radar frames is equivalent to the fusion of target objects in the target radar frames collected by multiple radar devices at the same time. In some embodiments, the radar device with the largest number of target objects is selected as the reference radar device, if there are multiple radar devices with the largest number of target objects, the radar device with the smaller device number is selected as the reference radar device, and so on, if all radar devices have the same number of target objects, the radar device with device number 1 is selected as the reference radar device, and the other radar devices are selected as the comparison radar devices.

[0071] After selecting the reference radar device and the comparison radar device, the target objects in the comparison radar device are aligned with the corresponding target objects in the reference radar device using the ground plane radar coordinates of the reference radar device and the ground plane radar coordinates of the comparison radar device, and the aligned radar coordinates of each target object in each radar device are used as the aligned radar coordinates. In some embodiments, when aligning each target object in the comparison radar device with the corresponding target object in the reference radar device, the formula is used to calculate the Euclidean distance between each target object in the comparison radar device and the corresponding target object in the reference radar device, where represents the ground plane radar coordinates of the target object in the reference radar device, represents the ground plane radar coordinates of the target object in the comparison radar device, and a target object matching cost table between the comparison radar device and the reference radar device is constructed, the size of the cost table is According to the formula of the above Euclidean distance, the column index j corresponding to the minimum cost value of each row i is calculated, if i≠j, the id of the target object in the comparison radar device is exchanged with the id of the corresponding target object in the reference radar device, otherwise no processing is performed, thereby realizing the alignment of the target objects in the comparison radar device with the corresponding target objects in the reference radar device. For example, aligning the 6 target objects of the second radar device (comparison radar device) with the 5 target objects of the first radar (reference radar device), if the first target object of the first radar device is closest to the third target object of the second radar device, the id of the third target object in the second radar device needs to be modified to 1.

[0072] The aligned ground-plane radar coordinates of each target object are used as the aligned radar coordinates. The mean of all aligned radar coordinates for each target object is calculated to obtain the fused radar coordinates of each target object. Then, the fused radar coordinates of each target object are fused into the same radar frame to obtain the single-source target radar frame result at a certain time, thus achieving single-source fusion of single-source radar frames. The set of single-source radar frames at different times is denoted as: Where 'c' represents the result after single-source fusion, and 's' in the R subscript, coordinate superscript, and target object category superscript also represents the result after single-source fusion, the different symbols used here are only for distinction.

[0073] In some embodiments, step S400 involves fusing the single-source target video frame result and the single-source target radar frame result at the same time to obtain a cross-device radar-visual fusion result, specifically including:

[0074] At the same time, cross-device fusion matching is performed using the fused video pixel coordinates of each target object in the single-source target video frame result and the fused radar coordinates of each target object in the single-source target radar frame result.

[0075] The successfully matched target radar object and target video object are treated as the same target object, and the average value of the fused radar coordinates of the target radar object and the fused video pixel coordinates of the target video object is calculated as the cross-device fused coordinates of the same target object.

[0076] The cross-device radar-visual fusion result is obtained from the cross-device fusion coordinates of the same target object.

[0077] As can be seen from the aforementioned step S300, the target video frame and target radar frame at the same time have been synchronized. After obtaining the single-source target video frame result and the single-source target radar frame result, cross-device fusion between the video device and the radar device can be performed. Specifically, cross-device fusion matching can be performed using the fused video pixel coordinates of each target object in the single-source target video frame result and the fused radar coordinates of each target object in the single-source target radar frame result, i.e., using the formula... A single-source fusion data matching cost table is constructed for single-source target video frame results and single-source target radar frame results. The size of this cost table is denoted as . Using the Euclidean distance formula described above, the column index j corresponding to the minimum cost of each row i in the single-source fusion data matching cost table is found as the target object id in the target radar frame to be matched. r According to id r Reverse search for the row index corresponding to the minimum cost value as the target object ID in the target video frame for matching. v When ID vWhen = i, it indicates that the target objects of the single-source target video frame result and the single-source target radar frame result have been successfully matched. The average value of the fused radar coordinates of the target radar object and the fused video pixel coordinates of the target video object is calculated, and the cross-device fused coordinates (i.e., cross-device radar-visual fused pixel coordinates) are output. The cross-device fused coordinates of each successfully matched target radar object and target video object are then fused into the same data frame to obtain the cross-device radar-visual fusion result. Otherwise, it indicates that the matching has failed, and the single-source target video frame result and the single-source target radar frame result are retained. Further, the cross-device radar-visual fusion results obtained at different times are summarized to obtain the cross-device radar-visual fusion result set, denoted as: in, The number of target objects in the cross-device radar-visual fusion result is indicated by the subscript f for the coordinates and target object category, which represents the result of the target objects after cross-device radar-visual fusion.

[0078] S500, the tracking trajectory of each target object is obtained based on the fusion coordinates of each target object in the cross-device radar-visual fusion result at different times.

[0079] Specifically, in some embodiments, step S500 includes: constructing a trajectory tracker for each of the target objects, wherein each trajectory tracker stores the historical cross-device fused coordinates of the corresponding target object.

[0080] Trajectory matching is performed using the cross-device fusion coordinates of each target object in the cross-device radar-visual fusion result and the historical cross-device fusion coordinates.

[0081] Add the successfully matched cross-device fused coordinates to the matched trajectory tracker.

[0082] Add the unmatched cross-device fusion coordinates to the newly created trajectory tracker.

[0083] Similarly, the cross-device fused coordinates of each target object at all times are added to the corresponding trajectory tracker to obtain the tracking trajectory of each target object.

[0084] After obtaining the cross-device radar-visual fusion results at different times in step S400, trajectory matching is performed using the cross-device fusion coordinates of each target object in the cross-device radar-visual fusion results and the historical cross-device fusion coordinates. Specifically, redundant calculations can be avoided by calculating the Euclidean distance between the current target object's cross-device fusion coordinates and the previous historical cross-device fusion coordinates in each trajectory tracker, and a fusion data cost table [n] is constructed between each target object in each cross-device radar-visual fusion result and each target object in all trajectory trackers. f ,n t ], where n f n represents the number of target objects in each cross-device radar-visual fusion result.t This represents the total number of trajectory trackers, expressed using the aforementioned Euclidean distance formula in the fusion data cost table [n]. f ,n t Find the column index corresponding to the minimum cost value in each row i as the matching trajectory tracker id. t Based on the matched trajectory tracker ID t Reverse search for the row index corresponding to the minimum cost value as the target object ID for matching radar-visual fusion. f When ID f When the value is equal to i, it indicates that a matching trajectory tracker exists for the target object. The cross-device fused coordinates of the target object are added to the corresponding node in the matching trajectory tracker. Otherwise, it indicates that the match has failed, a new tracker trajectory is created, and the cross-device fused coordinates of the target object are added to the new tracker trajectory. Finally, the trajectory numbers and node list of all target objects are output, completing the tracking of the target objects in the radar-visual fusion. Each trajectory tracker corresponds to a trajectory number, which is implemented by assigning numbers to the trajectory trackers for each target object when constructing them.

[0085] Optionally, in some embodiments, the step of constructing the trajectory tracker for each target object in the above embodiments is not a necessary step and can be constructed in advance. Each target object corresponds to a trajectory tracker, each trajectory tracker corresponds to a trajectory number, and each trajectory tracker stores the historical cross-device fusion coordinates of the corresponding target object. The cross-device fusion coordinates of each target object at different times are sequentially filled into the corresponding node of the corresponding trajectory tracker through the aforementioned matching rules to achieve trajectory tracking of the target object over a period of time.

[0086] According to the target tracking method based on radar-visual fusion provided in the embodiments of this application, target video frames containing target objects from different video devices and target radar frames containing target objects from different radar devices are acquired. The target video pixel coordinates corresponding to the target objects in the target video frames and the target latitude and longitude information of the target objects in the target radar frames are converted into geodetic video coordinates and geodetic radar coordinates in the same geodetic coordinate system, respectively. Then, the target video frames from different video devices at the same time are synchronized, the target radar frames from different radar devices at the same time are synchronized, and the target video frames and target radar frames at the same time are synchronized to obtain a time-synchronized frame dataset with a time synchronization correspondence. Furthermore, the target objects in the synchronized target video frames from different video devices at the same time are fused. The method obtains a fused single-source target video frame result, fuses target objects in synchronous radar frames from different radar devices at the same time, and obtains a fused single-source target radar frame result. Then, it fuses the single-source target video frame result and the single-source target radar frame result at the same time to obtain a cross-device radar-visual fusion result. Finally, it obtains the tracking trajectory of each target object based on the cross-device fused coordinates of each target object in the cross-device radar-visual fusion result at different times. This method can achieve accurate and efficient target fusion of different sensors. Due to the complementary advantages of fusing video data and radar data, it can not only improve the detection accuracy of target objects, but also adapt to complex scenarios such as severe weather. Furthermore, this method can simultaneously adapt to the tracking of multiple target objects and can be applied to large-scale scenarios such as intersections, and achieve adaptive tracking of target objects.

[0087] The target tracking method based on radar-visual fusion provided in this application can be executed by a target tracking device based on radar-visual fusion. This application uses the example of a target tracking device based on radar-visual fusion executing the target tracking method based on radar-visual fusion to illustrate the target tracking device based on radar-visual fusion provided in this application.

[0088] This application also provides a target tracking device based on radar-visual fusion, such as... Figure 2 As shown, the target tracking device based on radar-visual fusion includes: a radar-visual data acquisition module, used to acquire video data from multiple video devices and radar data from multiple radar devices at a traffic intersection, wherein each video device corresponds to multiple frames of video data and each radar device corresponds to multiple frames of radar data.

[0089] The target coordinate extraction module is used to extract target video frames containing target objects from the multi-frame video data, extract target video pixel coordinates corresponding to each target object from each target video frame, extract target radar frames containing target objects from the multi-frame radar data, and extract target latitude and longitude information corresponding to each target object from each target radar frame.

[0090] The conversion synchronization module is used to convert the target video pixel coordinates and the target latitude and longitude information into ground plane video coordinates and ground plane radar coordinates, respectively. It also obtains a time synchronization frame dataset based on the timestamps of the target video frames and the target radar frames, as well as the ground plane video coordinates and the ground plane radar coordinates. The time synchronization frame dataset includes synchronized target video frames with time synchronization correspondences from different video devices and synchronized target radar frames with time synchronization correspondences from different radar devices. The synchronized target video frames and the synchronized target radar frames also have time synchronization correspondences.

[0091] The radar-visual fusion module is used to fuse target objects in synchronous target video frames from different video devices at the same time to obtain a fused single-source target video frame result. It also fuses target objects in synchronous target radar frames from different radar devices at the same time to obtain a fused single-source target radar frame result. Finally, it fuses the single-source target video frame result and the single-source target radar frame result at the same time to obtain a cross-device radar-visual fusion result.

[0092] The trajectory tracking module is used to obtain the tracking trajectory of each target object based on the cross-device fusion coordinates of each target object in the cross-device radar-visual fusion result at different times.

[0093] According to the target tracking device based on radar-visual fusion provided in this application embodiment, target video frames containing target objects from different video devices and target radar frames containing target objects from different radar devices are acquired. The target video pixel coordinates corresponding to the target objects in the target video frames and the target latitude and longitude information of the target objects in the target radar frames are converted into geodetic video coordinates and geodetic radar coordinates in the same plane coordinate system, respectively. Then, the target video frames from different video devices at the same time are synchronized, the target radar frames from different radar devices at the same time are synchronized, and the target video frames and target radar frames at the same time are synchronized, resulting in a time-synchronized frame dataset with a time synchronization correspondence. Furthermore, the target objects in the synchronized target video frames from different video devices at the same time are fused to obtain... The fused single-source target video frame result is obtained by fusing target objects in synchronous radar frames from different radar devices at the same time. Then, the single-source target video frame result and the single-source target radar frame result at the same time are fused to obtain the cross-device radar-visual fusion result. Finally, the tracking trajectory of each target object is obtained based on the cross-device fused coordinates of each target object in the cross-device radar-visual fusion result at different times. This device can achieve accurate and efficient fusion of target objects from different sensors. Due to the complementary advantages of fusing video data and radar data, it can not only improve the detection accuracy of target objects, but also adapt to complex scenarios such as severe weather. Furthermore, this method can simultaneously adapt to the tracking of multiple target objects and can be applied to large-scale scenarios such as intersections, and can achieve adaptive tracking of target objects.

[0094] In some embodiments, the target coordinate extraction module is further configured to extract each target object from each target video frame; and use the bottom midpoint pixel coordinates of each target object as the target video pixel coordinates.

[0095] In some embodiments, the conversion synchronization module is further configured to: determine an initial radar matching frame within a preset range corresponding to the target video frame at the same time based on the timestamps of the target video frame and the target radar frame; calculate the Euclidean distance between each target object in the target video frame at the same time and all target objects in the initial radar matching frame based on the ground plane video coordinates and the ground plane radar coordinates; take the initial radar matching frame where each target object is located when the Euclidean distance is the smallest as the radar matching frame at the time corresponding to the target video frame; and take all target video frames at all times and their corresponding radar matching frames as the time synchronization frame dataset.

[0096] In some embodiments, the radar-visual fusion module is further configured to: select a reference video device from the video devices at a certain moment, and use the remaining video devices as comparison video devices; align each target object in the comparison video devices with the corresponding target object in the reference video devices using the geodetic plane video coordinates of the reference video devices and the geodetic plane video coordinates of the comparison video devices, and use the geodetic plane video coordinates of each target object in each video device after alignment as the aligned video coordinates; calculate the mean of all aligned video coordinates of each target object to obtain the fused video pixel coordinates of each target object; and obtain the single-source target video frame result at the certain moment based on the fused video pixel coordinates of each target object.

[0097] In some embodiments, the radar-visual fusion module is further configured to: select a reference radar device from the radar devices at a certain moment, and use the remaining radar devices as comparison radar devices; align each target object in the comparison radar device with the corresponding target object in the reference radar device using the geodetic radar coordinates of the reference radar device and the geodetic radar coordinates of the comparison radar devices; use the geodetic radar coordinates of each target object in each radar device after alignment as the aligned radar coordinates; calculate the mean of all aligned radar coordinates of each target object to obtain the fused radar coordinates of each target object; and obtain the single-source target radar frame result at the certain moment based on the fused radar coordinates of each target object.

[0098] In some embodiments, the radar-visual fusion module is further configured to perform cross-device fusion matching at the same time using the fused video pixel coordinates of each target object in the single-source target video frame result and the fused radar coordinates of each target object in the single-source target radar frame result; to regard the successfully matched target radar object and target video object as the same target object, and to calculate the average value of the fused radar coordinates of the target radar object and the fused video pixel coordinates of the target video object as the cross-device fusion coordinates of the same target object; and to obtain the cross-device radar-visual fusion result from the cross-device fusion coordinates of each of the same target objects.

[0099] In some embodiments, the trajectory tracking module is further configured to construct trajectory trackers for each target object, each trajectory tracker storing historical cross-device fusion coordinates of the corresponding target object; perform trajectory matching using the cross-device fusion coordinates of each target object and the historical cross-device fusion coordinates in the cross-device radar fusion result; add successfully matched cross-device fusion coordinates to the matched trajectory tracker; add unmatched cross-device fusion coordinates to the newly created trajectory tracker; and so on, adding the cross-device fusion coordinates of each target object at all times to the corresponding trajectory tracker to obtain the tracking trajectory of each target object.

[0100] The target tracking device based on radar-visual fusion in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific type of device.

[0101] The target tracking device based on radar-visual fusion in this application embodiment can be a device with an operating system. This operating system can be Microsoft (Windows), Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0102] The target tracking device based on radar-visual fusion provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0103] In some embodiments, such as Figure 3As shown, this application embodiment also provides an electronic device 300, including a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the program is executed by the processor 301, it implements the various processes of the above-described target tracking method embodiment based on radar-visual fusion and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0104] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0105] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described target tracking method embodiment based on radar-visual fusion and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0106] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0107] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described target tracking method based on radar-visual fusion.

[0108] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0109] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described target tracking method embodiment based on radar-visual fusion, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0110] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0111] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0112] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0113] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0114] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0115] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A target tracking method based on radar-visual fusion, characterized in that, The method includes: The system acquires video data from multiple video devices and radar data from multiple radar devices at a traffic intersection. Each video device corresponds to multiple frames of video data, and each radar device corresponds to multiple frames of radar data. Extract target video frames containing target objects from the multi-frame video data, and extract target video pixel coordinates corresponding to each target object from each target video frame; extract target radar frames containing target objects from the multi-frame radar data, and extract target latitude and longitude information corresponding to each target object from each target radar frame. The target video pixel coordinates and the target latitude and longitude information are converted into ground plane video coordinates and ground plane radar coordinates, respectively. A time synchronization frame dataset is obtained based on the timestamps of the target video frames and the target radar frames, as well as the ground plane video coordinates and the ground plane radar coordinates. The time synchronization frame dataset includes synchronized target video frames with time synchronization correspondences from different video devices and synchronized target radar frames with time synchronization correspondences from different radar devices. There is also a time synchronization correspondence between the synchronized target video frames and the synchronized target radar frames. The target objects in the synchronous target video frames of different video devices at the same time are fused to obtain the fused single-source target video frame result. The target objects in the synchronous target radar frames of different radar devices at the same time are fused to obtain the fused single-source target radar frame result. The single-source target video frame result and the single-source target radar frame result at the same time are fused to obtain the cross-device radar-visual fusion result. The tracking trajectory of each target object is obtained based on the cross-device fusion coordinates of each target object in the cross-device radar-visual fusion results at different times.

2. The method according to claim 1, characterized in that, Extracting the target video pixel coordinates corresponding to each target object from each of the target video frames includes: Extract each target object from each of the target video frames; The bottom midpoint pixel coordinates of each target object are used as the target video pixel coordinates.

3. The method according to claim 2, characterized in that, The process of obtaining the time-synchronized frame dataset based on the timestamps of the target video frame and the target radar frame, as well as the ground plane video coordinates and the ground plane radar coordinates, includes: Based on the timestamps of the target video frame and the target radar frame, determine the initial radar matching frame within a preset range corresponding to the target video frame at the same time. Calculate the Euclidean distance between each target object in the target video frame and all target objects in the initial radar matching frame at the same time based on the ground plane video coordinates and the ground plane radar coordinates. Take the initial radar matching frame where each target object is located when the Euclidean distance is the smallest as the radar matching frame at the corresponding time of the target video frame. The target video frames and their corresponding radar matching frames at all times are used as the time synchronization frame dataset.

4. The method according to claim 3, characterized in that, The process of fusing target objects in synchronized target video frames from different video devices at the same time to obtain a fused single-source target video frame result includes: At a certain moment, a reference video device is selected from the video devices, and the remaining video devices are used as comparison video devices; Using the geodetic plane video coordinates of the reference video device and the geodetic plane video coordinates of the comparison video device, each target object in the comparison video device is aligned with the corresponding target object in the reference video device, and the geodetic plane video coordinates of each target object after alignment are used as the aligned video coordinates; Calculate the mean of all aligned video coordinates of each target object to obtain the fused video pixel coordinates of each target object; The single-source target video frame result at a certain moment is obtained based on the fused video pixel coordinates of each target object.

5. The method according to claim 4, characterized in that, The step of fusing target objects in synchronous radar frames from different radar devices at the same time to obtain a fused single-source target radar frame result includes: At a certain moment, a reference radar device is selected from the radar devices, and the remaining radar devices are used as comparison radar devices. The target objects in the comparison radar device are aligned with the corresponding target objects in the reference radar device using the geodetic radar coordinates of the reference radar device and the geodetic radar coordinates of the comparison radar device. The geodetic radar coordinates of the aligned target objects are then used as the alignment radar coordinates. Calculate the mean of all aligned radar coordinates of each target object to obtain the fused radar coordinates of each target object; The single-source target radar frame result at a certain moment is obtained based on the fused radar coordinates of each target object.

6. The method according to claim 5, characterized in that, The process of fusing the single-source target video frame results and the single-source target radar frame results at the same time to obtain the cross-device radar-visual fusion result includes: At the same time, cross-device fusion matching is performed using the fused video pixel coordinates of each target object in the single-source target video frame result and the fused radar coordinates of each target object in the single-source target radar frame result; The successfully matched target radar object and target video object are treated as the same target object, and the average value of the fused radar coordinates of the target radar object and the fused video pixel coordinates of the target video object is calculated as the cross-device fused coordinates of the same target object. The cross-device radar-visual fusion result is obtained from the cross-device fusion coordinates of the same target object.

7. The method according to claim 6, characterized in that, The step of obtaining the tracking trajectory of each target object based on the cross-device fusion coordinates of each target object in the cross-device radar-visual fusion results at different times includes: Construct trajectory trackers for each of the target objects, and each trajectory tracker stores the historical cross-device fused coordinates of the corresponding target object; Trajectory matching is performed using the cross-device fusion coordinates and historical cross-device fusion coordinates of each target object in the cross-device radar-visual fusion result; Add the successfully matched cross-device fused coordinates to the matched trajectory tracker; Add the unmatched cross-device fusion coordinates to the newly created trajectory tracker; Similarly, the cross-device fused coordinates of each target object at all times are added to the corresponding trajectory tracker to obtain the tracking trajectory of each target object.

8. A target tracking device based on radar-visual fusion, characterized in that, The device includes: The radar data acquisition module is used to acquire video data from multiple video devices and radar data from multiple radar devices at traffic intersections. Each video device corresponds to multiple frames of video data, and each radar device corresponds to multiple frames of radar data. The target coordinate extraction module is used to extract target video frames containing target objects from the multi-frame video data, extract target video pixel coordinates corresponding to each target object from each target video frame, extract target radar frames containing target objects from the multi-frame radar data, and extract target latitude and longitude information corresponding to each target object from each target radar frame. The conversion synchronization module is used to convert the target video pixel coordinates and the target latitude and longitude information into ground plane video coordinates and ground plane radar coordinates, respectively, and to obtain a time synchronization frame dataset based on the timestamps of the target video frames and the target radar frames, as well as the ground plane video coordinates and the ground plane radar coordinates. The time synchronization frame dataset includes synchronized target video frames with time synchronization correspondences from different video devices and synchronized target radar frames with time synchronization correspondences from different radar devices. There is also a time synchronization correspondence between the synchronized target video frames and the synchronized target radar frames. The radar-visual fusion module is used to fuse target objects in synchronous target video frames of different video devices at the same time to obtain a fused single-source target video frame result, fuse target objects in synchronous target radar frames of different radar devices at the same time to obtain a fused single-source target radar frame result, and fuse the single-source target video frame result and the single-source target radar frame result at the same time to obtain a cross-device radar-visual fusion result. The trajectory tracking module is used to obtain the tracking trajectory of each target object based on the cross-device fusion coordinates of each target object in the cross-device radar-visual fusion result at different times.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the target tracking method based on radar-visual fusion as described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the target tracking method based on radar-visual fusion as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Target detection method and device based on Leiyu fusion and readable storage medium

    CN115082712A

  • Expressway multi-target tracking method based on Leiyu fusion

    CN116863382A