Intersection view blind area vehicle tracking method based on fusion perception technology
Through multi-way traffic camera recognition and image coordinate conversion, the target-level fusion map of intersections is established, which solves the problem of intersection traffic accidents caused by driver blind spots, and achieves effective collision avoidance warning and traffic safety improvement.
Patent Information
- Application Number
- CN202510320481.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, traffic accidents at intersections caused by driver blind spots occur frequently, especially at intersections where adjacent lanes cannot be identified due to blockage of adjacent lanes, and there is a lack of effective collision avoidance and early warning equipment and technology.
Using a method based on fusion perception technology, vehicle recognition is performed through multiple traffic cameras, and an intersection target-level fusion map is established using image and real coordinate transformation, target tracking and early warning is sent.
Effective tracking of vehicles in the blind spots of the intersection’s visual field has been achieved, the collision avoidance warning at the intersection has been promoted, and traffic safety has been improved.
Smart Images

Figure CN120374679A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation technologies, and particularly to a method and device for tracking vehicles in the blind area of intersection vision based on fusion perception technology. Background Art
[0002] Traffic accidents caused by driver blind spots occur frequently. The so-called "ghost vehicle" problem at intersections, that is, pedestrians, vehicles, etc. suddenly break into the blind area of vision or the braking reaction time is insufficient, resulting in collision accidents. For example, at a crossroads, due to the blockage of large vehicles in adjacent lanes, vehicle drivers and intelligent vehicles cannot identify and avoid vehicles and vulnerable road users that suddenly appear in front, and there are no corresponding devices and relevant research at present intersections. Summary of the Invention
[0003] This application aims to solve at least one of the technical problems in the related technologies to some extent.
[0004] To this end, the first object of this application is to propose a method for tracking vehicles in the blind area of intersection vision based on fusion perception technology, which realizes the tracking of vehicles in the blind area of intersection vision and effectively promotes collision avoidance warning and traffic safety at intersections.
[0005] The second object of this application is to propose a device for tracking vehicles in the blind area of intersection vision based on fusion perception technology.
[0006] To achieve the above object, an embodiment of the first aspect of this application proposes a method for tracking vehicles in the blind area of intersection vision based on fusion perception technology, including: using a target detection algorithm to respectively perform vehicle recognition on the video data of each road collected by multiple traffic cameras at intersections to obtain recognition results, where the recognition results are the image coordinates of vehicles on each road; using a coordinate conversion method between images and the real world to process the recognition results to obtain the real-world coordinates of each vehicle; based on the real-world coordinates of each vehicle, mapping the vehicles one by one onto the top-view map of the intersection through the mapping relationship between the top-view map coordinates and the real-world coordinates to obtain the target-level fusion map of the intersection; performing target tracking on the vehicles at the intersection based on their positions in the target-level fusion map of the intersection, and sending the tracking data to the target vehicles waiting at the intersection.
[0007] The method for tracking vehicles in the blind area of intersection vision according to the embodiment of this application uses a target detection algorithm to perform target recognition on multiple traffic cameras, and then uses coordinate conversion technology between images and the real world to perform multi-sensor target-level mapping at intersections. Thus, from two aspects of target detection and multi-sensor mapping, a target-level fusion map for intersections is established, providing an important theoretical basis and technical support for the research and development of collision avoidance warning technology at intersections and the improvement of traffic safety.
[0008] Optionally, in an embodiment of the present application, the image coordinates of the target vehicle are p(u, v), and the real-world coordinates of the target vehicle are X C , Y C , Z C which are expressed as:
[0009]
[0010] where X W , Y W , Z W are the positions of the camera in the world coordinate system, R is the rotation matrix, and T is the transformation matrix.
[0011] Optionally, in an embodiment of the present application, the expression of the mapping relationship between the top-view map coordinates and the real-world coordinates is:
[0012] x cw = x rw - L x
[0013] y cw = y rw + L y
[0014] where x cw , y cw are the top-view map coordinates of the vehicle, x rw , y rw represent the real-world coordinates of the vehicle, and L x , y cw represents the space between the intersection points of the real-world coordinate system and the top-view map coordinate system on the x-axis and y-axis;
[0015] The formula for converting the real-world coordinates of the target vehicle into the coordinates in the intersection target-level fusion map is:
[0016]
[0017] where H and θ are the installation height and pitch angle of the multi-lane traffic camera, C x , C y are the offsets of the optical axis, L x , L y is the space between the intersection points of the real-world coordinate system and the top-view map coordinate system on the x-axis and y-axis, and x rw , y rw are the real-world coordinates of the detection target.
[0018] Optionally, in an embodiment of the present application, selecting the target vehicle includes:
[0019] Select a target vehicle according to the position of the vehicle and the traffic signal conditions at the intersection;
[0020] Perform target tracking based on the position of the target vehicle in the target-level fusion map, including:
[0021] Track the target vehicle through a target tracking algorithm to obtain the actual position data of the target vehicle;
[0022] Use the actual position data of the target vehicle to predict the trajectory of the target vehicle to obtain the driving state and position of the target vehicle;
[0023] Send an alarm through roadside equipment and send the target vehicle data to the waiting vehicles.
[0024] To achieve the above object, an embodiment of the second aspect of the present invention proposes a vehicle tracking device in the blind area of the intersection vision based on the fusion perception technology, including a target recognition module, a fusion module, a mapping module, and a target tracking module, where,
[0025] The target recognition module is used to use a target detection algorithm to separately perform vehicle recognition on the video data of each road collected by the multi-way traffic cameras at the intersection to obtain a recognition result, where the recognition result is the image coordinates of the vehicles on each road;
[0026] The fusion module is used to process the recognition result by using a coordinate conversion method between the image and the real world to obtain the real world coordinates of each vehicle;
[0027] The mapping module is used to map each vehicle to the intersection top view map one by one based on the real world coordinates of each vehicle through the mapping relationship between the top view map coordinates and the real world coordinates to obtain the intersection target-level fusion map;
[0028] The target tracking module is used to perform target tracking on the vehicles at the intersection based on the position in the intersection target-level fusion map and send the tracking data to the target vehicles waiting at the intersection.
[0029] Optionally, in an embodiment of the present application, the image coordinates of the target vehicle are p(u, v), and the real world coordinates X C 、Y C 、Z C Are expressed as:
[0030]
[0031] Among them, X W 、Y w 、Z w Is the position of the camera in the world coordinate system, R is the rotation matrix, and T is the transformation matrix.
[0032] Optionally, in an embodiment of the present application, the expression of the mapping relationship between the top-view map coordinates and the real-world coordinates is:
[0033] x cw = x rw - L x
[0034] y cw = y rw + L y
[0035] Wherein, x cw , y cw are the top-view map coordinates of the vehicle, x rw , y rw represent the real-world coordinates of the vehicle, and L x , y cw represents the space between the intersection of the real-world coordinate system and the top-view map coordinate system on the x-axis and the y-axis;
[0036] The formula for converting the real-world coordinates of the target vehicle into the coordinates in the intersection target-level fusion map is:
[0037]
[0038] Wherein, H and θ are the installation height and pitch angle of the multi-lane traffic camera, C x , C y are the offsets of the optical axis, L x , L y is the space between the intersection of the real-world coordinate system and the top-view map coordinate system on the x-axis and the y-axis, and x rw , y rw are the real-world coordinates of the detection target.
[0039] Optionally, in an embodiment of the present application, selecting the target vehicle includes:
[0040] Selecting the target vehicle according to the position of the vehicle and the traffic signal conditions at the intersection;
[0041] Target tracking based on the position of the target vehicle in the target-level fusion map includes:
[0042] Tracking the target vehicle through the target tracking algorithm to obtain the actual position data of the target vehicle;
[0043] Predicting the trajectory of the target vehicle using the actual position data of the target vehicle to obtain the driving state and position of the target vehicle;
[0044] Issuing an alarm through roadside equipment and sending the target vehicle data to the waiting vehicles.
[0045] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0046] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0047] Figure 1 It is a schematic flow chart of a method for tracking vehicles in the blind area of intersection vision based on fusion perception technology provided in Embodiment 1 of the present application;
[0048] Figure 2 It is a flow chart for constructing a map of the intersection in the embodiment of the present application;
[0049] Figure 3 It is a schematic diagram of the conversion relationship between the detection target coordinate system and the mapped image coordinate system of the intersecting image in the embodiment of the present application;
[0050] Figure 4 It is a flow chart for target tracking in the embodiment of the present application;
[0051] Figure 5 It is a schematic structural diagram of a device for tracking vehicles in the blind area of intersection vision based on fusion perception technology provided in the embodiment of the present application; Detailed Description of the Embodiments
[0052] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0053] The method and device for tracking vehicles in the blind area of intersection vision based on fusion perception technology in the embodiments of the present application will be described below with reference to the drawings.
[0054] Figure 1 It is a schematic flow chart of a method for tracking vehicles in the blind area of intersection vision based on fusion perception technology provided in Embodiment 1 of the present application.
[0055] As Figure 1 shown, the method for tracking vehicles in the blind area of intersection vision based on fusion perception technology includes the following steps:
[0056] Step 101, use a target detection algorithm to respectively identify vehicles in each road of the video data collected by multiple traffic cameras at the intersection, and obtain the identification result, where the identification result is the image coordinates of the vehicles on each road;
[0057] Step 102: Process the recognition result using the coordinate conversion method between the image and the real world to obtain the real-world coordinates of each vehicle;
[0058] Step 103: Based on the real-world coordinates of each vehicle, map the vehicles one by one onto the intersection top-view map through the mapping relationship between the top-view map coordinates and the real-world coordinates to obtain the intersection target-level fusion map;
[0059] Step 104: Track the vehicles at the intersection based on their positions in the intersection target-level fusion map, and send the tracking data to the target vehicles waiting at the intersection.
[0060] The method for tracking vehicles in the blind area of the intersection vision based on the fusion perception technology in the embodiment of the present application uses a target detection algorithm to perform target recognition on multiple traffic cameras, and then adopts the coordinate conversion technology between the image and the real world to perform multi-sensor target-level mapping at the intersection. By integrating pedestrians, motor vehicles, etc. at the intersection into the intersection target-level fusion map in the intersection traffic scenario, it solves the collision problem in the blind area of the intersection vision, and establishes a target-level fusion map in other multi-sensor target-level fusion scenarios to solve problems such as multi-camera target tracking.
[0061] Optionally, in an embodiment of the present application, the image coordinates of the target vehicle are p(u, v), and the real-world coordinates X C 、Y C 、Z C are expressed as:
[0062]
[0063] where X W 、Y W 、Z W are the positions of the camera in the world coordinate system, R is the rotation matrix, and T is the transformation matrix.
[0064] Optionally, in an embodiment of the present application, the expression of the mapping relationship between the top-view map coordinates and the real-world coordinates is:
[0065] x cw = x rw - L x
[0066] y cw = y rw + L y
[0067] where x cw 、y cw are the top-view map coordinates of the vehicle, x rw 、yrw Represents the real - world coordinates of the vehicle, L x , y cw Represents the space between the intersection points of the real - world coordinate system and the top - view map coordinate system on the x - axis and y - axis;
[0068] The formula for converting the real - world coordinates of the target vehicle into coordinates in the intersection target - level fusion map is:
[0069]
[0070] where H and θ are the installation height and pitch angle of the multi - traffic camera, C x , C y is the offset of the optical axis, L x , L y is the space between the intersection points of the real - world coordinate system and the top - view map coordinate system on the x - axis and y - axis, x rw , y rw are the real - world coordinates of the detection target.
[0071] Optionally, in an embodiment of the present application, selecting the target vehicle includes:
[0072] Selecting the target vehicle according to the position of the vehicle and the traffic signal conditions at the intersection;
[0073] Based on the position of the target vehicle in the target - level fusion map for target tracking, including:
[0074] Tracking the target vehicle through a target - tracking algorithm to obtain the actual position data of the target vehicle;
[0075] Using the actual position data of the target vehicle to predict the trajectory of the target vehicle to obtain the driving state and position of the target vehicle;
[0076] Sending an alarm through roadside equipment and sending the target vehicle data to the waiting vehicles.
[0077] The following describes in detail the method for tracking vehicles in the blind area of the intersection vision based on the fusion perception technology of the present application through two specific embodiments.
[0078] Embodiment 1
[0079] This embodiment introduces the concept of "intersection target - level fusion map", combines the existing sensors at the intersection, target detection technology and image coordinate conversion technology to establish an intersection target - level fusion map.
[0080] Such as Figure 2As shown in (a), select 4 existing surveillance videos at this intersection, and use computer vision technology to perform object recognition on the 4 surveillance videos at this intersection respectively. Use the conversion algorithm from image coordinates to real-world coordinates to calculate the real-world position of the vehicle object and the mapping relationship between the top-view map coordinates and the real-world coordinates.
[0081] By calculating the position of the target vehicle in the real-world coordinates, fuse the targets in the four images shown in Figure 2 (b) and Figure 2 (c). Then, through the mapping relationship between the top view and the real-world coordinates, map the fused targets to the intersection top-view map one by one to obtain the intersection map image shown in Figure 2 (d).
[0082] Embodiment 2
[0083] Taking a common four-way intersection as an example, the process of tracking the target vehicle includes:
[0084] (1) Based on the four cameras at the intersection, perform object detection on the four images respectively to obtain the pixel coordinates p(u, v) of each target relative to each target point; (2)
[0086]
[0087] Calculate the transformation of a point from the image coordinate system to the world coordinate system by the above formula, where R is the rotation matrix, T is the transformation matrix, Z_C is the (non-zero) scale factor, is the effective focal length (the distance from the optical center to the image plane), is the point in the homogeneous coordinate space in the camera coordinate system, and is the homogeneous coordinate point of the image in the image coordinate system. dx, dy are the pixel sizes, and u0, v0 are the image centers. Are the normalized focal lengths on the x and y axes respectively. X W Y W Z w : The position of the camera in the world coordinate system. X C Y C Δ C : The position of a point in the image in the world coordinate system. Through the above formula, X C , Y C , Z C can be calculated from p(u, v).
[0088] (3) Figure 3 Shows the conversion relationship between the detected target coordinate system and the mapped image coordinate system of the intersecting image. Among them, O cw -x cw y cw z cw: Coordinate system of the intersection image map, where the origin is the projection point of the intersection image map on the ground. O rw -x rw y rw z rw : Detection target coordinate system. The intersection image map coordinate system and the camera projection coordinate system are two parallel coordinate systems in space, and their spatial relative relationship is shown in the figure. O p -x p y p : Intersection image mapping image coordinate system, with the origin located at the upper left corner of the image; O c -x c y c z c : Detection camera coordinate system. According to Figure 2 , the formula for converting an object in the camera coordinate system to the intersection image mapping image coordinate system is:
[0089]
[0090] H, θ are the installation height and pitch angle of the camera, c x , C y is the offset of the optical axis, L x , L y is the space between the intersection of the target coordinate system and the image mapping image coordinate system on the x-axis and y-axis. The position of the target in the target-level fusion map can be calculated by the formula, and then target tracking can be carried out.
[0091] (4) Figure 4 is the single-target trajectory prediction process. Select the target vehicle according to the position of the target vehicle and the traffic signal conditions at the intersection. The target tracking algorithm is used to track the target vehicle and store its actual position data. Then, the trajectory of the target vehicle is predicted using the position database of the target vehicle to obtain the driving state and position of the target vehicle; finally, an alarm is issued using roadside equipment, and the target vehicle data is sent to the waiting vehicles.
[0092] such as Figure 4 shown, the tracking and prediction algorithm includes:
[0093] Create the corresponding trajectory of the target vehicle detected in the first frame. Initialize the motion variables of the Kalman filter and predict its position in the next frame through the Kalman filter.
[0094] In the second frame, perform IOU matching on the position of the target detection in this frame and the position predicted through the trajectory in the previous frame one by one, and then calculate its cost matrix through the result of the IOU matching.
[0095] Taking the cost matrix as the input of the Hungarian algorithm, three linear matching results are obtained. The first is trajectory mismatch. The mismatched uncertain trajectories are deleted, and the certain trajectories need to wait for a period of time before being deleted. The second is detection mismatch. This detection is initialized as a new trajectory and marked as a confirmed trajectory. The third is that the detected target and the predicted target are successfully paired, indicating successful tracking. The trajectory is marked as a confirmed trajectory, and the corresponding detection position is used to update the corresponding trajectory variable through Kalman filtering to obtain the predicted trajectory.
[0096] To implement the above embodiments, the present application also proposes a vehicle tracking device in the blind area of intersection vision based on fusion perception technology.
[0097] Figure 5 It is a schematic structural diagram of a vehicle tracking device in the blind area of intersection vision based on fusion perception technology provided by an embodiment of the present application.
[0098] As Figure 5 shown, the vehicle tracking device in the blind area of intersection vision based on fusion perception technology includes a target recognition module, a fusion module, a mapping module, and a target tracking module. Among them,
[0099] The target recognition module is used to use a target detection algorithm to separately perform vehicle recognition on the video data of each road collected by the multi-way traffic cameras at the intersection, and obtain the recognition results. Among them, the recognition results are the image coordinates of the vehicles on each road.
[0100] The fusion module is used to process the recognition results by using the coordinate conversion method between the image and the real world to obtain the real world coordinates of each vehicle.
[0101] The mapping module is used to map each vehicle to the intersection top view map one by one based on the real world coordinates of each vehicle through the mapping relationship between the top view map coordinates and the real world coordinates, and obtain the intersection target-level fusion map.
[0102] The target tracking module is used to perform target tracking on the vehicles at the intersection based on their positions in the intersection target-level fusion map, and send the tracking data to the target vehicle waiting at the intersection.
[0103] Optionally, in an embodiment of the present application, the image coordinates of the target vehicle are p(u, v), and the real world coordinates X C 、Y C 、Z C are expressed as:
[0104]
[0105] Among them, X W 、Y W 、ZW Where \(P\) is the position of the camera in the world coordinate system, \(R\) is the rotation matrix, and \(T\) is the transformation matrix.
[0106] Optionally, in an embodiment of the present application, the expression of the mapping relationship between the top - view map coordinates and the real - world coordinates is:
[0107] x cw = x rw - L x
[0108] y cw = y rw + L y
[0109] Where \(x\) cw , \(y\) cw are the top - view map coordinates of the vehicle, \(x\) rw , \(y\) rw represent the real - world coordinates of the vehicle, and \(L\) x , \(y\) cw represents the space between the intersection of the real - world coordinate system and the top - view map coordinate system on the \(x\) - axis and \(y\) - axis;
[0110] The formula for converting the real - world coordinates of the target vehicle into the coordinates in the intersection target - level fusion map is:
[0111]
[0112] Where \(H\) and \(\theta\) are the installation height and pitch angle of the multi - lane traffic camera, \(C\) x , \(C\) y are the offsets of the optical axis, \(L\) x , \(L\) y is the space between the intersection of the real - world vehicle coordinate system and the top - view map coordinate system on the \(x\) - axis and \(y\) - axis, and \(x\) rw , \(y\) rw are the real - world coordinates of the detection target.
[0113] Optionally, in an embodiment of the present application, selecting the target vehicle includes:
[0114] Selecting the target vehicle according to the position of the vehicle and the traffic signal conditions at the intersection;
[0115] Target tracking based on the position of the target vehicle in the target - level fusion map includes:
[0116] Tracking the target vehicle through a target - tracking algorithm to obtain the actual position data of the target vehicle;
[0117] Predicting the trajectory of the target vehicle using the actual position data of the target vehicle to obtain the driving state and position of the target vehicle;
[0118] An alarm is issued through roadside equipment, and the target vehicle data is sent to the waiting vehicles.
[0119] It should be noted that the foregoing explanation of the embodiment of the vehicle tracking method in the blind area of the intersection vision based on the fusion perception technology is also applicable to the vehicle tracking device in the blind area of the intersection vision based on the fusion perception technology of this embodiment, and will not be repeated here.
[0120] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "examples", "specific examples" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0121] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0122] Any process or method description in the flowchart or described in other ways herein may be understood to represent a module, segment or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0123] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0124] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0125] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0126] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0127] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A vehicle tracking method in the blind area of intersection vision based on fusion perception technology, characterized in that It includes the following steps: Use the object detection algorithm to separately identify vehicles in the video data of each road collected by the multi-way traffic cameras at the intersection, and obtain the recognition results, where the recognition results are the image coordinates of the vehicles on each road; Adopt the coordinate conversion method between the image and the real world to process the recognition results, and obtain the real-world coordinates of each vehicle; Based on the real-world coordinates of each vehicle, map the vehicles one by one onto the top-view map of the intersection through the mapping relationship between the top-view map coordinates and the real-world coordinates, and obtain the target-level fusion map of the intersection; Track the vehicles at the intersection based on their positions in the target-level fusion map of the intersection, and send the tracking data to the target vehicles waiting at the intersection.
2. The method according to claim 1, wherein The image coordinates of the target vehicle are p(u, v), and the real-world coordinates X C , Y C , Z C are expressed as: Among them, X W , Y w , Z W are the positions of the camera in the world coordinate system, R is the rotation matrix, and T is the transformation matrix.
3. The method according to claim 2, wherein The expression of the mapping relationship between the top-view map coordinates and the real-world coordinates is: x cw = x rw - L x y cw = y rw + L y where x cw and y cw are the map coordinates of the top view of the vehicle, and x rw and y rw represent the real-world coordinates of the vehicle. L x and y cw represent the space between the intersections of the real-world coordinate system and the top-view map coordinate system on the x-axis and y-axis; The formula for converting the real-world coordinates of the target vehicle into the coordinates in the target-level fusion map of the intersection is: Among them, H and θ are the installation height and pitch angle of the multi-lane traffic camera, C x and C y are the offsets of the optical axis, L x and L y is the space between the intersections of the real-world coordinate system and the top-view map coordinate system on the x-axis and y-axis, x rw and y rw are the real-world coordinates of the detection target.
4. The method according to claim 1, wherein The selection of the target vehicle includes: Select the target vehicle according to the position of the vehicle and the traffic signal conditions at the intersection; The target tracking based on the position of the target vehicle in the target-level fusion map includes: Track the target vehicle through the target tracking algorithm to obtain the actual position data of the target vehicle; Use the actual position data of the target vehicle to predict the trajectory of the target vehicle, and obtain the driving state and position of the target vehicle; Send out an alarm through the roadside device and send the target vehicle data to the waiting vehicles.
5. A vehicle tracking device in the blind area of intersection vision based on fusion perception technology, characterized in that, It includes an object recognition module, a fusion module, a mapping module, and a target tracking module, where The object recognition module is used to use the object detection algorithm to separately identify vehicles in the video data of each road collected by the multi-way traffic cameras at the intersection, and obtain the recognition results, where the recognition results are the image coordinates of the vehicles on each road; The fusion module is used to adopt the coordinate conversion method between the image and the real world to process the recognition results, and obtain the real-world coordinates of each vehicle; The mapping module is used to, based on the real-world coordinates of each vehicle, map the vehicles one by one onto the top-view map of the intersection through the mapping relationship between the top-view map coordinates and the real-world coordinates, and obtain the target-level fusion map of the intersection; The target tracking module is used to track the vehicles at the intersection based on their positions in the target-level fusion map of the intersection, and send the tracking data to the target vehicles waiting at the intersection.
6. The device according to claim 5, characterized in that The image coordinates of the target vehicle are p(u, v), and the real-world coordinates X C , Y C , Z C are expressed as: Among them, X W , Y W , Z W are the positions of the camera in the world coordinate system, R is the rotation matrix, and T is the transformation matrix.
7. The device according to claim 5, characterized in that, The expression of the mapping relationship between the top-view map coordinates and the real-world coordinates is: x cw = x rw - L x y cw = y rw + L y Among them, x cw and y cw are the map coordinates of the top view of the vehicle, and x rw and y rw represent the real-world coordinates of the vehicle. L x and y cw represent the space between the intersections of the real-world coordinate system and the top-view map coordinate system on the x-axis and the y-axis; The formula for converting the real-world coordinates of the target vehicle into the coordinates in the target-level fusion map of the intersection is: Among them, H and θ are the installation height and pitch angle of the multi-lane traffic camera, C x , C y are the offsets of the optical axis, L x , L y is the space between the intersections of the real-world coordinate system and the top-view map coordinate system on the x-axis and y-axis, x rw , y rw are the real-world coordinates of the detection target.
8. The device according to claim 5, characterized in that, The selection of the target vehicle includes: Select the target vehicle according to the position of the vehicle and the traffic signal conditions at the intersection; The target tracking based on the position of the target vehicle in the target-level fusion map includes: Track the target vehicle through the target tracking algorithm to obtain the actual position data of the target vehicle; Use the actual position data of the target vehicle to predict the trajectory of the target vehicle, and obtain the driving state and position of the target vehicle; Send out an alarm through the roadside device and send the target vehicle data to the waiting vehicles.