Multi-perspective collaborative perception method and device for highway traffic load spatiotemporal distribution pattern enhanced by boundary reasoning
By combining side and top-view video acquisition devices with data processing, and utilizing boundary inference algorithms and multi-target tracking detection, accurate correction of vehicle trajectories and effective integration of load information are achieved, solving the trajectory distortion problem at the boundaries of the field of view in existing systems and providing reliability and cost-effectiveness for traffic control and infrastructure assessment.
Patent Information
- Application Number
- CN202510136962.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The existing highway traffic monitoring system cannot effectively monitor the vehicle trajectory, especially after the vehicle moves to the edge of the top-down field of view, which leads to trajectory distortion and load identification anomalies. In addition, the existing system cannot effectively integrate vehicle geometric characteristics and spatial position information.
By combining side and top-view video acquisition equipment with data processing equipment, and using a boundary inference enhancement algorithm, the vehicle tire image is identified and the vehicle trajectory is inferred. A non-contact weighing system is used to obtain load information. Combined with multi-target tracking detection and trajectory correction algorithms, collaborative perception of the spatiotemporal distribution of vehicle loads is achieved.
It achieves accurate correction of vehicle trajectories and effective integration of load information, provides reliability and cost-effectiveness for traffic control and infrastructure assessment, and solves the problem of trajectory recognition anomalies at the boundaries of the field of view in traditional methods.
Smart Images

Figure CN119942420B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of highway vehicle load identification, and in particular to a boundary reasoning-enhanced multi-perspective collaborative perception method and device for spatiotemporal distribution patterns of highway traffic loads. Background Art
[0002] With the increasing demand for road transportation, new and renovated highways are trending towards wider pavement and increased lane count. To more effectively study the impact of traffic loads on highway and bridge infrastructure and implement traffic flow control and anomaly warnings, it's necessary not only to obtain vehicle deadweight but also to understand the movement trajectory of each vehicle within a specific range to establish a dynamic distribution model of traffic loads for key sections.
[0003] Vehicle-in-motion weighing (WIM) technology has been implemented on many highways and urban roads in China. However, existing contact / non-contact WIM systems only measure vehicle weight and are unable to monitor vehicle trajectory. To address this issue, road surveillance cameras can be used as auxiliary visual sensing devices to identify vehicles and track their trajectories using pre-trained neural networks. However, existing road surveillance system cameras typically use an oblique viewing angle, making trajectory quantification difficult. Mounting the camera perpendicular to the road surface (using a vertical, top-down perspective) facilitates trajectory quantification. However, when a vehicle reaches the edge of the top-down perspective, portions of the vehicle body are outside the visible range, causing the length and width of the vehicle detection frame to continuously change. Directly extracting the center trajectory of the vehicle detection frame using existing multi-target tracking detection techniques can produce anomalous results, such as distorted trajectory shape and unexpected reductions in calculated vehicle speed. Furthermore, current video-based vehicle tracking and identification technologies mostly focus solely on the location of the detected object, without identifying and analyzing the object's geometric features. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention proposes a multi-perspective collaborative perception method and device for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning. The specific technical solutions are as follows:
[0005] A multi-view collaborative perception device for the spatiotemporal distribution pattern of highway traffic loads enhanced by boundary reasoning, a side-view video acquisition device, a top-view video acquisition device, and a data processing device;
[0006] The side-view video acquisition device is installed on the side of each lane of the immutable road section to identify the vehicle load sequence of the corresponding lane; the camera installation of the side-view video acquisition device must meet the following conditions:
[0007] (1) The camera's image plane is perpendicular to the road surface;
[0008] (2) The main optical axis of the camera is perpendicular to the direction of the road;
[0009] (3) The camera should be installed at a height that allows the complete capture of tire images when a vehicle passes through its field of view.
[0010] The overhead video capture device is installed along the road extension direction and immediately after the immutable road segment, serving as an overhead information source to obtain the vehicle's driving trajectory. The installation of the overhead video capture device meets the following conditions:
[0011] (1) The principal optical axis of the lens is perpendicular to the road surface, and the horizontal axis of the image coordinate axis is parallel to the longitudinal axis of the road;
[0012] (2) The size of the visible area is more than twice the full length of a conventional vehicle in the longitudinal direction of the road, and covers all lanes in the same direction on one side of the road in the transverse direction;
[0013] (3) The vertical position of the top-view video capture device is immediately behind the side view area covered by the side view video capture device, and the vehicle within the visual range of the top-view video capture device enters a non-laneable area of a preset length on one side, so as to ensure that the lane where the vehicle is located within the side view range can be obtained in the top-view video, which is convenient for subsequent data association;
[0014] The data processing device has a built-in non-contact vehicle dynamic weighing analysis system, which obtains the real-time inflation pressure of the tires of a moving vehicle and the vehicle tire image data obtained from the side view through the vehicle-road cooperative network to solve the vehicle's single wheel load and the vehicle load; the vehicle running trajectory is solved by obtaining the vehicle video information from the top view, and the vehicle load information obtained from the side view is associated with the vehicle spatial trajectory obtained from the top view to obtain the temporal and spatial distribution of the vehicle load.
[0015] Furthermore, the bird's-eye view video acquisition device has multiple bird's-eye view angles along the longitudinal direction of the road, and adjacent bird's-eye view angles satisfy: the field of view of each bird's-eye view angle overlaps to a certain extent with the next bird's-eye view angle, and each bird's-eye view video time frame is recorded using the same benchmark, and each bird's-eye view angle uses a unified road surface coordinate system.
[0016] A boundary reasoning-enhanced multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads is implemented based on a boundary reasoning-enhanced multi-perspective collaborative perception device for highway traffic load distribution patterns, and includes the following steps:
[0017] S1: The side view video capture device of each lane collects tire images of each vehicle passing through the lane, and calculates the single wheel load of each vehicle based on the tire images, thereby obtaining the total load of the vehicle;
[0018] S2: When several vehicles exit the side view's non-lane-changeable zone and enter the top view's visible range, the top view video capture device begins continuous video capture and uses a pre-trained convolutional neural network to identify the vehicles frame by frame, obtaining detection frames and corresponding confidence scores. Subsequently, a multi-target tracking and detection algorithm based on the BYTE strategy is used to identify the vehicles' original motion trajectories.
[0019] S3: Use the top-down perspective boundary vehicle trajectory inference algorithm to obtain the translational features of the vehicle detection frame boundary and the vehicle's in-plane rotational features when the vehicle is completely within the field of view. This algorithm infers and improves the missing information when the vehicle is partially outside the field of view boundary, thereby correcting the original motion trajectory in the image coordinate system and obtaining the motion trajectory when the vehicle is partially outside the field of view boundary.
[0020] S4: Correlate the vehicle load information obtained from the side view with the corrected motion trajectory of the corresponding vehicle from the top view to obtain the spatiotemporal distribution of the traffic load.
[0021] Furthermore, in step S2, the original motion trajectory of the vehicle is identified by a multi-target tracking detection algorithm based on the BYTE strategy, specifically including:
[0022] S201: After identifying each frame of the video captured from the bird's-eye view, a vehicle detection frame and its detection score are obtained in each frame, and the detection score is compared with two thresholds η. high and η low Comparison, 0<η low <η high <1, thus grouping the detection frames into a high-resolution group, a low-resolution group, and a background group, and directly discarding the detection frames of the background group;
[0023] The specific grouping strategy is:
[0024] When the detection score η corresponding to the j-th detection frame in the i-th frame i,j Satisfy η high ≤η i,j <1, the detection frame corresponding to the detection score is divided into the high group G high,i ;
[0025] When η i,j Satisfy η low ≤η i,j <η high When , the detection frame corresponding to the detection score is divided into the low group G low,i ;
[0026] When η i,j Satisfying 0<η i,j <η low , then the detection frame corresponding to the detection score is divided into the background group G back,i ;
[0027] S202: Set the vehicle detection frame obtained in the first frame as the initial trajectory of the vehicle. From the second frame onwards, use the Kalman filter to predict the new position of each vehicle trajectory in the current frame. The predicted trajectory set corresponding to the i-th frame is recorded as Γ i The initial trajectory and predicted trajectory both refer to the vehicle rectangular detection frame containing length, width and center coordinate information;
[0028] S203: Group high-frequency high,i The set of all detection boxes and trajectories in Γ i Perform the first round of matching, and the unmatched high-scoring detection frames in this step are stored in the set G remain,i In the , unmatched trajectories are stored in the set Γ remain,i middle;
[0029] S204: Group low-level low,i The set of all detection boxes and trajectories in Γ remain,i Perform the second round of matching. The unmatched detection frames in this step are regarded as background and no longer used. The unmatched trajectories are stored in the set Γ lost,i middle;
[0030] S205: Output the successfully matched detection box ID and trajectory ID as the matching result of the i-th frame and the i-1 frame; use the method described in S202 to obtain the predicted trajectory set Γ corresponding to the i+1 frame i+1 ; Set the regenerated video frame number threshold n lost , the i+1-nth lost The set of all unmatched trajectories from the frame image to the i-th frame Join Γ i+1 ; G of the i-th frame described in step S203 remain,i Each detection box in the set is added to Γ i+1 , regarded as the initialization of the new trajectory;
[0031] S206: Repeat S201-S205 to complete the matching between all adjacent frames, and obtain the detection frame and its center position of each frame as the original motion trajectory of the vehicle.
[0032] Furthermore, the step S3 specifically includes the following sub-steps:
[0033] S301: Extract the critical start time t of the vehicle from the top-view video information and the original motion trajectory of the vehicle. c and critical end time t e , including the critical start time t when the vehicle enters the picture c,in and critical end time t e,in , and the critical starting time t of the vehicle leaving the screen c,out and critical end time t e,out;
[0034] For the case where a vehicle enters the picture, the critical starting time t c,in is the first frame of the trajectory, the critical end time t e,in The vehicle detection frame's approach direction frame line leaves the field of view boundary and starts to move longitudinally;
[0035] For the case where the vehicle leaves the screen, the critical starting time t c,out The critical end time t is when the vehicle detection frame reaches the boundary of the visual field and stops moving longitudinally. e,out is the last frame of the trajectory;
[0036] S302: At t e,in to t c,out During the period between, the vehicle detection frame includes the complete vehicle body, and the vehicle length and width are obtained based on the vehicle geometric feature contour homogenization recognition algorithm;
[0037] S303: When the vehicle partially enters or exits the visual field boundary, estimate the longitudinal length w of the vehicle virtual detection frame v and the acute angle θ formed between the vehicle's axis of symmetry and the direction in which the road extends;
[0038] S304: Calculating the length l of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary v ;
[0039] S305: Correcting the initial trajectory of the vehicle when it partially enters or exits the field of view boundary;
[0040] S306: Convert the corrected trajectory to the road coordinate system using a homography transformation method to obtain the actual trajectory of the vehicle.
[0041] Furthermore, in S302, the vehicle length and width are obtained according to the vehicle geometric feature contour homogenization recognition algorithm, which specifically includes the following sub-steps:
[0042] S3021: When the monitoring frame includes the complete vehicle and the vehicle is within the set range in the center of the image, the image within the detection frame is cropped and converted to a grayscale image. Canny edge detection is performed on the grayscale image, followed by three image processing operations: morphological dilation, hole filling, and morphological erosion. This eliminates various fine dots and lines in the background area and holes in the foreground area, so that the foreground is merged into one or more continuous and complete areas.
[0043] S3022: If the detection frame in S3021 is converted into multiple foreground regions, the largest foreground region U is selected. max And find the center of mass P c ; If only one foreground area is converted, then its centroid is directly calculated; from the centroid P cDeparture at U max Find a point P on each of the two long sides of the boundary j , j = 1, 2, so that P c To P j The Euclidean distance is the shortest, and the shortest distances are d min,j ; In P j Search for a series of pixels around each point so that they all meet P c To P j The Euclidean distance is (1+ζ d )d min The condition where ζ d is the search magnification factor, which is 5%-10%; j The multiple pixel points obtained by searching around are defined as the candidate perpendicular point set Ψ j ;
[0044] S3023: Calculate the uniformity index of the candidate perpendicular foot point set to further determine the perpendicular foot position:
[0045] (1) Ψ j All pixel points in the image are sorted by ascending horizontal coordinates as the first criterion and ascending vertical coordinates as the second criterion. If multiple points with adjacent labels have the same horizontal coordinates, the vertical coordinates are averaged and the average result is used as a new point to replace the previous points with the same horizontal coordinates. The following preprocessed candidate perpendicular foot point set is obtained:
[0046]
[0047] Among them, m j for The number of rows in the point set;
[0048] (2) Calculation The slope vector S j , S j The kth element in is:
[0049]
[0050] Where k = 2, 3…, m j ;S j,1 =0;
[0051] (3) Statistical slope vector S j The number of positive and negative elements in the array is n. +j , the number of negative elements is n -j , calculate the candidate perpendicular point set Ψ j The uniformity index I ev,j as follows:
[0052]
[0053] Among them, ∑m j is the sum of the number of points in all candidate perpendicular point sets after preprocessing, ζ n is the point set density weight coefficient;
[0054] S3024: Select I ev The largest set of candidate perpendicular points is used to calculate the average coordinates of all points as the identification perpendicular point P. ft ; Connect the centroid P c With P ft , and make a line through the centroid and perpendicular to P c P ft The straight line is used as the longitudinal symmetry axis of the vehicle; the Euclidean distance between any two intersection points of the longitudinal symmetry axis and all foreground areas in the vehicle detection frame is calculated; the maximum value of the Euclidean distance is selected as the vehicle recognition length l, P c P ft The twice of the Euclidean distance is the vehicle recognition width w.
[0055] Furthermore, the S303 specifically includes:
[0056] For the case where the vehicle partially drives out of the boundary, a quadratic polynomial is used to calculate t e,in to t c,out The horizontal coordinate x of the exit direction frame line bl Fitting to x out =f1(t), and x out = f1(t) extrapolated to t c,out to t e,out t c,out to t e,out The vertical length of the virtual detection frame w v (t)=|f1(t)-x b,out |, where x b,out is the horizontal coordinate of the boundary of the visual field in the exit direction;
[0057] For the case where the vehicle partially enters the boundary, a quadratic polynomial is used to calculate t e,in to t c,out The horizontal coordinate of the driving direction frame between the two is fitted as x in =f2(t), and x in = f2(t) extrapolated to t c,out to t e,out t c,in to t e,in The vertical length w of the virtual detection frame between v (t)=|f2(t)-x b,in |;
[0058] For the case where the vehicle is completely within the detection frame, a quadratic polynomial is used to convert the height of the detection frame h into a series of changes when the vehicle is completely within the picture. t Fitting to h t =f3(t), and extrapolate to t c,in to t e,out The acute angle θ between the vehicle's axis of symmetry and the road extension direction is calculated using the following formula:
[0059]
[0060] in, w and l are the width and length of the complete vehicle respectively.
[0061] Furthermore, the length of the vehicle of the invisible portion when the vehicle partially enters or exits the visual field boundary in S304 is l v The calculation formula is as follows:
[0062]
[0063] in, is the angle between the vehicle's driving direction and the positive direction of the horizontal axis of the image coordinate system, The corresponding relationship with θ is as follows: If the vehicle is traveling in the first quadrant, then If it points to the second quadrant, If it points to the third quadrant, If it points to the fourth quadrant, then w t (t) is the longitudinal length of the detection frame of the visible part of the vehicle that has not gone out of the boundary at time t.
[0064] Furthermore, the S305 specifically includes:
[0065] When the vehicle partially drives out of the boundary, at t c,out ≤t≤t e,out Time periods include:
[0066] When the vehicle partially enters the boundary, at t c,in ≤t≤t e,in Time periods include:
[0067] Among them, x box (t), y box (t) is the center coordinate of the detection frame obtained by S2, that is, the original motion trajectory; and Correct the trajectory for the boundary.
[0068] Furthermore, the S4 is specifically:
[0069] In the image information obtained from the side view, vehicles are assigned IDs in the order in which they appear in each lane;
[0070] In the top-down image information, vehicles are assigned IDs in the order in which their trajectories appear within the field of view, and the top-down vehicle IDs are sorted by lane.
[0071] The vehicles in a certain lane in the top-down view are sequentially associated with the vehicles in the corresponding lane in the side view.
[0072] The present invention has the following beneficial effects:
[0073] (1) The present invention can achieve uniform recognition of the geometric features of vehicles from a bird's-eye view, obtain the geometric characteristics of vehicles in any driving posture on the road, and improve the utilization of bird's-eye view traffic video data. The resulting recognition data can be used for subsequent trajectory correction and as an information source for detecting and warning special traffic conditions (such as width and length restrictions).
[0074] (2) The vehicle trajectory inference algorithm based on the top-down perspective boundary in the present invention infers and improves the missing information when part of the vehicle is outside the field of view boundary, eliminating the original motion trajectory recognition anomaly caused by the detection frame being unable to include the entire vehicle body when the vehicle reaches the field of view boundary. In principle, it overcomes the defects of the traditional method of trajectory recognition based on the center position of the detection frame.
[0075] (3) The present invention associates the side-view vehicle load identification results with the top-view motion trajectory of the variable lane area, integrating the vehicle spatial position information with attribute information such as load, and obtains the spatiotemporal distribution pattern of traffic load on the road section of interest only through a non-contact system, providing a reliable, durable and low-cost solution for traffic control, traffic load data accumulation, and service status assessment of road and bridge infrastructure. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 Schematic diagram of the multi-perspective collaborative perception device for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning provided in this embodiment.
[0077] Figure 2 A schematic diagram of the process and principle of a multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning provided in an embodiment.
[0078] Figure 3 The flowchart and principle diagram for vehicle length and width recognition from a top-down perspective provided in this embodiment, where (a) is the recognition flowchart, (b) is the grayscale image, (c) is the detected edge, (d) is the result after morphological dilation, (e) is the result after the hole filling operation, and (f) is the recognition result after morphological erosion.
[0079] Figure 4 Figure 3 illustrates the effect of different vehicle trajectory recognition and original trajectory reasoning and correction in an embodiment of the present invention. (a) is a flowchart of the vehicle trajectory reasoning algorithm for the top-down perspective boundary, (b) is a schematic diagram of the geometric parameters of the vehicle exiting the boundary, and (c) is a schematic diagram of the virtual detection box width inference principle.
[0080] Figure 5 Schematic diagrams of the original and corrected trajectories of vehicles No. 115 and No. 15 in the embodiment, wherein (a) is the original trajectory of vehicle No. 115, (b) is the original trajectory of vehicle No. 15, (c) is the correction result of the trajectory of vehicle No. 115 using the method of the present invention, and (d) is the correction result of the trajectory of vehicle No. 15 using the method of the present invention. DETAILED DESCRIPTION
[0081] The present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments, and the purpose and effects of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0082] like Figure 1 As shown, the boundary reasoning-enhanced multi-perspective collaborative perception device for spatiotemporal distribution patterns of highway traffic loads of the present invention includes: a side-perspective video acquisition device, a top-perspective video acquisition device, a vehicle-road collaborative perception device, and a data processing device.
[0083] The side-view video acquisition device is installed on the side of each lane of the immutable road section to identify the vehicle load sequence of the corresponding lane. The camera installation of the side-view video acquisition device must meet the following conditions:
[0084] (1) The camera's image plane is perpendicular to the road surface;
[0085] (2) The main optical axis of the camera is perpendicular to the direction of the road;
[0086] (3) The camera should be installed at a height that allows the complete capture of tire images when most conventional vehicles pass through its field of view.
[0087] Because side-view video capture equipment requires a certain amount of space across the highway, it's recommended to be deployed in locations such as toll booths and checkpoints. This also ensures that vehicles within the side-view detection area cannot change lanes at will. Similarly, vehicle-road cooperative sensing equipment can be installed near the camera to obtain actual tire pressure data from the tire pressure monitoring system (TPMS).
[0088] The overhead video capture device located above the road section of interest serves as an overhead information source to obtain the vehicle's driving trajectory. The installation of the overhead video capture device should meet the following conditions:
[0089] (1) The principal optical axis of the lens is perpendicular to the road surface, and the horizontal axis of the image coordinate axis is parallel to the longitudinal axis of the road;
[0090] (2) The visual zone is more than twice the length of most conventional vehicles in the longitudinal direction of the road, and covers all lanes in the same direction on one side of the road in the transverse direction;
[0091] (3) The longitudinal position of the top-view video acquisition device is installed immediately after the side view area covered by the side view video acquisition device, and the side where the vehicle enters within the visible range of the top-view video acquisition device includes a preset non-lane change area of a certain length, so as to ensure that the lane where the vehicle is located within the side view range can be obtained in the top-view video, which is convenient for subsequent data association.
[0092] The data processing equipment has a built-in non-contact vehicle dynamic weighing analysis system, which obtains the real-time inflation pressure of the tires of a moving vehicle and the vehicle tire image data obtained from the side view through the vehicle-road cooperative network to solve the vehicle's single wheel load and the vehicle load. It solves the vehicle's running trajectory by obtaining the vehicle video information from the top view, and correlates the vehicle load information obtained from the side view with the vehicle's spatial trajectory obtained from the top view to obtain the temporal and spatial distribution of the vehicle load.
[0093] like Figure 2 As shown, the boundary reasoning-enhanced multi-perspective collaborative perception method for spatiotemporal distribution pattern of highway traffic load in this embodiment is implemented based on the above-mentioned multi-perspective collaborative perception device for spatiotemporal distribution pattern of highway traffic load, and proposes a top-down vehicle geometric feature contour homogenization recognition algorithm, which realizes the use of monocular top-down image to extract the geometric characteristics of vehicles with arbitrary driving postures on the surface; proposes a top-down boundary vehicle trajectory reasoning algorithm, which infers and completes the missing information when the vehicle partially drives out of the field of view boundary, and solves the problem that traditional trajectory recognition methods are prone to abnormal results at the field of view boundary; relying only on non-contact perception methods, the vehicle load information obtained from the side perspective and the vehicle trajectory distribution obtained from the top perspective are associated and integrated, and the vehicle load distribution pattern of the entire field can be obtained. Specifically, it includes the following steps:
[0094] S1: The side view video acquisition device of each lane collects tire images of each vehicle passing through the lane, and calculates the wheel load and deadweight information of each vehicle based on the tire images.
[0095] In one implementation, side-view vehicle tire image data, along with actual tire pressure data from the tire pressure monitoring system (TPMS) obtained through a vehicle-road cooperative network built using vehicle-road cooperative sensing devices, is transmitted to a non-contact WIM analysis system to determine single-wheel and vehicle loads. This step aims to obtain a sequence of vehicle loads for each lane, so any system capable of distinguishing between different vehicles and obtaining their own weight can be substituted, including non-contact visual WIM systems and contact piezoelectric WIM systems.
[0096] A non-contact WIM system that meets these requirements processes images and TPMS signals as follows: Tire images are segmented into multiple regions of interest (ROIs). Key information is extracted from images of the tire-ground contact area, rim area, and tire sidewall text marking area using machine vision algorithms. The system then quantifies the visual characteristics of tire deformation. The system then obtains the standard tire inflation pressure and various cross-sectional geometric parameters, rim parameters, and tread parameters through direct recognition or standard database retrieval. Based on these parameters and actual inflation pressure data, a contact mechanics model (such as the tire pressure balance model, radial compression model, Hertzian contact model, or elliptical constraint Hertzian deformation quantification model) is used to correlate tire deformation characteristics with external loads, thereby determining the single-wheel load and total vehicle load.
[0097] S2: When one or more vehicles exit the side view non-lane change zone and enter the top view range, the top view video acquisition device starts continuous video acquisition and uses the pre-trained YoloV8n network to identify the vehicles in the picture frame by frame, obtaining the detection frame and the corresponding confidence score; then, the original motion trajectory of the vehicle is identified through the multi-target tracking detection algorithm based on the BYTE strategy.
[0098] The original motion trajectory of the vehicle is a series of detection frames and their center coordinates (x box ,y box ). The origin of the image coordinate system is O IM Located in the upper left corner of the screen, x IM The positive direction of the axis points to the right, y IM The positive direction of the axis points downward.
[0099] The steps of the multi-target tracking detection algorithm based on BYTE strategy are as follows:
[0100] S201: After identifying each frame of the video captured from the bird's-eye view, a vehicle detection frame and its detection score are obtained in each frame, and the detection score is compared with two thresholds η. high and η low Comparison, 0<η low <η high<1, thus grouping the detection frames into a high-resolution group, a low-resolution group, and a background group, and directly discarding the detection frames of the background group.
[0101] The specific grouping strategy is:
[0102] When the detection score η corresponding to the j-th detection frame in the i-th frame i,j Satisfy η high ≤η i,j <1, the detection frame corresponding to the detection score is divided into the high group G high,i ;
[0103] When η i,j Satisfy η low ≤η i,j <η high When , the detection frame corresponding to the detection score is divided into the low group G low,i ;
[0104] When η i,j Satisfying 0<η i,j <η low , then the detection frame corresponding to the detection score is divided into the background group G back,i .
[0105] Since the vertical viewing angle reduces the uncertainty of recognition, the detection frame with low confidence is likely to be a false detection of the background, so G back,i Abandonment is reasonable.
[0106] S202: Set the vehicle detection frame obtained in the first frame as the initial trajectory of the vehicle. From the second frame onwards, use the Kalman filter to predict the new position of each trajectory in the current frame. The set of predicted trajectories in the i-th frame is recorded as Γ i ; Both the initial trajectory and the predicted trajectory refer to the vehicle rectangular detection box containing the length, width and center coordinate information;
[0107] S203: Group high-frequency high,i The set of all detection boxes and trajectories in Γ i Perform the first round of matching, and the unmatched high-scoring detection frames in this step are stored in the set G remain,i In the , unmatched trajectories are stored in the set Γ remain,i middle;
[0108] S204: Group low-level low,i The set of all detection boxes and trajectories in Γ remain,i Perform the second round of matching. The unmatched detection frames in this step are regarded as background and no longer used. The unmatched trajectories are stored in the set Γ lost,i middle.
[0109] The two-round detection mechanism of steps S203 and S204 increases the chance of matching and has better applicability to situations where there may be accidental occlusion of the target, sudden changes in the size of the target detection frame, etc.
[0110] S205: Output the successfully matched detection box ID and trajectory ID as the matching result of the i-th frame and the i-1 frame; use the method described in S202 to obtain the predicted trajectory set Γ corresponding to the i+1 frame i+1 ; Set the regenerated video frame number threshold n lost , the i+1-nth lost The set of all unmatched trajectories from the frame image to the i-th frame Join Γ i+1 ; G of the i-th frame described in step S203 remain,i Each detection box in the set is added to Γ i+1 , regarded as the initialization of the new trajectory.
[0111] Because G remain,i Although the detection box in has a high detection score, it cannot match any of the aforementioned trajectories. Therefore, it is usually a vehicle that suddenly merges from both sides, and it is necessary to initialize it as a new trajectory.
[0112] S206: Repeat steps S201-S205 to complete the matching between all adjacent frames, and obtain the detection frame and its center position of each frame as the original motion trajectory of the vehicle.
[0113] In the above process, it is necessary to determine the specific solution for optimally matching multiple detection boxes with multiple predicted trajectories in each round. In practical applications, the intersection-over-union (IoU) similarity can be used as the matching basis, and the matching is completed using the classic Hungarian algorithm, as follows:
[0114] (1) Calculate the intersection over union (IoU) similarity of all detection boxes and predicted trajectories to be matched in the same frame to form a cost matrix. The IoU similarity calculation method between the jth detection box and the kth predicted trajectory in the i-th frame is:
[0115]
[0116] Where S(·) represents the area of the region of interest.
[0117] (2) Let the number of detection frames required to be matched in a certain round be n G , the number of predicted trajectories required to match is n Γ , if the two are not equal, assume n Γ If the value is the larger one, the IoU similarity calculation results can be summarized as n Γ ×n Γ The cost matrix is:
[0118]
[0119] That is, if n G ≠n Γ , the missing elements in the corresponding positions in the square matrix are filled with 0; if the two are equal, the square matrix is directly constructed.
[0120] (3) The matching of detection box and predicted trajectory is compared to the “assignment problem”, and the Hungarian algorithm is used to process the cost matrix C IoU,i , you can get the best matching solution.
[0121] S3: Use the top-down perspective boundary vehicle trajectory inference algorithm to obtain the translational characteristics of the vehicle detection frame boundary and the vehicle's in-plane rotational characteristics when the vehicle is completely within the field of view. Infer and improve the missing information when the vehicle is partially outside the field of view boundary, and then correct the original motion trajectory in the image coordinate system to obtain the motion trajectory of the vehicle when it is partially outside the field of view boundary.
[0122] S301: Extract the critical start time t of the vehicle from the top-view video information and the original motion trajectory of the vehicle. c and critical end time t e , including the critical start time t when the vehicle enters the picture c,in and critical end time t e,in , and the critical starting time t of the vehicle leaving the screen c,out and critical end time t e,out ;
[0123] For the case where a vehicle enters the picture, the critical starting time t c,in is the first frame of the trajectory, the critical end time t e,in The vehicle detection frame's approach direction frame line leaves the field of view boundary and starts to move longitudinally;
[0124] For the case where the vehicle leaves the screen, the critical starting time t c,out The critical end time t is when the vehicle detection frame reaches the boundary of the visual field and stops moving longitudinally. e,out is the last frame of the trajectory.
[0125] A feasible method for obtaining critical moments is given below.
[0126] (1) For the complete trajectory of a vehicle in the picture, obtain the horizontal coordinate x of the entry direction frame line of the detection frame corresponding to each point on it br and the horizontal coordinate x of the exit direction frame bl , get the horizontal coordinate x of the driving direction frame br Sequence and exit direction frame horizontal coordinate x bl sequence.
[0127] For the sake of convenience, it is assumed that the vehicle enters the screen from right to left, so the frame line in the direction of entry is the "right frame line" and the frame line in the direction of exit is the "left frame line". e ) and the detection box width (w b ) Calculate the horizontal coordinate x of the driving direction frame line br and the horizontal coordinate x of the exit direction frame bl The method is as follows:
[0128] x bl =x e -0.5w b
[0129] x br =x e +0.5w b
[0130] So far, we have obtained the horizontal coordinate x of the entry direction frame line at equal time intervals. br Sequence and exit direction frame horizontal coordinate x bl sequence.
[0131] (2) Calculate the horizontal coordinate x of the entry direction frame line br Sequence and exit direction frame horizontal coordinate x bl The first-order difference of the sequence, Δx bl Sequence and Δx br Sequence, and according to the length of the sequence, select the appropriate period to perform moving average on the first-order difference sequence to enhance the smoothness of its non-mutation segment, and obtain the sequence Δx bl and Δx br ;
[0132] (3) Continue to calculate Δx after moving average processing bl Sequence and Δx br The difference of the sequence, that is, the second-order difference Δ 2 x bl Sequence and Δ 2 x br Sequence, find the moment when the absolute value of the second-order difference sequence is the largest, |Δ 2 x br The moment corresponding to the maximum value is the critical end moment t when the vehicle enters the screen e,in ,|Δ 2 x bl The moment corresponding to the maximum value is the critical starting moment t when the vehicle leaves the screen c,out .
[0133] Here x br Taking the sequence as an example (assuming that the initial number of elements is n), the specific calculation methods of (2) and (3) are given:
[0134] The specific method of the first-order difference mentioned above is: Δx bl,i =x bl,i+1 -x bl,i (i=2,3…n-1);
[0135] Assuming that the number of periods used in the above moving average method is N, then The reason for nN here is that if the moving average uses N periods, the sequence length will be shortened by N-1.
[0136] The specific method of the second-order difference mentioned above is:
[0137] S302: At t e,in to t c,out During the period between , the vehicle detection frame includes the complete vehicle body, and the vehicle length and width are obtained based on the vehicle geometric feature contour homogenization recognition algorithm, such as Figure 3 As shown, it specifically includes the following sub-steps:
[0138] S3021: When the monitoring frame includes the complete vehicle and the vehicle is within the set range in the center of the image, the image within the detection frame is cropped and converted to a grayscale image. Canny edge detection is performed on the grayscale image, followed by three image processing operations: morphological dilation, hole filling, and morphological erosion. This eliminates various fine dots and lines in the background area and holes in the foreground area, so that the foreground is merged into one or more continuous and complete areas.
[0139] S3022: If the detection frame in S3021 is converted into multiple foreground regions, the largest foreground region U is selected. max And find the center of mass P c If only a foreground area is converted, then its center of mass is directly calculated (from the vehicle structure characteristics, U max Approximately a rectangle, with the long sides of both sides parallel to the direction of vehicle travel); from the center of mass P c Departure at U max Find a point P on each of the two long sides of the boundary j , j = 1, 2, so that P c To P j The Euclidean distance is the shortest, and the shortest distances are d min,j ; In P j Search for a series of pixels around each point so that they all meet P c To P j The Euclidean distance is (1+ζ d )d min The condition where ζ d is the search magnification factor, which can be 5%-10%. j The multiple pixel points obtained by searching around are defined as the candidate perpendicular point set Ψj .
[0140] S3023: Calculate the uniformity index of the candidate perpendicular foot point set to further determine the perpendicular foot position. The steps are as follows: j All pixel points in the image are sorted by ascending horizontal coordinates as the first criterion and ascending vertical coordinates as the second criterion. If multiple points with adjacent labels have the same horizontal coordinates, the vertical coordinates are averaged and the average result is used as a new point to replace the previous points with the same horizontal coordinates. The following preprocessed candidate perpendicular foot point set is obtained:
[0141]
[0142] At this time there must be Not equal, m j for The number of rows in the point set.
[0143] calculate The slope vector S j . S j The kth element in (k=2,3…,m j )for:
[0144]
[0145] In particular, another S j,1 = 0. If There is only one point in the same way, so let S j,1 =0;
[0146] Statistical slope vector S j The number of positive and negative elements in the array is n. +j , the number of negative elements is n -j Finally, the candidate perpendicular point set Ψ is calculated j The uniformity index is as follows:
[0147]
[0148] where ∑m j is the sum of the number of points in all candidate perpendicular point sets after preprocessing, ζ n is the point set density weight coefficient. The first term of the expression represents the monotonicity and uniformity of the boundary change near the candidate perpendicular foot, reflecting the degree of interference in boundary identification. The second term reflects the proportion of the number of candidate perpendicular feet contained in the point set.
[0149] S3024: Select I ev The largest set of candidate perpendicular points is used to calculate the average coordinates of all points as the identification perpendicular point P. ft ; Connect the centroid P c With P ft, and make a line through the centroid and perpendicular to P c P ft The straight line is used as the longitudinal symmetry axis of the vehicle. The longitudinal symmetry axis may have 2 or more intersections with all foreground areas in the vehicle detection frame. The Euclidean distance between any two intersections is calculated, and the maximum Euclidean distance is selected as the vehicle recognition length l; the vehicle recognition width w is P c P ft Twice the Euclidean distance.
[0150] In actual operation, you can e,in to t c,out Randomly capture 3-5 frames for analysis and take the average value of the length and width calculation results.
[0151] S303: When the vehicle partially enters or exits the field of view boundary v Make inferences with θ;
[0152] When a vehicle partially enters or exits the visual field boundary, first obtain the longitudinal length of the detection frame w t , horizontal height is h t Assume that the part of the vehicle body outside the field of view is also inscribed in a rectangular virtual detection frame. The virtual detection frame is the circumscribed rectangle of the invisible part of the vehicle body, and the two sides of the rectangle are parallel to the two coordinate axes of the top-view image coordinate system. Calculate the longitudinal length of the vehicle virtual detection frame as w v , horizontal height is h v , and the acute angle θ formed by the vehicle's symmetry axis and the road extension direction (the horizontal axis of the image coordinate system) when the vehicle partially enters or exits.
[0153] like Figure 4 As shown, when the vehicle partially enters or exits the field of view, w v The inference is made with θ, which is expressed as follows:
[0154] According to the horizontal coordinate x of the frame line in the direction of entry br and the horizontal coordinate x of the exit direction frame bl Regarding the time variation curve, we can know that the horizontal coordinate x of the exit direction frame line bl In t c,in to t c,out The segment changes continuously, and this segment is fitted with a quadratic polynomial as x out =f1(t); horizontal coordinate t of the exit direction frame c,out to t e,out The horizontal line is a horizontal line, and its abscissa value is the abscissa x of the boundary of the field of view in the direction of driving out. b,out . The fitting result x out = f1(t) extrapolated to t c,out to t e,out When t c,out ≤t≤te,out , the vertical length of the virtual detection frame can be calculated as: w v (t)=|f1(t)-x b,out |;
[0155] Similarly, we can know that the horizontal coordinate of the driving direction frame line is at t e,in to t e,out This segment is fitted with a quadratic polynomial as x in =f2(t); horizontal coordinate t of the driving direction frame line c,in to t e,in The horizontal line is a horizontal line, and its abscissa value is the abscissa x of the visual field boundary in the driving direction. b,in . The fitting result x in = f2(t) extrapolated to t c,in to t e,in When t c,in ≤t≤t e,in , the vertical length of the virtual detection frame can be calculated as: w v (t)=|f2(t)-x b,in |;
[0156] When the vehicle is completely within the frame (ie t e,in ≤t≤t c,out ), a series of changing detection frame heights h can be obtained t , which is fitted with a quadratic polynomial to h t =f3(t), and extrapolated to t c,in to t e,out Time period. Let the acute angle between the vehicle's symmetry axis and the road extension direction (the horizontal axis of the image coordinate system) be θ, then:
[0157]
[0158] in w and l are the vehicle recognition length and recognition width obtained in step 3 respectively.
[0159] S304: Calculating the length l of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary v as follows:
[0160]
[0161] in, is the angle between the vehicle's driving direction and the positive direction of the horizontal axis of the image coordinate system, The corresponding relationship with θ is as follows: If the vehicle is traveling in the first quadrant, then If it points to the second quadrant, If it points to the third quadrant, If it points to the fourth quadrant,
[0162] S305: Correct the initial trajectory of the vehicle when it partially enters or exits the field of view using the following formula:
[0163] When the vehicle partially drives out of the boundary, at t c,out ≤t≤t e,out Time periods include:
[0164] Similarly, when the vehicle partially enters the boundary, at t c,in ≤t≤t e,in Time periods include:
[0165] Among them, x box (t), y box (t) is the center coordinate of the detection frame obtained in step 2, i.e., the initial trajectory; and Correcting trajectories for boundaries;
[0166] S306: Since the boundary correction trajectory is in the image coordinate system x IM O IM y IM Therefore, it is necessary to transform the corrected trajectory, that is, the coordinate position of the vehicle at different times, from the image coordinate system to the road coordinate system through the homography transformation method to obtain the real trajectory of the vehicle.
[0167] Among them, the homography transformation method is as follows:
[0168] Assume that the coordinates of a point on the vehicle trajectory after boundary correction in the image coordinate system are (x IM ,y IM ), the coordinates in the road coordinate system are (x RD ,y RD ), then the two have the following relationship:
[0169]
[0170] The above relationship can be summarized as: RD =HP IM , H is the homography matrix. Using (H 31 x IM +H 32 y IM +1) and multiply the left side of the first two rows of the above formula, and shift the terms to get:
[0171] H 11 x IM +H 12 y IM +H 13 -H 31 x IMx RD -H 32 y IM x RD =x RD
[0172] H 21 x IM +H 22 y IM +H 23 -H 31 x IM y RD -H 32 y IM y RD =y RD
[0173] Therefore, the eight unknowns in the homography matrix can be calculated using four marker points whose coordinates are known, within the visible range of the top-down perspective, and not collinear. At this point, eight independent equations can solve all the unknown elements.
[0174] After obtaining the actual trajectory of the vehicle, the vehicle speed can be calculated through a series of point coordinates in the actual trajectory and its time frame. The origin of the road coordinate system can be determined according to the on-site features, and the coordinate axis x r Parallel to the road extension direction, coordinate axis y r Perpendicular to the direction of road extension.
[0175] S4: The side-view vehicle load identification results in the non-laneable area are associated with the following top-view motion trajectory according to the principle of sequential matching. By integrating the vehicle trajectory with attribute information such as load and axle load, the spatiotemporal distribution of traffic load is obtained.
[0176] All vehicles in the top-down view are sorted according to the lanes they enter. That is, the unified global vehicle ID in the top-down view is reconstructed according to the order of vehicles entering each lane, and then associated with the vehicles identified in the side view.
[0177] The specific column order matching principles are:
[0178] (1) ID numbering rule: In the image information obtained from the side view, the vehicles are assigned IDs in the order in which they appear in each lane, for example, ID side =(i,j) represents the jth vehicle that appears in lane i; in the top-down image information, the vehicles are assigned IDs in the order in which their trajectories appear in the field of view, for example, ID vert =k represents the kth vehicle that appears within the visible range of the top-down perspective;
[0179] (2) Lane determination rule: Since there is a certain area of non-laneable area (solid road line) in the direction of vehicle approach from the top-down perspective, if the first three frames of the vehicle's actual trajectory obtained in step 4 are all located in a certain lane, the vehicle is considered to come from that lane;
[0180] (3) Association criteria: Based on rule (2), the top-view vehicle IDs can be sorted by lane. vert = k is determined to be from lane m and its position in the lane is n, then it is compared with ID side =(m,n) vehicle association.
[0181] If there is a side view and several subsequent top-down views, the vehicle trajectories between adjacent top-down views are matched in the following manner: (1) Ensure that the field of view of each top-down view overlaps with that of the next top-down view, and that the time frames of the top-down view videos are recorded using the same reference, and that each top-down view uses a unified road coordinate system; (2) Obtain the boundary-corrected vehicle trajectories in each top-down view; (3) If the two vehicle trajectories obtained from adjacent top-down views match within the overlapping field of view, or if the Euclidean distance between the detection frame centers calculated frame by frame is less than a certain threshold, the two trajectories are automatically matched.
[0182] To demonstrate the effectiveness of the method, this example uses a DJI Phantom 4 RTK drone as a bird's-eye view video capture device. Its camera is a 1-inch CMOS sensor with a 3840×2160 resolution and 30 frames per second. The video was captured at a highway toll station in Zhejiang Province, with the drone hovering vertically above the lane-changing area behind the toll station.
[0183] The bird's-eye view area set in this embodiment includes the first and second lanes and half of the third lane of the road section after the toll station, where half of the third lane is used to increase interference and recognition difficulty. The above-mentioned first lane is located on the roadside, and the third lane is close to the center line of the road. When performing the test, there was an engineering vehicle parked in the first lane about 60m away from the toll station. Although it was outside the bird's-eye view range, it caused a large number of vehicles entering from the first lane to change lanes, increasing the richness of the trajectory. The drone's field of view is slightly larger than the pre-set bird's-eye view test area, which avoids incompleteness of the test area due to wind-induced vibration.
[0184] The machine vision program of this embodiment is based on the pre-trained YoloV8n network. The training method is: first pre-training using the MS-COCO large image dataset, and then migrating to the 1230 overhead vehicle database established in the early research of this embodiment for fine-tuning. All images in the overhead vehicle database are manually labeled using the LabelImg tool. The data processing equipment (industrial control computer) equipped with the neural network and algorithm uses NVIDIA TMThe GeForce RTX 4060 GPU is used as the computing core. After testing, the multi-target tracking and detection algorithm based on the BYTE strategy of the present invention takes an average of 11.7ms per frame, which is significantly faster than the video frame rate, indicating that the software and hardware system has the ability of real-time tracking.
[0185] like Figure 5 As shown, Figure 5 Figures (a) and (b) show the detection frames and initial motion trajectories of vehicles 115 and 15, respectively, identified during the test period. Vehicle 115 is a light-duty flatbed truck, while vehicle 15 is a heavy-duty flatbed truck, the latter significantly longer than the former. The trajectory curves show that the longitudinal length of the detection frame decreases as it approaches the edge of the field of view, resulting in unusual bends in the trajectory in this area, with the anomaly becoming more pronounced with longer vehicles.
[0186] Figure 5 Figures (c) and (d) show the trajectories of vehicles 115 and 15, respectively, corrected using the proposed top-view boundary vehicle trajectory inference algorithm. Compared to the original trajectories, the unusual bends near the boundary of the field of view have been corrected. Furthermore, the original trajectories have a denser distribution of points at the boundary of the field of view, indicating a decrease in velocity. The corrected trajectories exhibit a constant velocity or acceleration, more accurately reflecting the vehicle's actual driving conditions.
[0187] Those skilled in the art will understand that the foregoing descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art will still be able to modify the technical solutions described in the foregoing examples or substitute equivalents for some of the technical features therein. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the invention shall be included within the scope of protection of the invention.
Claims
1. A multi-perspective collaborative perception device for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning, characterized in that: Side-view video acquisition equipment, overhead video acquisition equipment, and data processing equipment; The side-view video acquisition device is installed on the side of each lane of the immutable road section to identify the vehicle load sequence of the corresponding lane; the camera installation of the side-view video acquisition device must meet the following conditions: (1) The camera's image plane is perpendicular to the road surface; (2) The main optical axis of the camera is perpendicular to the direction of the road; (3) The camera should be installed at a height that allows the complete capture of tire images when a vehicle passes through its field of view. The overhead video capture device is installed along the road extension direction and immediately after the immutable road segment, serving as an overhead information source to obtain the vehicle's driving trajectory. The installation of the overhead video capture device meets the following conditions: (1) The principal optical axis of the lens is perpendicular to the road surface, and the horizontal axis of the image coordinate axis is parallel to the longitudinal axis of the road; (2) The size of the visible area is more than twice the full length of a conventional vehicle in the longitudinal direction of the road, and covers all lanes in the same direction on one side of the road in the transverse direction; (3) The vertical position of the top-view video capture device is immediately behind the side view area covered by the side view video capture device, and the vehicle within the visual range of the top-view video capture device enters a non-laneable area of a preset length on one side, so as to ensure that the lane where the vehicle is located within the side view range can be obtained in the top-view video, which is convenient for subsequent data association; The data processing device has a built-in non-contact vehicle dynamic weighing analysis system, which obtains the real-time inflation pressure of the tires of a moving vehicle and the vehicle tire image data obtained from the side view through the vehicle-road cooperative network to solve the vehicle's single wheel load and the vehicle load. The vehicle's running trajectory is solved by obtaining the vehicle video information from the top view, and the vehicle load information obtained from the side view is correlated with the vehicle's spatial trajectory obtained from the top view to obtain the vehicle's spatiotemporal distribution of the load. The method of obtaining the vehicle's running trajectory by using the obtained vehicle video information from a bird's-eye view includes: Using the top-down perspective boundary vehicle trajectory inference algorithm, the translational characteristics of the vehicle detection frame boundary and the vehicle's in-plane rotation characteristics are obtained when the vehicle is completely within the field of view. The missing information when the vehicle is partially outside the field of view is inferred and improved, and the original motion trajectory in the image coordinate system is corrected to obtain the motion trajectory of the vehicle when it is partially outside the field of view.
2. The multi-perspective collaborative perception device for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 1 is characterized in that: The bird's-eye view video acquisition device has multiple bird's-eye view angles along the longitudinal direction of the road, and adjacent bird's-eye view angles satisfy the following conditions: the field of view of each bird's-eye view angle overlaps with the next bird's-eye view angle to a certain extent, and each bird's-eye view video time frame is recorded using the same reference, and each bird's-eye view angle uses a unified road surface coordinate system.
3. A multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning, characterized by: The method is implemented based on the boundary reasoning-enhanced multi-perspective collaborative perception device for highway traffic load distribution patterns according to claim 1, comprising the following steps: S1: The side view video capture device of each lane collects tire images of each vehicle passing through the lane, and calculates the single wheel load of each vehicle based on the tire images, thereby obtaining the total load of the vehicle; S2: When several vehicles exit the side view's non-lane-changeable zone and enter the top view's visible range, the top view video capture device begins continuous video capture and uses a pre-trained convolutional neural network to identify the vehicles frame by frame, obtaining detection frames and corresponding confidence scores. Subsequently, a multi-target tracking and detection algorithm based on the BYTE strategy is used to identify the vehicles' original motion trajectories. S3: Use the top-down perspective boundary vehicle trajectory inference algorithm to obtain the translational features of the vehicle detection frame boundary and the vehicle's in-plane rotational features when the vehicle is completely within the field of view. This algorithm infers and improves the missing information when the vehicle is partially outside the field of view boundary, thereby correcting the original motion trajectory in the image coordinate system and obtaining the motion trajectory when the vehicle is partially outside the field of view boundary. S4: Correlate the vehicle load information obtained from the side view with the corrected motion trajectory of the corresponding vehicle from the top view to obtain the spatiotemporal distribution of the traffic load.
4. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 3 is characterized in that: In step S2, the original motion trajectory of the vehicle is identified by a multi-target tracking detection algorithm based on the BYTE strategy, which specifically includes: S201: After identifying the video captured from the bird's-eye view frame by frame, a vehicle detection frame and its detection score are obtained in each frame, and the detection score is compared with two thresholds η. high and η low Comparison, 0<η low <η high <1, thus grouping the detection frames into a high-resolution group, a low-resolution group, and a background group, and directly discarding the detection frames of the background group; The specific grouping strategy is: When the detection score η corresponding to the j-th detection frame in the i-th frame i,j Satisfy η high ≤η i,j <1, the detection frame corresponding to the detection score is divided into the high group G high,i ; When η i,j Satisfy η low ≤η i,j <η high When , the detection frame corresponding to the detection score is divided into the low group G low,i ; When η i,j Satisfying 0<η i,j <η low , then the detection frame corresponding to the detection score is divided into the background group G back,i ; S202: Set the vehicle detection frame obtained in the first frame as the initial trajectory of the vehicle. From the second frame onwards, use the Kalman filter to predict the new position of each vehicle trajectory in the current frame. The predicted trajectory set corresponding to the i-th frame is recorded as Γ i The initial trajectory and predicted trajectory both refer to the vehicle rectangular detection frame containing length, width and center coordinate information; S203: Group high-frequency high,i The set of all detection boxes and trajectories in Γ i Perform the first round of matching, and the unmatched high-scoring detection frames in this step are stored in the set G remain,i In the , unmatched trajectories are stored in the set Γ remain,i middle; S204: Group low-level low,i The set of all detection boxes and trajectories in Γ remain,i Perform the second round of matching. The unmatched detection frames in this step are regarded as background and no longer used. The unmatched trajectories are stored in the set Γ lost,i middle; S205: Output the successfully matched detection box ID and trajectory ID as the matching result of the i-th frame and the i-1 frame; use the method described in S202 to obtain the predicted trajectory set Γ corresponding to the i+1 frame i+1 ; Set the regenerated video frame number threshold n lost , the i+1-nth lost The set of all unmatched trajectories from the frame image to the i-th frame Join Γ i+1 ; G of the i-th frame described in step S203 remain,i Each detection box in the set is added to Γ i+1 , regarded as the initialization of the new trajectory; S206: Repeat S201-S205 to complete the matching between all adjacent frames, and obtain the detection frame and its center position of each frame as the original motion trajectory of the vehicle.
5. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 4 is characterized in that: The step S3 specifically includes the following sub-steps: S301: Extract the critical start time t of the vehicle from the top-view video information and the original motion trajectory of the vehicle. c and critical end time t e , including the critical start time t when the vehicle enters the picture c,in and critical end time t e,in , and the critical starting time t of the vehicle leaving the screen c,out and critical end time t e,out ; For the case where a vehicle enters the picture, the critical starting time t c,in is the first frame of the trajectory, the critical end time t e,in The vehicle detection frame's approach direction frame line leaves the field of view boundary and starts to move longitudinally; For the case where the vehicle leaves the screen, the critical starting time t c,out The critical end time t is when the vehicle detection frame reaches the boundary of the visual field and stops moving longitudinally. e,out is the last frame of the trajectory; S302: At t e,in to t c,out During the period between, the vehicle detection frame includes the complete vehicle body, and the vehicle length and width are obtained based on the vehicle geometric feature contour homogenization recognition algorithm; S303: When the vehicle partially enters or exits the visual field boundary, estimate the longitudinal length w of the vehicle virtual detection frame v and the acute angle θ formed between the vehicle's axis of symmetry and the direction in which the road extends; S304: Calculating the length l of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary v ; S305: Correcting the initial trajectory of the vehicle when it partially enters or exits the field of view boundary; S306: Convert the corrected trajectory to the road coordinate system using a homography transformation method to obtain the actual trajectory of the vehicle.
6. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 5 is characterized in that: In S302, the vehicle length and width are obtained based on the vehicle geometric feature contour homogenization recognition algorithm, which specifically includes the following sub-steps: S3021: When the monitoring frame includes the complete vehicle and the vehicle is within the set range in the center of the image, the image within the detection frame is cropped and converted to a grayscale image. Canny edge detection is performed on the grayscale image, followed by three image processing operations: morphological dilation, hole filling, and morphological erosion. This eliminates various fine dots and lines in the background area and holes in the foreground area, so that the foreground is merged into one or more continuous and complete areas. S3022: If the detection frame in S3021 is converted into multiple foreground regions, the largest foreground region U is selected. max And find the center of mass P c ; If only a foreground area is converted, its centroid is directly calculated; From the center of mass P c Departure at U max Find a point P on each of the two long sides of the boundary j , j = 1, 2, so that P c To P j The Euclidean distance is the shortest, and the shortest distances are d min,j ; In P j Search for a series of pixels around each point so that they all meet P c To P j The Euclidean distance is (1+ζ d )d min The condition where ζ d is the search magnification factor, which is 5%-10%; j The multiple pixel points obtained by searching around are defined as the candidate perpendicular point set Ψ j ; S3023: Calculate the uniformity index of the candidate perpendicular foot point set to further determine the perpendicular foot position: (1) Ψ j All pixel points in the image are sorted by ascending horizontal coordinates as the first criterion and ascending vertical coordinates as the second criterion. If multiple points with adjacent labels have the same horizontal coordinates, the vertical coordinates are averaged and the average result is used as a new point to replace the previous points with the same horizontal coordinates. The following preprocessed candidate perpendicular foot point set is obtained: Among them, m j for The number of rows in the point set; (2) Calculation The slope vector S j , S j The kth element in is: Where k = 2, 3…, m j ;S j,1 =0; (3) Statistical slope vector S j The number of positive and negative elements in the array is n. +j , the number of negative elements is n -j , calculate the candidate perpendicular point set Ψ j The uniformity index I ev,j as follows: Among them, ∑m j is the sum of the number of points in all candidate perpendicular point sets after preprocessing, ζ n is the point set density weight coefficient; S3024: Select I ev The largest set of candidate perpendicular points is used to calculate the average coordinates of all points as the identification perpendicular point P. ft ; Connect the centroid P c With P ft , and make a line through the center of mass and perpendicular to P c P ft The straight line is used as the longitudinal symmetry axis of the vehicle; the Euclidean distance between any two intersection points of the longitudinal symmetry axis and all foreground areas in the vehicle detection frame is calculated; the maximum value of the Euclidean distance is selected as the vehicle recognition length l, P c P ft The twice of the Euclidean distance is the vehicle recognition width w.
7. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 6 is characterized in that: The S303 specifically includes: For the case where the vehicle partially drives out of the boundary, a quadratic polynomial is used to calculate t e,in to t c,out The horizontal coordinate x of the exit direction frame line bl Fitting to x out =f1(t), and x out = f1(t) extrapolated to t c,out to t e,out t c,out to t e,out The vertical length of the virtual detection frame w v (t)=|f1(t)-x b,out |, where x b,out is the horizontal coordinate of the boundary of the visual field in the exit direction; For the case where the vehicle partially enters the boundary, a quadratic polynomial is used to calculate t e,in to t c,out The horizontal coordinate of the driving direction frame between the two is fitted as x in =f2(t), and x in = f2(t) extrapolated to t c,out to t e,out t c,in to t e,in The vertical length w of the virtual detection frame between v (t)=|f2(t)-x b,in |; For the case where the vehicle is completely within the detection frame, a quadratic polynomial is used to convert the height of the detection frame h into a series of changes when the vehicle is completely within the picture. t Fitting to h t =f3(t), and extrapolate to t c,in to t e,out The acute angle θ between the vehicle's axis of symmetry and the road extension direction is calculated using the following formula: in, w and l are the width and length of the complete vehicle respectively.
8. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 7 is characterized in that: The length of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary in S304 is l v The calculation formula is as follows: in, is the angle between the vehicle's driving direction and the positive direction of the horizontal axis of the image coordinate system, The corresponding relationship with θ is as follows: If the vehicle is traveling in the first quadrant, then If it points to the second quadrant, If it points to the third quadrant, If it points to the fourth quadrant, then w t (t) is the longitudinal length of the detection frame of the visible part of the vehicle that has not gone out of the boundary at time t.
9. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 8 is characterized in that: The S305 specifically includes: When the vehicle partially drives out of the boundary, at t c,out ≤t≤t e,out Time periods include: When the vehicle partially enters the boundary, at t c,in ≤t≤t e,in Time periods include: Among them, x box (t), y box (t) is the center coordinate of the detection frame obtained by S2, that is, the original motion trajectory; and Correct the trajectory for the boundary.
10. The multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning according to claim 9 is characterized in that: The S4 is specifically: In the image information obtained from the side view, vehicles are assigned IDs in the order in which they appear in each lane; In the top-down image information, vehicles are assigned IDs in the order in which their trajectories appear within the field of view, and the top-down vehicle IDs are sorted by lane. The vehicles in a certain lane in the top-down view are sequentially associated with the vehicles in the corresponding lane in the side view.
Citation Information
Patent Citations
Lane line feature extraction method based on visual-correlation double spaces
CN107153823A
Spatial distribution monitoring system of whole bridge deck moving load based on dynamic weighing and multi-video information fusion
CN109167956A