Boundary reasoning enhanced road traffic load spatial-temporal distribution mode multi-view collaborative sensing method and device
Through the multi-view collaborative perception method enhanced by boundary reasoning, combined with side view angle and top view video acquisition equipment, the problem that existing systems cannot effectively monitor the vehicle's motion trajectory is solved, and the accurate identification and analysis of the spatio-temporal distribution pattern of traffic loads is achieved.
Patent Information
- Application Number
- CN202510136962.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Existing road traffic load recognition systems cannot effectively monitor the vehicle's movement trajectory, especially when the vehicle moves to the camera's field of view, resulting in abnormal trajectory recognition.
A multi-view collaborative perception method enhanced by boundary inference is adopted, and a pre-trained neural network and multi-objective tracking detection algorithm are used to identify the vehicle and correct the trajectory to obtain the load space-time distribution of the vehicle through a combination of side view and top view video acquisition equipment.
Accurate identification and correction of vehicle motion trajectory is achieved, identification abnormalities at the visual field boundary are eliminated, vehicle load and trajectory information are integrated, and the spatiotemporal distribution mode of traffic load is obtained.
Smart Images

Figure CN119942420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of highway vehicle load identification, and in particular to a boundary reasoning-enhanced highway traffic load spatiotemporal distribution pattern multi-perspective collaborative perception method and device. Background Art
[0002] With the increasing demand for road transportation, new and rebuilt roads have a trend of widening the road surface and increasing the number of lanes. In order to more effectively study the impact of traffic load on road and bridge infrastructure, and to achieve traffic flow control and abnormal warning, it is necessary not only to obtain the vehicle's own weight, but also to grasp the movement trajectory of each vehicle within a certain range, so as to establish a dynamic distribution model of traffic load for key sections.
[0003] Vehicle dynamic weighing technology (WIM) has been put into use on many highways and urban roads in China, but the current contact / non-contact WIM system only has the function of obtaining the vehicle's own weight, but cannot monitor the vehicle's running trajectory. In response to the above problems, road monitoring cameras can be used as auxiliary visual sensing equipment to identify vehicles and track trajectories through pre-trained neural networks, but the existing road monitoring system cameras generally use tilted viewing angles, which makes it difficult to quantify the trajectory. Installing the camera's viewing angle perpendicular to the road surface (using a vertical overlooking angle) is conducive to trajectory quantification, but when the vehicle moves to the boundary of the overlooking angle of view, part of the vehicle body is outside the visible range, and the length and width of the vehicle detection frame will continue to change. If the center trajectory of the vehicle detection frame is directly extracted according to the existing multi-target tracking detection technology, abnormal results will be obtained, such as shape distortion of the trajectory, unexpected reduction in vehicle speed calculated based on the trajectory, etc. In addition, most of the current vehicle tracking and recognition technologies based on video recognition only focus on the location information of the detection object, without identifying and analyzing the geometric characteristics of the object itself. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention proposes a multi-perspective collaborative perception method and device for the spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning. The specific technical solution is as follows:
[0005] A multi-view collaborative perception device for the spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning, a side-view video acquisition device, a top-view video acquisition device, and a data processing device;
[0006] The side-view video acquisition device is installed on the side of each lane of the immutable road section to identify the vehicle load sequence of the corresponding lane; the camera installation of the side-view video acquisition device needs to meet the following conditions:
[0007] (1) The camera's image plane is perpendicular to the road surface;
[0008] (2) The main optical axis of the camera is perpendicular to the direction of the road;
[0009] (3) The camera should be installed at a height that allows the tire image to be fully captured when a vehicle passes through its field of view;
[0010] The overhead video acquisition device is installed along the road extension direction and immediately after the immutable road segment, and serves as an overhead information source to obtain the vehicle driving trajectory. The installation of the overhead video acquisition device meets the following conditions:
[0011] (1) The principal optical axis of the lens is perpendicular to the road surface, and the horizontal axis of the image coordinate axis is parallel to the longitudinal axis of the road;
[0012] (2) The size of the visible area is more than twice the length of a conventional vehicle in the longitudinal direction of the road and covers all lanes in the same direction on one side of the road in the transverse direction;
[0013] (3) The vertical position of the top-view video acquisition device is immediately behind the side-view area covered by the side-view video acquisition device, and the vehicle within the visible range of the top-view video acquisition device enters a non-lane-changeable area of a preset length on one side, so as to ensure that the lane where the vehicle is located within the side-view range can be obtained in the top-view video, which is convenient for subsequent data association;
[0014] The data processing device has a built-in non-contact vehicle dynamic weighing analysis system, which obtains the real-time inflation pressure of the tires of a moving vehicle and the vehicle tire image data obtained from the side view through the vehicle-road cooperative network to solve the vehicle's single wheel load and the vehicle load; the vehicle's running trajectory is solved by obtaining the vehicle video information from the top view, and the vehicle load information obtained from the side view is associated with the vehicle's spatial trajectory obtained from the top view to obtain the vehicle's temporal and spatial distribution.
[0015] Furthermore, the bird's-eye view video acquisition device has multiple bird's-eye view angles along the longitudinal direction of the road, and adjacent bird's-eye view angles satisfy: the field of view of each bird's-eye view angle overlaps to a certain extent with the next bird's-eye view angle, and each bird's-eye view video time frame is recorded using the same reference, and each bird's-eye view angle uses a unified road surface coordinate system.
[0016] A method for multi-perspective collaborative perception of spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning is implemented based on a multi-perspective collaborative perception device for highway traffic load distribution patterns enhanced by boundary reasoning, and includes the following steps:
[0017] S1: The side view video acquisition device of each lane collects the tire image of each vehicle passing through the lane, and calculates the single wheel load of each vehicle based on the tire image, thereby obtaining the total load of the vehicle;
[0018] S2: When several cars drive out of the non-changeable lane area of the side view and enter the visible range of the top view, the top view video acquisition device starts continuous video acquisition and uses a pre-trained convolutional neural network to identify the vehicles in the picture frame by frame to obtain the detection frame and the corresponding confidence score; then, the original motion trajectory of the vehicle is identified through the multi-target tracking detection algorithm based on the BYTE strategy;
[0019] S3: Use the top-down perspective boundary vehicle trajectory inference algorithm to obtain the translational features of the vehicle detection frame boundary and the vehicle in-plane rotation features when the vehicle is completely within the field of view, infer and improve the missing information when the vehicle is partially outside the field of view boundary, and then correct the original motion trajectory in the image coordinate system, and obtain the motion trajectory when the vehicle is partially outside the field of view boundary;
[0020] S4: Associating the load information of the vehicle obtained from the side view with the corrected motion trajectory of the corresponding vehicle from the top view to obtain the spatiotemporal distribution of the traffic load.
[0021] Furthermore, in step S2, the original motion trajectory of the vehicle is identified by a multi-target tracking detection algorithm based on the BYTE strategy, which specifically includes:
[0022] S201: After identifying each frame of the video captured from the bird's-eye view, a vehicle detection frame and its detection score are obtained in each frame, and the detection score is compared with two thresholds η high With η low Comparison, 0<η low <η high <1, thus grouping the detection frames into a high-resolution group, a low-resolution group, and a background group, and directly discarding the detection frames of the background group;
[0023] The specific grouping strategy is:
[0024] When the detection score η corresponding to the jth detection box in the i-th frame i,j Satisfy η high ≤η i,j <1, the detection frame corresponding to the detection score is assigned to the high group G high,i ;
[0025] When η i,j Satisfy η low ≤η i,j <η high When , the detection frame corresponding to the detection score is divided into the low group G low,i ;
[0026] When η i,j Satisfying 0<η i,j <η low , then the detection frame corresponding to the detection score is divided into the background group G back,i ;
[0027] S202: The vehicle detection frame obtained in the first frame is set as the initial trajectory of the vehicle. From the second frame onwards, the Kalman filter is used to predict the new position of each vehicle trajectory in the current frame. The predicted trajectory set corresponding to the i-th frame is recorded as Γ i ; The initial trajectory and predicted trajectory both refer to the vehicle rectangular detection frame containing the length, width and center coordinate information;
[0028] S203: Group high high,i The set of all detection boxes and trajectories in i Perform the first round of matching. The unmatched high-scoring detection frames in this step are stored in the set G. remain,i In the above example, unmatched trajectories are stored in the set Γ. remain,i middle;
[0029] S204: Group low-level low,i The set of all detection boxes and trajectories in remain,i Perform the second round of matching. The unmatched detection frames in this step are regarded as background and no longer used. The unmatched trajectories are stored in the set Γ lost,i middle;
[0030] S205: Output the successfully matched detection box ID and trajectory ID as the matching result of the i-th frame and the i-1 frame; use the method described in S202 to obtain the predicted trajectory set Γ corresponding to the i+1 frame i+1 ; Set the regenerated video frame number threshold n lost , the i+1-nth lost The set of all unmatched trajectories from the frame image to the i-th frame Join Gamma i+1 ; G of the i-th frame described in step S203 remain,i Each detection box in the set is added to Γ i+1 , regarded as the initialization of the new trajectory;
[0031] S206: Repeat S201-S205 to complete the matching between all adjacent frames, and obtain the detection frame and its center position of each frame as the original motion trajectory of the vehicle.
[0032] Furthermore, the step S3 specifically includes the following sub-steps:
[0033] S301: Extract the critical start time t of the vehicle from the bird's-eye view video information and the original motion trajectory of the vehicle. c and critical end time t e , including the critical start time t of the vehicle entering the picture c,in and critical end time t e,in , and the critical start time t of the vehicle leaving the screen c,out and critical end time t e,out;
[0034] For the case where a vehicle enters the picture, the critical start time t c,in is the first frame of the trajectory, the critical end time t e,in The entry direction frame line of the vehicle detection frame leaves the field of view boundary and starts to move longitudinally;
[0035] For the case where the vehicle drives out of the picture, the critical start time t c,out The critical end time t is when the vehicle detection frame reaches the boundary of the visual field and stops moving longitudinally. e,out is the last frame of the trajectory;
[0036] S302: At t e,in to c,out During the period between, the vehicle detection frame includes the complete vehicle body, and the length and width of the vehicle are obtained based on the vehicle geometric feature contour homogenization recognition algorithm;
[0037] S303: When the vehicle partially enters or exits the visual field boundary, estimate the longitudinal length w of the vehicle virtual detection frame v and the acute angle θ between the vehicle's symmetry axis and the road extension direction;
[0038] S304: Calculating the length l of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary v ;
[0039] S305: Correcting the initial trajectory of the vehicle when it partially enters or exits the boundary of the field of view;
[0040] S306: Convert the corrected trajectory to the road coordinate system through a homography transformation method to obtain the actual trajectory of the vehicle.
[0041] Furthermore, in S302, the vehicle length and width are obtained according to the vehicle geometric feature contour homogenization recognition algorithm, which specifically includes the following sub-steps:
[0042] S3021: When the monitoring frame includes the complete vehicle and the vehicle is located within the set range in the center of the picture, the image in the detection frame is cropped and converted into a grayscale image; Canny edge detection is performed on the grayscale image, and then three image processing operations, namely morphological dilation, hole filling, and morphological erosion, are performed in sequence to eliminate various fine dots and lines in the background area and holes in the foreground area, so that the foreground is merged into one or more continuous complete areas;
[0043] S3022: If the detection frame in S3021 is converted into multiple foreground regions, filter out the foreground region with the largest area U max And find the center of mass P c ; If only one foreground area is converted, then its centroid is directly calculated; from the centroid P cDeparture at U max Find a point P on each of the two long sides of the boundary j , j = 1, 2, so that P c To P j The Euclidean distance is the shortest, and the shortest distances are d min,j ; In P j Search a series of pixels around each to make them all satisfy P c To P j The Euclidean distance is (1+ζ d )d min The condition where ζ d is the search magnification factor, which is 5%-10%; j The multiple pixel points obtained by searching around are defined as the candidate perpendicular point set Ψ j ;
[0044] S3023: Calculate the uniformity index of the candidate perpendicular foot point set to further determine the perpendicular foot position:
[0045] (1) j All pixel points in the image are sorted according to the first standard of increasing horizontal coordinates and the second standard of increasing vertical coordinates. If multiple points with adjacent labels have the same horizontal coordinates, the vertical coordinates are averaged and the average result is used as a new point to replace the previous points with the same horizontal coordinates. The following preprocessed candidate perpendicular foot point set is obtained:
[0046]
[0047] Among them, m j for The number of rows in the point set;
[0048] (2) Calculation The slope vector S j , S j The kth element in is:
[0049]
[0050] Where k = 2, 3…, m j ; S j,1 =0;
[0051] (3) Statistical slope vector S j The number of positive and negative elements in the array, the number of positive elements is n +j , the number of negative elements is n -j , calculate the candidate perpendicular point set Ψ j The uniformity index I ev,j as follows:
[0052]
[0053] Among them, ∑m j is the sum of the number of points in all candidate perpendicular point sets after preprocessing, ζ n is the point set density weight coefficient;
[0054] S3024: Select I ev The largest set of candidate perpendicular points is calculated, and the average coordinates of all points are used as the identification perpendicular point P. ft ; Connect the centroid P c With P ft , and make a line through the centroid and perpendicular to P c P ft The straight line is used as the longitudinal symmetry axis of the vehicle; the Euclidean distance between any two intersection points of the longitudinal symmetry axis and all foreground areas in the vehicle detection frame is calculated; the maximum value of the Euclidean distance is selected as the vehicle recognition length l, P c P ft The twice of the Euclidean distance is the vehicle recognition width w.
[0055] Furthermore, the S303 specifically includes:
[0056] For the case where the vehicle partially goes out of the boundary, a quadratic polynomial is used to calculate t e,in to c,out The horizontal coordinate x of the exit direction frame line bl Fitting to x out =f 1 (t), and x out =f 1 (t) Extrapolated to t c,out to e,out t c,out to e,out The vertical length of the virtual detection frame w v (t)=|f 1 (t)-x b,out |, where x b,out is the horizontal coordinate of the visual field boundary in the exit direction;
[0057] For the case where the vehicle partially enters the boundary, a quadratic polynomial is used to calculate t e,in to c,out The horizontal coordinate of the entry direction frame line between the two is fitted as x in =f 2 (t), and x in =f 2 (t) Extrapolated to t c,out to e,out t c,in to e,in The vertical length w of the virtual detection frame between v (t)=|f 2 (t)-xb,in |;
[0058] When the vehicle is completely within the detection frame, a quadratic polynomial is used to calculate the height h of the detection frame when the vehicle is completely within the image. t Fitting to h t =f 3 (t), and extrapolated to t c,in to e,out The acute angle θ between the vehicle's symmetry axis and the road extension direction is calculated using the following formula:
[0059]
[0060] in, w and l are the width and length of the complete vehicle respectively.
[0061] Furthermore, the length of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary in S304 is l v The calculation formula is as follows:
[0062]
[0063] in, is the angle between the vehicle's driving direction and the positive direction of the horizontal axis of the image coordinate system, The corresponding relationship with θ is as follows: If the vehicle is traveling in the first quadrant, then If it points to the second quadrant, If it points to the third quadrant, If it points to the fourth quadrant, w t (t) is the longitudinal length of the detection frame of the visible part of the vehicle that has not gone out of the boundary at time t.
[0064] Furthermore, the S305 specifically includes:
[0065] When the vehicle partially drives out of the boundary, at t c,out ≤t≤t e,out The time periods are:
[0066] When the vehicle partially enters the boundary, at t c,in ≤t≤t e,in The time periods are:
[0067] Among them, x box (t), y box (t) is the center coordinate of the detection frame obtained by S2, that is, the original motion trajectory; and Correct the trajectory for the boundary.
[0068] Furthermore, the S4 is specifically:
[0069] In the image information obtained from the side view, the vehicles are assigned IDs in the order in which they appear in each lane;
[0070] In the image information obtained from the top-down perspective, vehicles are assigned IDs in the order in which their trajectories appear in the field of view, and the top-down perspective vehicle IDs are sorted by lane;
[0071] The vehicles in a lane in the top-down view are sequentially associated with the vehicles in the corresponding lane in the side view.
[0072] The present invention has the following beneficial effects:
[0073] (1) In the present invention, the geometric characteristic contour uniformity recognition of the top-down view vehicle can be realized, and the geometric characteristics of the vehicle in any driving posture on the road surface can be obtained, thereby improving the utilization of the top-down view vehicle flow video data. The obtained recognition data can be used for subsequent trajectory correction, and can also be used as an information source for detection and early warning of special traffic conditions (such as width and length limit sections).
[0074] (2) The vehicle trajectory inference algorithm at the top-view boundary in the present invention infers and improves the missing information when part of the vehicle is outside the boundary of the field of view, eliminating the original motion trajectory recognition anomaly caused by the detection frame failing to include the entire vehicle body when the vehicle reaches the boundary of the field of view. In principle, it overcomes the defect of the traditional method of trajectory recognition based on the center position of the detection frame.
[0075] (3) The present invention associates the side-view vehicle load identification results with the top-view motion trajectory of the variable lane area, integrating the vehicle spatial position information with attribute information such as load, and obtains the spatiotemporal distribution pattern of the traffic load on the road section of interest only through a non-contact system, providing a reliable, durable and low-cost solution for traffic control, traffic load data accumulation, and service status assessment of road and bridge infrastructure. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 Schematic diagram of a multi-perspective collaborative perception device for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning provided in this embodiment.
[0077] Figure 2 A schematic diagram of the process and principle of a multi-perspective collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning provided in an embodiment.
[0078] Figure 3The flowchart and principle diagram for vehicle length and width recognition from a top-down perspective provided in this embodiment, wherein (a) is the recognition flowchart, (b) is the grayscale image, (c) is the detected edge, (d) is the result after morphological dilation, (e) is the result after the hole filling operation, and (f) is the recognition result after morphological erosion.
[0079] Figure 4 The figures are effect diagrams of different vehicle trajectory recognition and original trajectory reasoning correction in the embodiments of the present invention, where (a) is a flowchart of the vehicle trajectory reasoning algorithm at the top-down perspective boundary, (b) is a schematic diagram of the geometric parameters of the vehicle driving out of the boundary, and (c) is a schematic diagram of the virtual detection box width inference principle.
[0080] Figure 5 Schematic diagram of the original trajectory and the corrected trajectory of vehicles No. 115 and No. 15 of the embodiment, wherein (a) is the original trajectory of vehicle No. 115, (b) is the original trajectory of vehicle No. 15, (c) is the correction result of the trajectory of vehicle No. 115 using the method of the present invention, and (d) is the correction result of the trajectory of vehicle No. 15 using the method of the present invention. DETAILED DESCRIPTION
[0081] The present invention will be described in detail below based on the accompanying drawings and preferred embodiments, and the purpose and effects of the present invention will become more clear. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0082] like Figure 1 As shown, the boundary reasoning-enhanced multi-perspective collaborative perception device for spatiotemporal distribution patterns of highway traffic loads of the present invention includes: a side-perspective video acquisition device, a top-perspective video acquisition device, a vehicle-road collaborative perception device, and a data processing device.
[0083] The side-view video acquisition device is installed on the side of each lane of the immutable road section to identify the vehicle load sequence of the corresponding lane. The camera installation of the side-view video acquisition device needs to meet the following conditions:
[0084] (1) The camera's image plane is perpendicular to the road surface;
[0085] (2) The main optical axis of the camera is perpendicular to the direction of the road;
[0086] (3) The camera should be installed at a height that allows the tire image to be fully captured when most conventional vehicles pass through its field of view.
[0087] Since the side-view video acquisition equipment needs to occupy a certain space in the horizontal direction of the highway, it is recommended to choose places such as highway toll stations and checkpoints for deployment. This scenario also ensures that vehicles in the side-view detection area cannot change lanes at will. The vehicle-road cooperative perception equipment can also be installed near the camera to obtain the actual inflation pressure data of the tire pressure monitoring system (TPMS).
[0088] The overhead video acquisition device located above the road section of interest serves as an overhead information source to obtain the vehicle's driving trajectory. The installation of the overhead video acquisition device should meet the following conditions:
[0089] (1) The principal optical axis of the lens is perpendicular to the road surface, and the horizontal axis of the image coordinate axis is parallel to the longitudinal axis of the road;
[0090] (2) The size of the visible area is more than twice the full length of most conventional vehicles in the longitudinal direction of the road, and covers all lanes in the same direction on one side of the road in the transverse direction;
[0091] (3) The vertical position of the top-view video acquisition device is installed immediately after the side view area covered by the side view video acquisition device, and the side where the vehicle enters within the visible range of the top-view video acquisition device includes a preset non-lane changeable area of a certain length, so as to ensure that the lane where the vehicle is located within the side view range can be obtained in the top-view video, which is convenient for subsequent data association.
[0092] The data processing equipment has a built-in non-contact vehicle dynamic weighing analysis system, which obtains the real-time inflation pressure of the tires of a moving vehicle and the vehicle tire image data obtained from the side view through the vehicle-road cooperative network to solve the vehicle's single wheel load and the vehicle load. The vehicle's running trajectory is solved by obtaining the vehicle video information from the top view, and the vehicle load information obtained from the side view is associated with the vehicle's spatial trajectory obtained from the top view to obtain the temporal and spatial distribution of the vehicle load.
[0093] like Figure 2 As shown, the boundary reasoning-enhanced multi-perspective collaborative perception method for spatiotemporal distribution pattern of highway traffic load in this embodiment is implemented based on the above-mentioned multi-perspective collaborative perception device for spatiotemporal distribution pattern of highway traffic load, and proposes a top-down vehicle geometric feature contour homogenization recognition algorithm, which realizes the use of monocular top-down image to extract the geometric characteristics of vehicles with arbitrary driving postures on the surface; proposes a top-down boundary vehicle trajectory reasoning algorithm, which infers and completes the missing information when the vehicle partially drives out of the field of view boundary, and solves the problem that traditional trajectory recognition methods are prone to produce abnormal results at the field of view boundary; relying only on non-contact perception methods, the vehicle load information obtained from the side perspective and the vehicle trajectory distribution obtained from the top perspective are associated and fused, and the vehicle load distribution pattern of the entire field can be obtained. Specifically, it includes the following steps:
[0094] S1: The side view video acquisition device of each lane acquires the tire image of each vehicle passing through the lane, and calculates the wheel load and deadweight information of each vehicle based on the tire image.
[0095] As one of the implementation methods, the vehicle tire image data obtained from the side view and the actual inflation pressure data of the tire pressure monitoring system (TPMS) obtained by the vehicle-road cooperative network built based on the vehicle-road cooperative perception device are transmitted to the non-contact WIM analysis system to solve the single wheel load and vehicle load. The purpose of this step is to obtain the vehicle load sequence of each lane, so any system that can distinguish different vehicles and obtain the vehicle's own weight can be used as a substitute, including non-contact visual WIM system, contact piezoelectric WIM system, etc.
[0096] A non-contact WIM system that can meet the above requirements processes images and TPMS signals as follows: segment the tire image into multiple interest regions, extract key information from the tire-ground contact area, rim area, and tire sidewall text mark area images through machine vision algorithms, quantify the visual characteristics of tire deformation, and obtain the standard tire inflation pressure and various cross-sectional geometric parameters, rim parameters, and tread parameters through direct recognition or standard database retrieval methods. On the basis of obtaining the above parameters and actual inflation pressure data, a certain contact mechanics model (such as tire pressure balance model, radial compression model, Hertz contact model, elliptical constraint Hertz deformation quantification model, etc.) is selected to establish the association between tire deformation characteristics and external loads, and then solve the single wheel load and vehicle total load.
[0097] S2: When one or more vehicles drive out of the non-changeable lane area of the side view and enter the visible range of the top view, the top view video acquisition device starts continuous video acquisition and uses the pre-trained YoloV8n network to identify the vehicles in the picture frame by frame to obtain the detection frame and the corresponding confidence score; then, the original motion trajectory of the vehicle is identified through the multi-target tracking detection algorithm based on the BYTE strategy.
[0098] The original motion trajectory of the vehicle is a series of detection frames and their center coordinates (x box ,y box ). The origin of the image coordinate system is O IM Located in the upper left corner of the screen, x IM The positive direction of the axis points to the right, y IM The positive direction of the axis points downward.
[0099] The steps of the multi-target tracking detection algorithm based on BYTE strategy are as follows:
[0100] S201: After identifying each frame of the video captured from the bird's-eye view, a vehicle detection frame and its detection score are obtained in each frame, and the detection score is compared with two thresholds ηhigh With η low Comparison, 0<η low <η high <1, thus grouping the detection frames into a high-frequency group, a low-frequency group, and a background group, and directly discarding the detection frames of the background group.
[0101] The specific grouping strategy is:
[0102] When the detection score η corresponding to the jth detection box in the i-th frame i,j Satisfy η high ≤η i,j <1, the detection frame corresponding to the detection score is assigned to the high group G high,i ;
[0103] When η i,j Satisfy η low ≤η i,j <η high When , the detection frame corresponding to the detection score is divided into the low group G low,i ;
[0104] When η i,j Satisfying 0<η i,j <η low , then the detection frame corresponding to the detection score is divided into the background group G back,i .
[0105] Since the vertical viewing angle reduces the uncertainty of recognition, the detection frame with low confidence is likely to be a false detection of the background, so G back,i Abandonment is reasonable.
[0106] S202: The vehicle detection frame obtained in the first frame is set as the initial trajectory of the vehicle. From the second frame onwards, the Kalman filter is used to predict the new position of each trajectory in the current frame. The predicted trajectory set of the i-th frame is recorded as Γ i ; Both the initial trajectory and the predicted trajectory refer to the vehicle rectangular detection box containing the length, width and center coordinate information;
[0107] S203: Group high high,i The set of all detection boxes and trajectories in i Perform the first round of matching. The unmatched high-scoring detection frames in this step are stored in the set G. remain,i In the above example, unmatched trajectories are stored in the set Γ. remain,i middle;
[0108] S204: Group low-level low,i The set of all detection boxes and trajectories in remain,i Perform the second round of matching. The unmatched detection frames in this step are regarded as background and no longer used. The unmatched trajectories are stored in the set Γ lost,i middle.
[0109] The two-round detection mechanism of steps S203 and S204 increases the chance of matching and is more applicable to situations where there may be accidental occlusion of the target, sudden changes in the size of the target detection frame, etc.
[0110] S205: Output the successfully matched detection box ID and trajectory ID as the matching result of the i-th frame and the i-1 frame; use the method described in S202 to obtain the predicted trajectory set Γ corresponding to the i+1 frame i+1 ; Set the regenerated video frame number threshold n lost , the i+1-nth lost The set of all unmatched trajectories from the frame image to the i-th frame Join Gamma i+1 ; G of the i-th frame described in step S203 remain,i Each detection box in the set is added to Γ i+1 , regarded as the initialization of the new trajectory.
[0111] Because G remain,i Although the detection box in has a high detection score, it cannot match any of the aforementioned trajectories. Therefore, it is usually a vehicle that suddenly merges from both sides, and it is necessary to initialize it as a new trajectory.
[0112] S206: Repeat steps S201-S205 to complete the matching between all adjacent frames, and obtain the detection frame and its center position of each frame as the original motion trajectory of the vehicle.
[0113] In the above process, it is necessary to determine the specific solution for optimal matching of multiple detection boxes and multiple predicted trajectories in each round. In practical applications, the intersection over union (IoU) similarity can be used as the matching basis, and the matching can be completed through the classic Hungarian algorithm, as follows:
[0114] (1) Calculate the IoU similarity of all detection boxes and predicted trajectories to be matched in the same frame to form a cost matrix. The IoU similarity calculation method between the jth detection box and the kth predicted trajectory in the i-th frame is:
[0115]
[0116] Where S(·) represents the area of the region of interest.
[0117] (2) Let the number of detection boxes required to be matched in a certain round be n G , the number of predicted trajectories required to match is n Γ , if the two are not equal, assume that n Γ is the larger one, the IoU similarity calculation results can be summarized as n Γ ×n Γ The cost matrix is:
[0118]
[0119] If n G ≠n Γ , the missing elements in the corresponding positions in the square matrix are filled with 0; if the two are equal, a square matrix is directly constructed.
[0120] (3) The matching of the detection box and the predicted trajectory is analogized to an “assignment problem” and the cost matrix C is processed using the Hungarian algorithm. IoU,i , you can get the best matching solution.
[0121] S3: Use the top-down perspective boundary vehicle trajectory inference algorithm to obtain the translational features of the vehicle detection frame boundary and the vehicle in-plane rotation features when the vehicle is completely within the field of view, infer and improve the missing information when the vehicle is partially outside the field of view boundary, and then correct the original motion trajectory in the image coordinate system, and obtain the motion trajectory when the vehicle is partially outside the field of view boundary.
[0122] S301: Extract the critical start time t of the vehicle from the bird's-eye view video information and the original motion trajectory of the vehicle. c and critical end time t e , including the critical start time t of the vehicle entering the picture c,in and critical end time t e,in , and the critical start time t of the vehicle leaving the screen c,out and critical end time t e,out ;
[0123] For the case where a vehicle enters the picture, the critical start time t c,in is the first frame of the trajectory, the critical end time t e,in The entry direction frame line of the vehicle detection frame leaves the field of view boundary and starts to move longitudinally;
[0124] For the case where the vehicle drives out of the picture, the critical start time t c,out The critical end time t is when the vehicle detection frame reaches the boundary of the visual field and stops moving longitudinally. e,out is the last frame of the trajectory.
[0125] A feasible method for obtaining critical moments is given below.
[0126] (1) For the complete trajectory of a vehicle in the image, obtain the horizontal coordinate x of the entry direction frame line of the detection frame corresponding to each point on it br and the horizontal coordinate x of the exit direction frame bl , get the horizontal coordinate x of the entry direction frame line br Sequence and exit direction frame horizontal coordinate x bl sequence.
[0127] For the sake of description, it is assumed that the vehicle enters the screen from right to left, so the frame line in the direction of entry is the "right frame line" and the frame line in the direction of exit is the "left frame line". e ) and the detection box width (w b ) Calculate the horizontal coordinate x of the entry direction frame line br and the horizontal coordinate x of the exit direction frame bl The method is as follows:
[0128] x bl =x e -0.5w b
[0129] x br =x e +0.5w b
[0130] So far, we get the horizontal coordinate x of the entry direction frame line at equal time intervals. br Sequence and exit direction frame horizontal coordinate x bl sequence.
[0131] (2) Calculate the horizontal coordinate x of the entry direction frame line br Sequence and exit direction frame horizontal coordinate x bl The first-order difference of the sequence, Δx bl Sequence and Δx br Sequence, and according to the length of the sequence, select the appropriate period to perform moving average on the first-order difference sequence to enhance the smoothness of its non-mutation segment, and obtain the sequence Δx bl With Δx br ;
[0132] (3) Continue to calculate Δx after moving average processing bl Sequence and Δx br The difference of the sequence, that is, the second-order difference Δ 2 x bl Sequence and Δ 2 x br Sequence, find the moment when the absolute value of the second-order difference sequence is the largest, |Δ 2 x br The moment corresponding to the maximum value is the critical end moment t of the vehicle entering the screen e,in , |Δ 2 x bl The time corresponding to the maximum value is the critical start time t of the vehicle leaving the screen c,out .
[0133] Here x br Taking the sequence as an example (assuming that the number of its initial elements is n), the specific calculation methods of (2) and (3) are given:
[0134] The specific method of the first-order difference mentioned above is: Δx bl,i =x bl,i+1 -x bl,i (i=2,3…n-1);
[0135] Assuming that the number of periods used in the above moving average method is N, then The reason why it is nN here is that if the moving average uses N periods, the sequence length will be shortened by N-1.
[0136] The specific method of the second-order difference mentioned above is:
[0137] S302: At t e,in to c,out During the period between , the vehicle detection frame includes the complete vehicle body, and the length and width of the vehicle are obtained based on the vehicle geometric feature contour homogenization recognition algorithm, such as Figure 3 As shown, it specifically includes the following sub-steps:
[0138] S3021: When the monitoring frame includes the complete vehicle and the vehicle is located within the set range in the center of the picture, the image in the detection frame is cropped and converted into a grayscale image; Canny edge detection is performed on the grayscale image, and then three image processing operations, namely morphological dilation, hole filling, and morphological erosion, are performed in sequence to eliminate various fine dots and lines in the background area and holes in the foreground area, so that the foreground is merged into one or more continuous complete areas;
[0139] S3022: If the detection frame in S3021 is converted into multiple foreground regions, filter out the foreground region with the largest area U max And find the center of mass P c ; If only a foreground area is converted, then its centroid is directly calculated (from the vehicle structure characteristics, U max Approximately a rectangle, and the long sides of both sides are nearly parallel to the direction of vehicle travel); from the center of mass P c Departure at U max Find a point P on each of the two long sides of the boundary j , j = 1, 2, so that P c To P j The Euclidean distance is the shortest, and the shortest distances are d min,j ; In P j Search a series of pixels around each to make them all satisfy P c To P j The Euclidean distance is (1+ζ d )d min The condition where ζ d is the search magnification factor, which can be 5%-10%. jThe multiple pixel points obtained by searching around are defined as the candidate perpendicular point set Ψ j .
[0140] S3023: Calculate the uniformity index of the candidate perpendicular foot point set to further determine the perpendicular foot position. The steps are as follows: j All pixel points in the image are sorted according to the first standard of increasing horizontal coordinates and the second standard of increasing vertical coordinates. If multiple points with adjacent labels have the same horizontal coordinates, the vertical coordinates are averaged and the average result is used as a new point to replace the previous points with the same horizontal coordinates. The following preprocessed candidate perpendicular foot point set is obtained:
[0141]
[0142] At this time there must be Each is not equal, m j for The number of rows in the point set.
[0143] calculate The slope vector S j . S j The kth element in (k=2,3…,m j )for:
[0144]
[0145] In particular, another j,1 = 0. If There is only one point in the same way. j,1 =0;
[0146] Statistical slope vector S j The number of positive and negative elements in the array, the number of positive elements is n +j , the number of negative elements is n -j Finally, the candidate perpendicular point set Ψ is calculated j The uniformity index is as follows:
[0147]
[0148] Where ∑m j is the sum of the number of points in all candidate perpendicular point sets after preprocessing, ζ n is the point set density weight coefficient. The first term of the expression represents the monotonicity and uniformity of the boundary change near the candidate perpendicular foot, reflecting the degree of interference in boundary identification. The second term reflects the proportion of the number of candidate perpendicular feet contained in the point set.
[0149] S3024: Select I ev The largest set of candidate perpendicular points is calculated, and the average coordinates of all points are used as the identification perpendicular point P. ft ; Connect the centroid P cWith P ft , and make a line through the centroid and perpendicular to P c P ft The straight line is taken as the longitudinal symmetry axis of the vehicle. The longitudinal symmetry axis may have 2 or more intersections with all foreground areas in the vehicle detection frame. The Euclidean distance between any two intersections is calculated, and the maximum value of the Euclidean distance is selected as the vehicle recognition length l; the vehicle recognition width w is P c P ft Twice the Euclidean distance.
[0150] In actual operation, you can e,in to c,out 3-5 frames are randomly captured for analysis, and the length and width calculation results are averaged.
[0151] S303: When the vehicle partially enters or exits the field of view boundary v Make inferences with θ;
[0152] When a vehicle partially enters or exits the boundary of the field of view, first obtain the longitudinal length of the detection frame w t , horizontal height is h t ; Assume that the part of the vehicle body outside the field of view is also inscribed in a rectangular virtual detection frame. The virtual detection frame is the circumscribed rectangle of the invisible part of the vehicle body, and the two sides of the rectangle are parallel to the two coordinate axes of the top-down view image coordinate system; calculate the longitudinal length of the vehicle virtual detection frame as w v , horizontal height is h v , and the acute angle θ formed by the vehicle's symmetry axis and the road extension direction (horizontal axis of the image coordinate system) when the vehicle partially enters or exits.
[0153] like Figure 4 As shown, when the vehicle partially enters or exits the field of view boundary, w v The specific expression is as follows:
[0154] According to the horizontal coordinate x of the frame line in the direction of entry br and the horizontal coordinate x of the exit direction frame bl Regarding the time variation curve, we can know that the horizontal coordinate x of the exit direction frame line bl In t c,in to c,out This segment is fitted with a quadratic polynomial as x out =f 1 (t); horizontal coordinate t of the frame line of the exit direction c,out to e,out The horizontal line is a horizontal line, and its abscissa value is the abscissa x of the boundary of the visual field in the direction of exit. b,out . The fitting result x out =f 1 (t) Extrapolated to t c,out toe,out When t c,out ≤t≤t e,out , the vertical length of the virtual detection frame can be calculated as: w v (t)=|f 1 (t)-x b,out |;
[0155] Similarly, we can know that the horizontal coordinate of the entry direction frame line is at t e,in to e,out This segment is fitted with a quadratic polynomial as x in =f 2 (t); horizontal coordinate of the frame line of the driving direction t c,in to e,in The horizontal line is a horizontal line, and its abscissa value is the abscissa x of the visual field boundary in the driving direction. b,in . The fitting result x in =f 2 (t) Extrapolated to t c,in to e,in When t c,in ≤t≤t e,in , the vertical length of the virtual detection frame can be calculated as: w v (t)=|f 2 (t)-x b,in |;
[0156] When the vehicle is completely within the frame (i.e. t e,in ≤t≤t c,out ), a series of changing detection box heights h can be obtained t , which is fitted with a quadratic polynomial as h t =f 3 (t), and extrapolated to t c,in to e,out Time period. Assume that the acute angle between the vehicle symmetry axis and the road extension direction (horizontal axis of the image coordinate system) is θ, then:
[0157]
[0158] in w and l are the vehicle recognition length and recognition width obtained in step 3 respectively.
[0159] S304: Calculating the length l of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary v as follows:
[0160]
[0161] in, is the angle between the vehicle's driving direction and the positive direction of the horizontal axis of the image coordinate system, The corresponding relationship with θ is as follows: If the vehicle is traveling in the first quadrant, then If it points to the second quadrant, If it points to the third quadrant, If it points to the fourth quadrant,
[0162] S305: Correct the initial trajectory of the vehicle when it partially enters or exits the boundary of the field of view using the following formula:
[0163] When the vehicle partially drives out of the boundary, at t c,out ≤t≤t e,out The time periods are:
[0164] Similarly, when the vehicle partially enters the boundary, at t c,in ≤t≤t e,in The time periods are:
[0165] Among them, x box (t), y box (t) is the center coordinate of the detection box obtained in step 2, i.e., the initial trajectory; and Correct the trajectory for the boundary;
[0166] S306: Since the boundary correction trajectory is in the image coordinate system x IM O IM y IM Therefore, it is necessary to transform the corrected trajectory, that is, the coordinate position of the vehicle at different times, from the image coordinate system to the road coordinate system through the homography transformation method to obtain the real trajectory of the vehicle.
[0167] Among them, the homography transformation method is as follows:
[0168] Assume that the coordinates of a point on the vehicle trajectory after boundary correction in the image coordinate system are (x IM ,y IM ), the coordinates in the road coordinate system are (x RD ,y RD ), then the two have the following relationship:
[0169]
[0170] The above relationship can be summarized as: RD =HP IM , H is the homography matrix. Using (H 31 x IM +H 32 y IM +1) and multiply the left side of the first two rows of the above formula, and shift the terms to get:
[0171] H 11 x IM +H 12 y IM +H 13 -H 31 x IM x RD -H 32 y IM x RD =x RD
[0172] H 21 x IM +H 22 y IM +H 23 -H 31 x IM y RD -H 32 y IM y RD =y RD
[0173] Therefore, the eight unknown quantities of the homography matrix can be calculated using four marker points whose road coordinates are known, within the visible range of the top-down view, and not collinear. At this point, eight independent equations can solve all the unknown elements.
[0174] After obtaining the actual trajectory of the vehicle, the vehicle speed can be calculated through a series of point coordinates in the actual trajectory and its time frame. The origin of the road coordinate system can be determined according to the on-site features, and the coordinate axis x r Parallel to the road extension direction, coordinate axis y r Perpendicular to the direction of the road.
[0175] S4: The side-view vehicle load identification results of the non-laneable area are associated with the following top-view motion trajectory according to the principle of matching in column order. The spatiotemporal distribution of traffic load is obtained by integrating vehicle trajectory with attribute information such as load and axle load.
[0176] All vehicles in the top-down view are sorted according to the lanes they enter, that is, the unified global vehicle ID in the top-down view is reconstructed according to the order of vehicles entering each lane, and then associated with each vehicle identified in the side view.
[0177] The specific column order matching principles are:
[0178] (1) ID numbering rule: In the image information obtained from the side view, vehicles are assigned IDs in the order in which they appear in each lane, for example, ID side =(i,j) represents the jth vehicle that appears in lane i; in the image information obtained from the top-down perspective, the vehicles are assigned IDs in the order in which their trajectories appear in the field of view, for example, IDvert = k represents the kth vehicle that appears within the visible range of the top-down perspective;
[0179] (2) Lane determination rule: Since there is a certain area of non-changeable lanes (solid road line) in the direction of vehicle entry from the top-down perspective, when the first three frames of the actual trajectory of the vehicle obtained in step 4 are all located in a certain lane, the vehicle is considered to come from that lane;
[0180] (3) Association criteria: Based on rule (2), the top-view vehicle IDs can be sorted by lane. vert = k is determined to be from lane m and its position in the lane is n, so it is compared with ID side =(m,n) vehicle association.
[0181] If there is a side view and several subsequent top-down views, the vehicle trajectories between adjacent top-down views are matched in the following manner: (1) Ensure that the field of view of each top-down view overlaps with that of the next top-down view, and that the video time frames of each top-down view are recorded using the same reference, and that each top-down view uses a unified road coordinate system; (2) Obtain the boundary-corrected vehicle trajectory in each top-down view; (3) If the two vehicle trajectories obtained from adjacent top-down views match within the overlapping field of view, or the Euclidean distance between the centers of the detection frames calculated frame by frame is less than a certain threshold, the two trajectories are automatically matched.
[0182] In order to demonstrate the application effect of the method of the present invention, this embodiment uses a DJI Phantom 4RTK drone as a bird's-eye view video collector, whose camera is CMOS (1 inch), with a video resolution of 3840×2160 and a frame rate of 30 frames per second. The video acquisition location is a highway toll station in Zhejiang Province, and the drone hovers vertically above the variable lane area behind the toll station.
[0183] The bird's-eye view area set in this embodiment includes the first and second lanes and half of the third lane of the road section after the toll station, where half of the third lane is used to increase interference and recognition difficulty. The above-mentioned first lane is located on the roadside, and the third lane is close to the center line of the road. When performing the test, there was an engineering vehicle parked in the first lane about 60m away from the toll station. Although it was outside the bird's-eye view range, it would cause a large number of vehicles entering from the first lane to change lanes, increasing the richness of the trajectory. The drone's field of view is slightly larger than the pre-set bird's-eye view test area, which avoids the incompleteness of the test area caused by wind-induced dithering.
[0184] The machine vision program of this embodiment is based on the pre-trained YoloV8n network. The training method is: first use the MS-COCO large image dataset for pre-training, and then migrate to the 1230 overhead view vehicle database established in the early research of this embodiment for fine-tuning. All images in the overhead view vehicle database are manually finely labeled using the LabelImg tool. The data processing equipment (industrial control computer) equipped with neural networks and algorithms uses NVIDIA TM The GeForce RTX 4060 GPU is the computing core. After testing, it is found that the multi-target tracking and detection algorithm based on the BYTE strategy of the present invention takes an average of 11.7ms per frame, which is significantly faster than the video frame rate, indicating that the software and hardware system has the ability of real-time tracking.
[0185] like Figure 5 As shown, Figure 5 (a) and (b) in the figure show the detection frames and initial motion trajectories of vehicles with ID numbers 115 and 15 respectively. Vehicle 115 is a light flatbed truck, and vehicle 15 is a heavy flatbed truck, the latter of which is significantly longer than the former. From the above trajectory curve, it can be seen that when approaching the edge of the field of view, the longitudinal length of the detection frame decreases, resulting in abnormal bending of the driving trajectory in this area, and the longer the vehicle body, the more significant the abnormality.
[0186] Figure 5 (c) and (d) in the figure respectively show the trajectories of vehicles with ID numbers 115 and 15 after being corrected by the vehicle trajectory inference algorithm at the top-view boundary proposed by the present invention. Compared with the original trajectory, the abnormal bends near the boundary of the field of view have been repaired; and the original trajectory has a denser distribution of trajectory points at the boundary of the field of view, which means that the speed is reduced, while the corrected trajectory shows a uniform speed or uniform acceleration state, which is more in line with the actual driving condition of the vehicle.
[0187] Those skilled in the art can understand that the above are only preferred examples of the invention and are not intended to limit the invention. Although the invention is described in detail with reference to the above examples, those skilled in the art can still modify the technical solutions recorded in the above examples or replace some of the technical features therein with equivalents. Any modification, equivalent replacement, etc. made within the spirit and principle of the invention shall be included in the protection scope of the invention.
Claims
1. A multi-view collaborative perception device for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning, characterized in that: Side-view video acquisition equipment, top-view video acquisition equipment and data processing equipment; The side-view video acquisition device is installed on the side of each lane of the immutable road section to identify the vehicle load sequence of the corresponding lane; the camera installation of the side-view video acquisition device needs to meet the following conditions: (1) The camera's image plane is perpendicular to the road surface; (2) The main optical axis of the camera is perpendicular to the direction of the road; (3) The camera should be installed at a height that allows the tire image to be fully captured when a vehicle passes through its field of view; The overhead video acquisition device is installed along the road extension direction and immediately after the immutable road segment, and serves as an overhead information source to obtain the vehicle driving trajectory. The installation of the overhead video acquisition device meets the following conditions: (1) The principal optical axis of the lens is perpendicular to the road surface, and the horizontal axis of the image coordinate axis is parallel to the longitudinal axis of the road; (2) The size of the visible area is more than twice the length of a conventional vehicle in the longitudinal direction of the road and covers all lanes in the same direction on one side of the road in the transverse direction; (3) The vertical position of the top-view video acquisition device is immediately behind the side-view area covered by the side-view video acquisition device, and the vehicle within the visible range of the top-view video acquisition device enters a non-lane-changeable area of a preset length on one side, so as to ensure that the lane where the vehicle is located within the side-view range can be obtained in the top-view video, which is convenient for subsequent data association; The data processing device has a built-in non-contact vehicle dynamic weighing analysis system, which obtains the real-time inflation pressure of the tires of a moving vehicle and the vehicle tire image data obtained from the side view through the vehicle-road cooperative network to solve the vehicle's single wheel load and the vehicle load; the vehicle's running trajectory is solved by obtaining the vehicle video information from the top view, and the vehicle load information obtained from the side view is associated with the vehicle's spatial trajectory obtained from the top view to obtain the vehicle's temporal and spatial distribution.
2. The multi-view collaborative perception device for spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 1 is characterized in that: The bird's-eye view video acquisition device has multiple bird's-eye view angles along the longitudinal direction of the road, and adjacent bird's-eye view angles satisfy: the field of view of each bird's-eye view angle overlaps with the next bird's-eye view angle to a certain extent, and each bird's-eye view video time frame is recorded using the same reference, and each bird's-eye view angle uses a unified road surface coordinate system.
3. A multi-view collaborative perception method for spatiotemporal distribution patterns of highway traffic loads enhanced by boundary reasoning, characterized in that: The method is implemented based on the boundary reasoning enhanced highway traffic load distribution pattern multi-view collaborative perception device according to claim 1, comprising the following steps: S1: The side view video acquisition device of each lane collects the tire image of each vehicle passing through the lane, and calculates the single wheel load of each vehicle based on the tire image, thereby obtaining the total load of the vehicle; S2: When several cars drive out of the non-changeable lane area of the side view and enter the visible range of the top view, the top view video acquisition device starts continuous video acquisition and uses a pre-trained convolutional neural network to identify the vehicles in the picture frame by frame to obtain the detection frame and the corresponding confidence score; then, the original motion trajectory of the vehicle is identified through the multi-target tracking detection algorithm based on the BYTE strategy; S3: Use the top-down perspective boundary vehicle trajectory inference algorithm to obtain the translational features of the vehicle detection frame boundary and the vehicle in-plane rotation features when the vehicle is completely within the field of view, infer and improve the missing information when the vehicle is partially outside the field of view boundary, and then correct the original motion trajectory in the image coordinate system, and obtain the motion trajectory when the vehicle is partially outside the field of view boundary; S4: Associating the load information of the vehicle obtained from the side view with the corrected motion trajectory of the corresponding vehicle from the top view to obtain the spatiotemporal distribution of the traffic load.
4. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 3 is characterized in that: In step S2, the original motion trajectory of the vehicle is identified by a multi-target tracking detection algorithm based on the BYTE strategy, which specifically includes: S201: After identifying each frame of the video captured from the bird's-eye view, a vehicle detection frame and its detection score are obtained in each frame, and the detection score is compared with two thresholds η high With η low Comparison, 0<η low <η high <1, thus grouping the detection frames into a high-resolution group, a low-resolution group, and a background group, and directly discarding the detection frames of the background group; The specific grouping strategy is: When the detection score η corresponding to the jth detection box in the i-th frame i,j Satisfy η high ≤η i,j <1, the detection frame corresponding to the detection score is assigned to the high group G high,i ; When η i,j Satisfy η low ≤η i,j <η high When , the detection frame corresponding to the detection score is divided into the low group G low,i ; When η i,j Satisfying 0<η i,j <η low , then the detection frame corresponding to the detection score is divided into the background group G back,i ; S202: The vehicle detection frame obtained in the first frame is set as the initial trajectory of the vehicle. From the second frame onwards, the Kalman filter is used to predict the new position of each vehicle trajectory in the current frame. The predicted trajectory set corresponding to the i-th frame is recorded as Γ i ; The initial trajectory and predicted trajectory both refer to the vehicle rectangular detection frame containing the length, width and center coordinate information; S203: Group high high,i The set of all detection boxes and trajectories in i Perform the first round of matching. The unmatched high-scoring detection frames in this step are stored in the set G. remain,i In the above example, unmatched trajectories are stored in the set Γ. remain,i middle; S204: Group low-level low,i The set of all detection boxes and trajectories in remain,i Perform the second round of matching. The unmatched detection frames in this step are regarded as background and no longer used. The unmatched trajectories are stored in the set Γ lost,i middle; S205: Output the successfully matched detection box ID and trajectory ID as the matching result of the i-th frame and the i-1 frame; use the method described in S202 to obtain the predicted trajectory set Γ corresponding to the i+1 frame i+1 ; Set the regenerated video frame number threshold n lost , the i+1-nth lost The set of all unmatched trajectories from the frame image to the i-th frame …Γ lost,i Join Gamma i+1 ; G of the i-th frame described in step S203 remain,i Each detection box in the set is added to Γ i+1 , regarded as the initialization of the new trajectory; S206: Repeat S201-S205 to complete the matching between all adjacent frames, and obtain the detection frame and its center position of each frame as the original motion trajectory of the vehicle.
5. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 4 is characterized in that: The step S3 specifically includes the following sub-steps: S301: Extract the critical start time t of the vehicle from the bird's-eye view video information and the original motion trajectory of the vehicle. c and critical end time t e , including the critical start time t of the vehicle entering the picture c,in and critical end time t e,in , and the critical start time t of the vehicle leaving the screen c,out and critical end time t e,out ; For the case where a vehicle enters the picture, the critical start time t c,in is the first frame of the trajectory, the critical end time t e,in The entry direction frame line of the vehicle detection frame leaves the field of view boundary and starts to move longitudinally; For the case where the vehicle drives out of the picture, the critical start time t c,out The critical end time t is when the vehicle detection frame reaches the boundary of the visual field and stops moving longitudinally. e,out is the last frame of the trajectory; S302: At t e,in to c,out During the period between, the vehicle detection frame includes the complete vehicle body, and the length and width of the vehicle are obtained based on the vehicle geometric feature contour homogenization recognition algorithm; S303: When the vehicle partially enters or exits the visual field boundary, estimate the longitudinal length w of the vehicle virtual detection frame v and the acute angle θ between the vehicle's symmetry axis and the road extension direction; S304: Calculating the length l of the invisible portion of the vehicle when the vehicle partially enters or exits the visual field boundary v ; S305: Correcting the initial trajectory of the vehicle when it partially enters or exits the boundary of the field of view; S306: Convert the corrected trajectory to the road coordinate system through a homography transformation method to obtain the actual trajectory of the vehicle.
6. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 5 is characterized in that: In S302, the vehicle length and width are obtained according to the vehicle geometric feature contour homogenization recognition algorithm, which specifically includes the following sub-steps: S3021: When the monitoring frame includes the complete vehicle and the vehicle is located within the set range in the center of the picture, the image in the detection frame is cropped and converted into a grayscale image; Canny edge detection is performed on the grayscale image, and then three image processing operations, namely morphological dilation, hole filling, and morphological erosion, are performed in sequence to eliminate various fine dots and lines in the background area and holes in the foreground area, so that the foreground is merged into one or more continuous complete areas; S3022: If the detection frame in S3021 is converted into multiple foreground regions, filter out the foreground region with the largest area U max And find the center of mass P c ; If it is only converted to a foreground area, then directly find its centroid; From the centroid P c Departure at U max Find a point P on each of the two long sides of the boundary j , j = 1, 2, so that P c To P j The Euclidean distance is the shortest, and the shortest distances are d min,j ; In P j Search a series of pixels around each to make them all satisfy P c To P j The Euclidean distance is (1+ζ d )d min The condition where ζ d is the search magnification factor, which is 5%-10%; j The multiple pixel points obtained by searching around are defined as the candidate perpendicular point set Ψ j ; S3023: Calculate the uniformity index of the candidate perpendicular foot point set to further determine the perpendicular foot position: (1) j All pixel points in the image are sorted according to the first criterion of increasing horizontal coordinates and the second criterion of increasing vertical coordinates. If multiple points with adjacent labels have the same horizontal coordinates, the vertical coordinates are averaged and the average result is used as a new point to replace the previous points with the same horizontal coordinates. The following preprocessed candidate perpendicular foot point set is obtained: Among them, m j for The number of rows in the point set; (2) Calculation The slope vector S j , S j The kth element in is: Where k = 2, 3…, m j ; S j,1 =0; (3) Statistical slope vector S j The number of positive and negative elements in the array, the number of positive elements is n +j , the number of negative elements is n -j , calculate the candidate perpendicular point set Ψ j The uniformity index I ev,j as follows: Among them, ∑m j is the sum of the number of points in all candidate perpendicular point sets after preprocessing, ζ n is the point set density weight coefficient; S3024: Select I ev The largest set of candidate perpendicular points is calculated, and the average coordinates of all points are used as the identification perpendicular point P. ft ; Connect the centroid P c With P ft , and make a line through the centroid and perpendicular to P c P ft The straight line is used as the longitudinal symmetry axis of the vehicle; the Euclidean distance between any two intersection points of the longitudinal symmetry axis and all foreground areas in the vehicle detection frame is calculated; the maximum value of the Euclidean distance is selected as the vehicle recognition length l, P c P ft The twice of the Euclidean distance is the vehicle recognition width w.
7. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 6 is characterized in that: The S303 specifically includes: For the case where the vehicle partially goes out of the boundary, a quadratic polynomial is used to calculate t e,in to c,out The horizontal coordinate x of the exit direction frame line bl Fitting to x out =f1(t), and x out = f1(t) extrapolated to t c,out to e,out t c,out to e,out The vertical length of the virtual detection frame w v (t)=|f1(t)-x b,out |, where x b,out is the horizontal coordinate of the visual field boundary in the exit direction; For the case where the vehicle partially enters the boundary, a quadratic polynomial is used to calculate t e,in to c,out The horizontal coordinate of the entry direction frame line between the two is fitted as x in =f2(t), and x in = f2(t) extrapolated to t c,out to e,out t c,in to e,in The vertical length w of the virtual detection frame between v (t)=|f2(t)-x b,in |; When the vehicle is completely within the detection frame, a quadratic polynomial is used to calculate the height h of the detection frame when the vehicle is completely within the image. t Fitting to h t = f3(t), and extrapolated to t c,in to e,out The acute angle θ between the vehicle's symmetry axis and the road extension direction is calculated using the following formula: in, w and l are the width and length of the complete vehicle respectively.
8. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 7 is characterized in that: The length of the invisible portion of the vehicle when the vehicle portion enters or exits the visual field boundary in S304 is l v The calculation formula is as follows: in, is the angle between the vehicle's driving direction and the positive direction of the horizontal axis of the image coordinate system, The corresponding relationship with θ is as follows: If the vehicle is traveling in the first quadrant, then If it points to the second quadrant, If it points to the third quadrant, If it points to the fourth quadrant, w t (t) is the longitudinal length of the detection frame of the visible part of the vehicle that has not gone out of the boundary at time t.
9. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 8 is characterized in that: The S305 specifically includes: When the vehicle partially drives out of the boundary, at t c,out ≤t≤t e,out The time periods are: When the vehicle partially enters the boundary, at t c,in ≤t≤t e,in The time periods are: Among them, x box (t), y box (t) is the center coordinate of the detection frame obtained by S2, that is, the original motion trajectory; and Correct the trajectory for the boundary.
10. The multi-view collaborative perception method of spatiotemporal distribution pattern of highway traffic load enhanced by boundary reasoning according to claim 9 is characterized in that: The S4 is specifically: In the image information obtained from the side view, the vehicles are assigned IDs in the order in which they appear in each lane; In the image information obtained from the top-down perspective, vehicles are assigned IDs in the order in which their trajectories appear in the field of view, and the top-down perspective vehicle IDs are sorted by lane; The vehicles in a lane in the top-down view are sequentially associated with the vehicles in the corresponding lane in the side view.
Citation Information
Patent Citations
Lane line feature extraction method based on visual-correlation double spaces
CN107153823A
Spatial distribution monitoring system of whole bridge deck moving load based on dynamic weighing and multi-video information fusion
CN109167956A
Bridge floor traffic flow load space-time distribution monitoring system and method based on computer vision
CN117036299A