Vehicle multi-source data matching method based on space-time constraint and cost function optimization

CN122618576BActive Publication Date: 2026-09-29TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611087781.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-29
Estimated Expiration
2046-07-22

AI Technical Summary

Technical Problem

[0008]本发明的目的就是为了克服上述现有技术存在的缺陷而提供一种基于时空约束与代价函数优化的车辆多源数据匹配方法,能够实现视觉观测数据与WIM数据的车辆级关联,并对匹配过程及匹配结果进行量化评价,解决现有技术中匹配准确性低、稳定性差、融合可靠性难以评价、匹配过程可解释性不足的问题

Benefits of technology

本发明首先针对视觉图像进行车辆检测与跟踪处理以得到视觉观测集合,从WIM记录中提取WIM信息并生成对应于目标WIM记录的参考时空位置;之后以目标WIM记录的参考时空位置为基准,在预设时间窗内,从视觉观测集合中筛选出视觉候选目标,构建候选匹配集合;最后通过计算候选匹配集合中各视觉候选目标的匹配总代价,并将匹配总代价最小的视觉候选目标作为最优匹配结果,将最优匹配结果的匹配总代价与预设匹配阈值进行比较,以确定是否匹配成功,若匹配成功,则建立目标WIM记录与最优匹配结果之间的对应关系,并输出相对应的匹配解释信息及量化评价结果,由此实现视觉数据与WIM数据的精准匹配关联,有效解决现有技术中匹配准确性低、稳定性差、可解释性不足的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618576B_ABST
    Figure CN122618576B_ABST
Patent Text Reader

Abstract

The application relates to a vehicle multi-source data matching method based on space-time constraints and cost function optimization, which comprises the following steps: vehicle detection and tracking processing are carried out on a bridge scene visual image to obtain a visual observation set; WIM information is extracted from a WIM record and a reference space-time position is generated to screen a plurality of visual candidate targets from the visual observation set within a preset time window; the matching total cost of each visual candidate target is calculated, the visual candidate target with the minimum matching total cost is taken as an optimal matching result, and the optimal matching result is compared with a preset matching threshold to determine whether matching is successful; if the matching is successful, a corresponding relationship between a target WIM record and the optimal matching result is established, and matching explanation information and quantitative evaluation results are output. Compared with the prior art, the application can realize vehicle-level correlation of visual observation data and WIM data, and solve the problems that the multi-source data fusion process depends on experience rules, the matching reliability is difficult to quantify, and the cause of the mis-matching is difficult to trace.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation technology, and in particular to a method for matching multi-source vehicle data based on spatiotemporal constraints and cost function optimization. Background Technology

[0002] As a core component of transportation infrastructure, bridges continuously bear various traffic loads during their long-term service. The frequent passage of heavy vehicles can easily exacerbate fatigue damage to bridge structures, increasing operational safety risks. Therefore, in bridge traffic monitoring scenarios, accurately identifying vehicle traffic status, obtaining vehicle load levels, and assessing the safety risks of heavy vehicle following are of significant engineering importance for bridge structural safety assessment, fatigue life prediction, and proactive early warning.

[0003] Currently, the two core technologies commonly used in bridge traffic monitoring are Weighing In Motion (WIM) systems and video vision systems. WIM systems can accurately collect load-related information such as timestamps, lane information, speed, weight, and vehicle type when vehicles pass through the detection area without stopping. They offer advantages such as accurate load data, continuous data collection, and no disruption to traffic flow. Video vision systems can capture bridge scene videos using high-definition cameras and output spatial behavior information in real time, including the vehicle's bounding box position, center point coordinates, trajectory, and detection confidence level. They offer advantages such as high spatial visualization, rich analytical dimensions, and easy scalability. By fusing WIM data with visual data, load information and spatial behavior information can be complemented, providing a more comprehensive and reliable data foundation for vehicle identification, heavy vehicle monitoring, safe distance analysis, and risk warning in bridge scenarios.

[0004] Furthermore, Bridge Weigh-In-Motion (BWIM) is another existing technological approach in the field of bridge vehicle load identification. This type of method typically uses bridge structural response, vehicle passage information, and corresponding identification models to infer vehicle load, enabling bridge load monitoring without entirely relying on traditional road surface weighing sensors. However, the key technical focus of BWIM is load or overload identification based on bridge structural response, and its application is still often affected by factors such as axle position, vehicle speed, bridge response characteristics, structural models, or simultaneous action by multiple vehicles. Therefore, BWIM cannot directly solve the problems of establishing vehicle-level correspondences between WIM records and video visual trajectories, or evaluating the reliability of these correspondences.

[0005] However, in practical engineering applications, there is no inherent one-to-one correspondence between WIM records and visual observations. WIM systems typically output recorded information such as vehicle passage time, lane, speed, weight, and vehicle type, while video vision systems output observational information such as image coordinates, detection boxes, trajectories, and confidence levels. These two types of data differ in sampling mechanisms, sampling frequencies, triggering methods, and spatial representation. Therefore, how to reliably correlate a WIM record with its corresponding visual vehicle trajectory and quantify this correlation becomes a key issue for the practical application of multi-source data fusion in bridge traffic monitoring.

[0006] Specifically, existing WIM recording and visual observation matching technologies mainly have the following five problems: 1. Poor time synchronization: The triggering time of the WIM system and the acquisition time of the video frame are usually not synchronized, and there is a certain time deviation. If only the same time alignment method is used for matching, it is very easy to cause mismatch. 2. Multiple candidate confusion: In the same lane and within the same time window, there may be multiple vehicles traveling at the same time. If matching is based solely on spatial nearest neighbor or a single threshold rule, the target WIM record may be incorrectly associated with neighboring vehicles, reducing matching accuracy. 3. Spatial deviation: Due to the installation angle, bridge monitoring images generally have obvious perspective distortion. The positional relationship of different lanes in the image plane is not completely consistent with the actual spatial relationship of the bridge surface. Directly comparing space in the original image coordinate system will significantly reduce the matching reliability. 4. Observation quality interference: Visual detection results are affected by factors such as occlusion and changes in lighting, resulting in problems such as differences in confidence and unstable trajectories. If observation quality constraints are not introduced during the matching process, the impact of low-quality detection results on the final association results will be amplified. 5. Lack of quantitative evaluation of fusion quality: Existing technical solutions mostly focus on the overall architecture design or result application of "joint use of vision system and WIM system", but lack a unified quantitative evaluation mechanism for why a certain WIM record can match a certain visual trajectory, how reliable the matching is, and what the reasons are for non-matching or mismatching. For key links such as candidate matching set construction, component cost calculation, threshold determination and global optimization allocation, there is also a lack of traceable and explainable methods to support them.

[0007] In summary, existing technologies, relying solely on simple time alignment, spatial nearest neighbor, or empirical threshold rules, are insufficient to meet the requirements of accuracy, stability, and interpretability of visual observation-WIM data matching in bridge traffic monitoring scenarios. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization. This method can realize vehicle-level association between visual observation data and WIM data, and quantitatively evaluate the matching process and matching results. It solves the problems of low matching accuracy, poor stability, difficulty in evaluating fusion reliability, and insufficient interpretability of the matching process in the existing technology.

[0009] The objective of this invention can be achieved through the following technical solution: a vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization, comprising the following steps: S1. Acquire video images and WIM records of the bridge scene respectively, perform vehicle detection and tracking processing on the video images, and obtain a visual observation set; Extract WIM information from WIM records and generate a reference spatiotemporal location corresponding to the target WIM record; S2. Based on the reference spatiotemporal location recorded by the target WIM, within a preset time window, visual candidate targets are selected from the visual observation set to construct a candidate matching set. S3. Calculate the total matching cost of each visual candidate target in the candidate matching set, and take the visual candidate target with the minimum total matching cost as the optimal matching result. The total matching cost is composed of different sub-costs through weighted combination. The sub-costs include two or more of the following: time difference cost, lateral spatial deviation cost, longitudinal spatial deviation cost, lane consistency penalty, and visual confidence penalty. S4. Compare the total matching cost of the optimal matching result with the preset matching threshold to determine whether the matching is successful. If the matching is successful, establish the correspondence between the target WIM record and the optimal matching result, and output the corresponding matching explanation information and quantitative evaluation results. The matching explanation information and quantitative evaluation results include the number of candidate matches, the total matching cost, the cost of each item, the matching acceptance threshold, the matching solution method identifier, and the judgment result of successful matching. Otherwise, output the conclusion message indicating that no match was found.

[0010] Further, S1 includes the following steps: S11. Acquire video images of the bridge scene, and map the target vehicle in the bridge scene to a unified reference plane through homography perspective correction to obtain the corrected image. Vehicle detection and tracking are performed on the corrected image to generate a visual observation set. The visual observation set includes multiple observation targets and their corresponding vehicle detection box coordinates, vehicle center point coordinates, vehicle detection confidence, vehicle trajectory identification information, and the timestamp corresponding to the current observation frame. S12. Read the vehicle passage records output by the WIM system and extract the WIM information corresponding to each WIM record. Then, combine the preset reference lateral position of the WIM sensor in the corrected image and the center longitudinal position of the target lane in the corrected image to generate the reference spatiotemporal position of the WIM record in the visual plane.

[0011] Furthermore, the WIM information includes vehicle passage time, lane number, driving speed, vehicle weight, and vehicle type information; The reference spatiotemporal position is determined by the lane number corresponding to the WIM record, the reference lateral position of the WIM sensor in the corrected image, the center longitudinal position of the target lane in the corrected image, and the timestamp corresponding to the WIM record.

[0012] Furthermore, S2 specifically involves caching the set of visual observations obtained within a set period within a preset time window for delayed matching with the target WIM record, thereby improving the matching stability in scenarios where time is not synchronized.

[0013] Furthermore, S2 specifically involves selecting visual candidate targets that meet preset constraints from the set of visual observations. These preset constraints include time difference constraints, lateral spatial deviation constraints, and longitudinal spatial deviation constraints.

[0014] Furthermore, the time difference constraint means that the time difference between visual observation and target WIM recording is not greater than a preset time threshold; The lateral spatial deviation constraint means that the lateral deviation between the visual observation center point and the reference lateral position is not greater than the preset lateral tolerance. The longitudinal spatial deviation constraint means that the longitudinal deviation between the visual observation center point and the longitudinal position of the target lane center should not be greater than the preset longitudinal tolerance.

[0015] Furthermore, the formula for calculating the total matching cost in S3 is as follows: in, The total matching cost ranges from 0 to 1. The smaller the value, the higher the matching degree between the visual candidate target and the WIM record. The time difference normalization cost is obtained by normalizing the difference between the timestamp of the candidate visual observation and the vehicle passage time recorded by WIM. The smaller the difference, the higher the cost. The smaller the value; The cost of lateral bias normalization is obtained by normalizing the deviation between the candidate visual observation center point and the reference lateral position. The smaller the deviation, the higher the cost. The smaller the value; The cost of longitudinal deviation normalization is obtained by normalizing the deviation between the candidate visual observation center point and the longitudinal position of the target lane center. The smaller the deviation, the higher the cost. The smaller the value; As a lane consistency penalty, if the lane of the candidate visual observation matches the lane corresponding to the WIM record, then... =0; if inconsistent, then =1; The confidence penalty cost is negatively correlated with the confidence level of the visual detection results; the higher the confidence level, the lower the confidence level. The smaller the value; , , , , These correspond to the weight coefficients of each cost item, with values ​​ranging from 0 to 1, and the sum of all weight coefficients is 1.

[0016] Furthermore, in step S4, if the total matching cost of the optimal matching result is less than a preset matching threshold, then the matching is considered successful. If the total cost of the optimal matching result is greater than or equal to the preset matching threshold, then the matching is considered unsuccessful.

[0017] Furthermore, in S4, if there are multiple WIM records and multiple visual observations within the same preset time window, a cost matrix is ​​constructed between the multiple WIM records and multiple visual candidate targets, and the Hungarian algorithm is used for global optimization allocation to obtain a one-to-one matching relationship, so as to avoid mismatch problems caused by local optima.

[0018] Compared with the prior art, the present invention has the following advantages: This invention first performs vehicle detection and tracking processing on visual images to obtain a visual observation set. WIM information is extracted from WIM records, and a reference spatiotemporal position corresponding to the target WIM record is generated. Then, using the reference spatiotemporal position of the target WIM record as a benchmark, visual candidate targets are selected from the visual observation set within a preset time window to construct a candidate matching set. Finally, the total matching cost of each visual candidate target in the candidate matching set is calculated, and the visual candidate target with the lowest total matching cost is taken as the optimal matching result. The total matching cost of the optimal matching result is compared with a preset matching threshold to determine whether a match is successful. If a match is successful, a correspondence is established between the target WIM record and the optimal matching result, and corresponding matching explanation information and quantitative evaluation results are output. This achieves accurate matching and association between visual data and WIM data, effectively solving the problems of low matching accuracy, poor stability, and insufficient interpretability in existing technologies.

[0019] This invention performs homography perspective correction processing on video images in bridge scenes, which can eliminate the impact of perspective distortion on matching accuracy. Then, vehicle detection and tracking processing is performed on the corrected image to generate a visual observation set including multiple observation targets and their corresponding vehicle detection box coordinates, vehicle center point coordinates, vehicle detection confidence, vehicle trajectory identification information, and the timestamp corresponding to the current observation frame, thereby ensuring the reliability of the matching between subsequent WIM data and visual observations.

[0020] This invention extracts the WIM information (including vehicle passage time, lane number, speed, vehicle weight, and vehicle type information) corresponding to each record from the vehicle passage records output by the WIM system. Then, it combines the preset reference lateral position of the WIM sensor in the corrected image and the center longitudinal position of the target lane in the corrected image to generate the reference spatiotemporal position of the WIM record in the visual plane. This reference spatiotemporal position is used to provide an accurate benchmark for subsequent screening of visual candidate targets.

[0021] This invention selects visual candidate targets from a set of visual observations that meet preset constraints (including time difference constraints, lateral spatial deviation constraints, and longitudinal spatial deviation constraints) to construct a candidate matching set. The time difference constraint ensures the temporal correlation between candidate visual observations and target WIM records; the lateral spatial deviation constraint ensures the consistency between the candidate visual observation center point and the reference lateral position in lateral space; and the longitudinal spatial deviation constraint ensures the correlation between the candidate visual observation center point and the target lane center longitudinal position in longitudinal space. This results in a candidate matching set that satisfies spatiotemporal constraints, effectively filtering out irrelevant vehicles, reducing the risk of mismatches, and ensuring the effectiveness of the candidate matching set. Furthermore, this invention designs a method to cache visual observations collected within a recently set period in a sliding time window for delayed matching with the current WIM record, thereby improving matching stability in time-asynchronous scenarios and solving the problem of asynchronous WIM triggering time and video frame acquisition time.

[0022] This invention calculates the total matching cost between each visual candidate target and the target WIM record for each visual candidate target in the candidate matching set. The total matching cost is composed of a weighted combination of different sub-costs. The sub-costs include two or more of the following: time difference cost, lateral spatial deviation cost, longitudinal spatial deviation cost, lane consistency penalty, and visual confidence penalty. This ensures the comprehensiveness and accuracy of the matching results.

[0023] This invention selects the visual candidate target with the minimum total matching cost as the optimal matching result for the current target WIM record, and compares it with a preset matching threshold. Only when the total matching cost of the optimal matching result is less than the preset matching threshold is the correspondence between the WIM record and the visual observation formally established, and the corresponding matching interpretation information is output, including the number of candidate matches, the total matching cost, the cost of each component, the matching acceptance threshold, and the matching solution method identifier. This not only achieves optimal matching, but also makes the matching process clearly interpretable.

[0024] This invention transforms the fusion relationship between visual observation data and WIM records from a simple empirical correspondence into a quantifiable, traceable, and verifiable matching evaluation result by outputting the number of candidate matches, the total matching cost, the cost of each component, the matching acceptance threshold, and the matching solution method identifier. This facilitates subsequent mismatch analysis, parameter adjustment, and engineering debugging.

[0025] This invention considers the situation where there are multiple WIM records and multiple visual observations within a preset time window. The design first constructs a cost matrix between multiple WIM records and multiple visual candidate targets, and then uses global optimization allocation algorithms such as the Hungarian algorithm to solve the one-to-one matching relationship. This can avoid mismatches caused by local optima and improve the matching accuracy in scenarios with multiple vehicles and multiple WIM records. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 A schematic diagram of the system structure built for this embodiment; Figure 3 This is a schematic diagram illustrating the application process of an example. Figure 4 The image shown is of the bridge scene before homography correction in this embodiment. Figure 5 The image shown is of the bridge scene after homography correction in the example. Figure 6 This is a schematic diagram illustrating the selection of visual candidate targets within a preset time window based on the target WIM record in the embodiment. Figure 7 This is a schematic diagram of a heavy vehicle safety distance warning based on WIM-visual matching results in the embodiment. Detailed Implementation

[0027] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0028] Example To address the technical challenges of temporal and spatial discrepancies, lane conflicts, confusion among multiple candidate vehicles, and uncertainties in visual detection between visual observations and WIM records in bridge traffic monitoring scenarios, this solution designs a complete matching process: acquiring visual observation data and WIM records, generating reference spatiotemporal locations, constructing a candidate matching set that satisfies spatiotemporal constraints, calculating the total matching cost weighted by multiple cost terms, selecting the optimal matching result, and outputting interpretability information. This approach aims to solve the following key problems: How to generate a scientific and reasonable target reference spatiotemporal position from the visual observation set based on lane and time information recorded in WIM, so as to provide an accurate benchmark for subsequent matching; How to construct a candidate matching set under preset time windows and spatial tolerance constraints, effectively filter out irrelevant vehicles, and reduce the risk of false matching; How to comprehensively consider factors such as time difference, lateral spatial deviation, longitudinal spatial deviation, lane consistency, and visual observation confidence to construct a scientific total matching cost function, and clarify the definition and value rules of each parameter; How to select the optimal matching result from multiple visual candidate targets and effectively control the false matching rate through the total cost threshold mechanism; How to output the total cost, cost of each component, comparison results of matching thresholds, and identification of matching solution methods for the matching results, so that the fusion results between each WIM record and visual observation have quantitative evaluation basis and interpretability, and provide a reliable foundation for subsequent mismatch analysis, parameter optimization, safety distance assessment, and risk warning.

[0029] To address this, this proposal suggests a vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization, such as... Figure 1 As shown, it includes the following steps: S1. Acquire video images and WIM records of the bridge scene respectively, perform vehicle detection and tracking processing on the video images, and obtain a visual observation set; Extract WIM information from WIM records and generate a reference spatiotemporal location corresponding to the target WIM record; S2. Based on the reference spatiotemporal location recorded by the target WIM, within a preset time window, visual candidate targets are selected from the visual observation set to construct a candidate matching set. S3. Calculate the total matching cost of each visual candidate target in the candidate matching set, and take the visual candidate target with the minimum total matching cost as the optimal matching result. The total matching cost is composed of different sub-costs through weighted combination. The sub-costs include two or more of the following: time difference cost, lateral spatial deviation cost, longitudinal spatial deviation cost, lane consistency penalty, and visual confidence penalty. S4. Compare the total matching cost of the optimal matching result with the preset matching threshold to determine whether the matching is successful. If the matching is successful, establish the correspondence between the target WIM record and the optimal matching result, and output the corresponding matching explanation information and quantitative evaluation results. The matching explanation information and quantitative evaluation results include the number of candidate matches, the total matching cost, the cost of each item, the matching acceptance threshold, the matching solution method identifier, and the judgment result of successful matching. Otherwise, output the conclusion message indicating that no match was found.

[0030] This embodiment applies the above solution, firstly building as follows: Figure 2 The system architecture shown includes a video acquisition device, a vision processing module, a WIM sensor, a WIM data processing module, and a multi-source data matching module. The video acquisition device is used to acquire video images of the bridge scene. The vision processing module performs homography perspective correction processing and vehicle detection and tracking processing to obtain visual observation data of vehicles in the bridge scene. WIM sensors are used to collect dynamic WIM data on the bridge. The WIM data processing module extracts the WIM record information. In this embodiment, WIM sensors are installed under each lane of the bridge to collect vehicle passage records in real time. For each WIM record, the WIM data processing module extracts the vehicle passage time, lane number, driving speed and vehicle weight information. Then, a multi-source data matching module is used to fuse visual observation data with WIM records. This multi-source data matching module includes: The reference spatiotemporal position generation unit is used to combine the WIM record, the WIM sensor reference position and the target lane center position to generate a reference spatiotemporal position corresponding to the target WIM record; Candidate set construction unit, used to construct candidate matching sets based on time and space constraints; The cost calculation unit is used to calculate the total matching cost and individual costs for each visual candidate target; The optimal matching unit is used to determine the optimal matching result and establish the correspondence between WIM records and visual observations.

[0031] Based on the above system architecture, the following can be achieved: Figure 3 The multi-source vehicle data matching process shown includes: I. Obtaining visual observation data in bridge scenarios First, video images or image sequences of the bridge scene are acquired. In this embodiment, multiple high-definition surveillance cameras are evenly installed above the bridge deck to acquire video image sequences in real time. To eliminate the impact of perspective distortion on matching accuracy in the original images, homography perspective correction processing can be selectively performed on the acquired images to map different lanes to a unified reference plane, such as... Figure 4 and Figure 5 As shown, by using preset calibration parameters, perspective distortion caused by the installation angle in the bridge monitoring image is eliminated, and different lanes are mapped to a unified reference plane, so that the spatial relationship in the image is consistent with the actual spatial relationship of the bridge surface. This effectively solves the problem of spatial deviation caused by perspective distortion and significantly improves the reliability and accuracy of spatial matching.

[0032] Subsequently, vehicle detection and tracking algorithms are applied to the corrected image to generate a visual observation set. In this embodiment, the YOLO vehicle detection algorithm and Kalman filter tracking algorithm are used to perform vehicle detection and tracking processing on the corrected image. To ensure the reliability of the matching between the subsequent WIM data and the visual observation, each observation target in the visual observation set must contain the following core information: vehicle detection box coordinates, vehicle center point coordinates, vehicle detection confidence (value range is 0 to 1), vehicle trajectory identification information, and the timestamp corresponding to the current observation frame.

[0033] In addition, to address the issue of asynchronous WIM triggering time and video frame acquisition time, this solution caches recently acquired visual observation data in a sliding time window. The length of the sliding time window can be flexibly adjusted based on the time synchronization deviation between the WIM system and the vision system (usually set to 5-10 seconds) for subsequent delayed matching with the current WIM record. This effectively solves the problem of asynchronous WIM triggering time and video frame acquisition time, improving the matching stability and success rate in asynchronous scenarios.

[0034] This embodiment uses a sliding time window caching mechanism to cache visual observation data within the most recent 10 seconds for subsequent delay matching.

[0035] II. Obtain WIM records and generate reference spatiotemporal locations The system reads vehicle passage records output by the WIM (Wide Motion Imaging) system and extracts core information for each record, including vehicle passage time, lane number, speed, vehicle weight, and vehicle type. Then, combining this information with a pre-defined reference lateral position of the WIM sensor in the corrected image and the center longitudinal position of the target lane in the corrected image, a reference spatiotemporal position for that WIM record in the visual plane is generated. This reference spatiotemporal position is determined by the lane number corresponding to the WIM record, the reference lateral position of the WIM sensor in the corrected image, the center longitudinal position of the target lane in the corrected image, and the timestamp corresponding to the WIM record, providing a unified and accurate benchmark for subsequent candidate selection.

[0036] III. Constructing a Candidate Matching Set Using the reference spatiotemporal location generated in step two as a benchmark, the visual observation set is filtered within a preset time window to construct a candidate matching set. To ensure the effectiveness of the candidate set, the filtering process must simultaneously satisfy the following three spatiotemporal constraints: Time difference constraint: The time difference between candidate visual observation and target WIM record is not greater than a preset time threshold to ensure that the two are correlated in the time dimension; Lateral spatial deviation constraint: The lateral deviation between the candidate visual observation center point and the reference lateral position is not greater than the preset lateral tolerance, ensuring the consistency of the two in lateral space; Longitudinal spatial deviation constraint: The longitudinal deviation between the candidate visual observation center point and the longitudinal position of the target lane center is not greater than the preset longitudinal tolerance, ensuring the correlation between the two in longitudinal space.

[0037] If a WIM record does not find a visual candidate target that meets the above spatiotemporal constraints within the current time window, it is marked as unmatched.

[0038] In this embodiment, the time window is set to ±5 seconds, the horizontal tolerance is 50 pixels, and the vertical tolerance is 100 pixels, to filter visual candidate targets as follows: Figure 6 As shown.

[0039] IV. Calculating the Cost of Candidate Matching For each visual candidate target in the candidate matching set, calculate the total matching cost between each visual candidate target and the target WIM record. The total matching cost needs to take into account multiple factors and must consist of at least two of the following: time difference cost, lateral spatial deviation cost, longitudinal spatial deviation cost, lane consistency penalty, and visual observation confidence penalty, in order to ensure the comprehensiveness and accuracy of the matching results.

[0040] In this embodiment, the total matching cost is calculated using a normalized weighted summation method. Each cost item includes time difference cost, lateral deviation cost, longitudinal deviation cost, lane consistency penalty, and confidence penalty. First, each cost item is normalized to a uniform value range of 0 to 1. Then, a weighted summation is performed according to preset weights to obtain the total matching cost; a smaller cost value indicates a higher degree of matching. Specifically, smaller time difference, smaller spatial deviation, higher lane consistency, and higher detection confidence correspond to lower costs.

[0041] To eliminate the influence of differences in the dimensions of each cost item, each cost item is first normalized (mapping the values ​​of each cost item to the interval of 0 to 1), and then a weighted summation method is used to calculate the total matching cost. The specific calculation formula is as follows: The specific definitions of each parameter are as follows: : Total matching cost, ranging from 0 to 1. The smaller the value, the higher the matching degree between the visual candidate target and the WIM record. The time difference normalization cost is obtained by normalizing the difference between the timestamp of the candidate visual observation and the vehicle passage time recorded by WIM. The smaller the difference, the higher the cost. The smaller the value; The lateral bias normalization cost is obtained by normalizing the bias between the candidate visual observation center point and the reference lateral position. The smaller the bias, the higher the cost. The smaller the value; The longitudinal deviation normalization cost is obtained by normalizing the deviation between the candidate visual observation center point and the longitudinal position of the target lane center. The smaller the deviation, the higher the cost. The smaller the value; Lane consistency penalty: If the lane of the candidate visual observation is consistent with the lane corresponding to the WIM record, then... =0; if inconsistent, then =1 (The specific value of the penalty coefficient can be adjusted according to the actual project requirements). The confidence penalty cost is negatively correlated with the confidence level of the visual detection result; the higher the confidence level, the lower the confidence level. The smaller the value, the more its calculation formula can be set as follows: ,in, The confidence level for visual detection ranges from 0 to 1. , , , , The weight coefficients for each cost item range from 0 to 1, and the sum of all weight coefficients is 1. The proportion of each weight can be flexibly adjusted according to the needs of the actual bridge monitoring scenario. For example, in a heavy vehicle monitoring scenario, the weight can be increased. , The weights are adjusted to enhance the accuracy of spatial matching.

[0042] V. Determining the optimal matching relationship After calculating the total matching cost of all visual candidate targets in the candidate matching set, the visual candidate target with the minimum total matching cost is selected as the optimal matching result of the current target WIM record.

[0043] Meanwhile, a preset matching threshold is set (the value ranges from 0 to 1 and can be adjusted according to the actual matching accuracy requirements). When the total matching cost of the best matching result is less than the preset matching threshold, the correspondence between the WIM record and the visual observation is formally established; if the total matching cost is greater than or equal to the preset threshold, it is marked as unmatched, and a second matching or manual review can be performed later.

[0044] In this embodiment, the matching threshold is set to 0.5. In addition, considering the situation that there are multiple WIM records and multiple visual observations within the preset time window, this solution designs and constructs a cost matrix between multiple WIM records and multiple visual candidate targets. Global optimization allocation algorithms such as the Hungarian algorithm are used to solve the one-to-one matching relationship, avoiding local optimal mismatch caused by matching a single WIM record with a candidate target, so as to improve the matching accuracy in multi-vehicle and multi-WIM record scenarios.

[0045] VI. Output the matching interpretation results and perform subsequent applications. After achieving optimal matching, the matching results, corresponding explanatory information, and quantitative evaluation results are output. The explanatory information and quantitative evaluation results include at least the number of candidate matches, the total final matching cost, the cost of each component, the matching acceptance threshold, the matching solution method identifier, and the judgment result of successful or unsuccessful matching. This ensures the matching process has clear interpretability and verifiability, facilitating subsequent mismatch analysis, parameter optimization, and engineering debugging. This embodiment, based on this, such as... Figure 7 As shown, the vehicle speed and weight information recorded in WIM are further utilized to calculate the dynamic safe distance of the vehicle (the faster the vehicle speed and the greater the weight, the longer the dynamic safe distance). At the same time, the position of the vehicle in front in the same lane is determined by visual observation, and the image pixel distance is converted into the actual road surface distance. If the actual following distance is less than the dynamic safe distance, it is determined that there is a dangerous following behavior, and a warning signal is output in time. This can give full play to the value of multi-source data fusion and realize the safety warning of heavy vehicles in bridge scenarios.

[0046] In summary, applying this solution to bridge traffic monitoring technology enables stable matching between visual observation data and WIM records in bridge scenarios. The matching accuracy is significantly improved compared to traditional time alignment or spatial nearest neighbor methods, and it maintains good stability even in multi-vehicle scenarios. Furthermore, by outputting matching interpretation information, the matching process possesses good interpretability, effectively supporting subsequent applications such as bridge heavy vehicle safety early warning.

[0047] This solution has the following significant advantages: High matching accuracy: By accurately generating reference spatiotemporal locations and strictly screening candidate matching sets, irrelevant vehicles are effectively filtered out, reducing the risk of mismatches from the source and significantly improving matching accuracy; High stability: The total matching cost is calculated by weighted summation of multiple cost terms, replacing the traditional simple nearest neighbor rule. It can comprehensively consider multiple factors such as time, space, lane, and observation quality, thereby improving the matching stability in multi-vehicle scenarios within the same time window. Good interpretability: By explicitly outputting the total matching cost and the cost information of each component, the definition and value rules of each parameter are clearly defined, making the matching process clear and traceable, which facilitates subsequent mismatch analysis, parameter optimization and engineering implementation debugging; High versatility: Through optimized technical solutions, it can be smoothly extended to the global optimal allocation between multiple WIM records and multiple visual trajectories, adapting to monitoring scenarios of different bridge types and different traffic flows, and has broad engineering application prospects; High engineering value: It provides a more reliable data association foundation for bridge heavy vehicle safety monitoring, risk early warning and digital twin display, realizes the effective integration of visual spatial information and WIM load information, and significantly improves the intelligence level and safety management capabilities of bridge traffic monitoring.

[0048] Furthermore, in bridge scenarios where both WIM data and visual data are available, the high-confidence matching results generated by this solution can serve as corresponding samples between vehicle visual features and load information, providing a data foundation for subsequent vehicle load level estimation, heavy vehicle risk identification, or auxiliary analysis model training under visual-only conditions. However, the above extended applications depend on specific training samples, bridge scenario conditions, and model validation results. This invention does not presuppose that visual data alone can directly replace WIM weighing results.

Claims

1. A method for matching multi-source vehicle data based on spatiotemporal constraints and cost function optimization, characterized in that, Includes the following steps: S1. Acquire video images and WIM records of the bridge scene respectively, perform vehicle detection and tracking processing on the video images, and obtain a visual observation set; Extract WIM information from WIM records and generate a reference spatiotemporal location corresponding to the target WIM record; S2. Based on the reference spatiotemporal location recorded by the target WIM, within a preset time window, visual candidate targets are selected from the visual observation set to construct a candidate matching set. S3. Calculate the total matching cost of each visual candidate target in the candidate matching set, and take the visual candidate target with the minimum total matching cost as the optimal matching result. The total matching cost is composed of different sub-costs through weighted combination. The sub-costs include two or more of the following: time difference cost, lateral spatial deviation cost, longitudinal spatial deviation cost, lane consistency penalty, and visual confidence penalty. S4. Compare the total matching cost of the optimal matching result with the preset matching threshold to determine whether the matching is successful. If the matching is successful, establish the correspondence between the target WIM record and the optimal matching result, and output the corresponding matching explanation information and quantitative evaluation results. The matching explanation information and quantitative evaluation results include the number of candidate matches, the total matching cost, the cost of each item, the matching acceptance threshold, the matching solution method identifier, and the judgment result of successful matching. Otherwise, output the conclusion message indicating that no match was found; S1 includes the following steps: S11. Acquire video images of the bridge scene, and map the target vehicle in the bridge scene to a unified reference plane through homography perspective correction to obtain the corrected image. Vehicle detection and tracking are performed on the corrected image to generate a visual observation set. The visual observation set includes multiple observation targets and their corresponding vehicle detection box coordinates, vehicle center point coordinates, vehicle detection confidence, vehicle trajectory identification information, and the timestamp corresponding to the current observation frame. S12. Read the vehicle passage records output by the WIM system and extract the WIM information corresponding to each WIM record. Then, combine the preset reference lateral position of the WIM sensor in the corrected image and the center longitudinal position of the target lane in the corrected image to generate the reference spatiotemporal position of the WIM record in the visual plane.

2. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 1, characterized in that, The WIM information includes vehicle passage time, lane number, speed, vehicle weight, and vehicle type information.

3. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 2, characterized in that, The reference spatiotemporal position is determined by the lane number corresponding to the WIM record, the reference lateral position of the WIM sensor in the corrected image, the center longitudinal position of the target lane in the corrected image, and the timestamp corresponding to the WIM record.

4. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 1, characterized in that, Specifically, S2 involves caching the set of visual observations obtained within a set period within a preset time window for delayed matching with the target WIM record, thereby improving the matching stability in scenarios where time is not synchronized.

5. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 1, characterized in that, Specifically, S2 involves selecting visual candidate targets from the set of visual observations that meet preset constraints, including time difference constraints, lateral spatial deviation constraints, and longitudinal spatial deviation constraints.

6. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 5, characterized in that, The time difference constraint means that the time difference between visual observation and target WIM recording is not greater than a preset time threshold. The lateral spatial deviation constraint means that the lateral deviation between the visual observation center point and the reference lateral position is not greater than the preset lateral tolerance. The longitudinal spatial deviation constraint means that the longitudinal deviation between the visual observation center point and the longitudinal position of the target lane center should not be greater than the preset longitudinal tolerance.

7. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 1, characterized in that, The formula for calculating the total matching cost in S3 is as follows: ; in, The total matching cost ranges from 0 to 1. The smaller the value, the higher the matching degree between the visual candidate target and the WIM record. The time difference normalization cost is obtained by normalizing the difference between the timestamp of the candidate visual observation and the vehicle passage time recorded by WIM. The smaller the difference, the higher the cost. The smaller the value; The cost of lateral bias normalization is obtained by normalizing the deviation between the candidate visual observation center point and the reference lateral position. The smaller the deviation, the higher the cost. The smaller the value; The cost of longitudinal deviation normalization is obtained by normalizing the deviation between the candidate visual observation center point and the longitudinal position of the target lane center. The smaller the deviation, the higher the cost. The smaller the value; As a lane consistency penalty, if the lane of the candidate visual observation matches the lane corresponding to the WIM record, then... =0; if inconsistent, then =1; The confidence penalty cost is negatively correlated with the confidence level of the visual detection results; the higher the confidence level, the lower the confidence level. The smaller the value; , , , , These correspond to the weight coefficients of each cost item, with values ​​ranging from 0 to 1, and the sum of all weight coefficients is 1.

8. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 1, characterized in that, In step S4, if the total cost of the optimal matching result is less than the preset matching threshold, then the matching is considered successful. If the total cost of the optimal matching result is greater than or equal to the preset matching threshold, then the matching is considered unsuccessful.

9. The vehicle multi-source data matching method based on spatiotemporal constraints and cost function optimization according to claim 1, characterized in that, In step S4, if there are multiple WIM records and multiple visual observations within the same preset time window, a cost matrix is ​​constructed between the multiple WIM records and multiple visual candidate targets, and the Hungarian algorithm is used for global optimization allocation to obtain a one-to-one matching relationship, so as to avoid mismatch problems caused by local optima.

Citation Information

Patent Citations

  • Camera-based high-performance collaborative awareness and distributed fusion positioning method

    CN117974807A

  • Multi-target tracking method and system for multi-view scene

    CN121330240A