This invention relates to the fields of
computer vision and multimodal
information processing technology, providing a vehicle matching method and apparatus based on cross-
modal multi-view perspectives. In this invention, visible light and
infrared images are first acquired.
Vehicle detection is performed on the visible light image to identify the target vehicle. Then, vehicles in the visible light and
infrared images are subjected to BEV
perspective transformation using a pre-calibrated
homography matrix to obtain normalized coordinates in the bird's-eye view space. Next, the left-right relationship between adjacent vehicles is determined based on the normalized coordinates, constructing left-right relationship sequences for vehicles in the visible light and
infrared images respectively. Then, an
infrared image frame that temporally matches the target vehicle is searched in the
infrared image sequence based on a
temporal search strategy. Finally, for each searched
infrared image frame, the lane number is calculated and lane consistency filtering is performed. A weighted
sequence matching method is then used to calculate the matching
score, and the candidate vehicle with the highest matching
score is selected as the target matching vehicle. This invention does not rely on visual features for matching, effectively eliminating the influence of multi-view
perspective distortion. It also solves the problem of low vehicle matching accuracy in visible light camera images and
infrared thermal imaging camera images, which prevents the detection of temperature anomalies. It is suitable for cross-
modal vehicle matching scenarios such as highway checkpoints and urban traffic monitoring.