A vehicle matching method and device based on cross-modal multi-view

CN122867078APending Publication Date: 2026-10-02ZHEJIANG JIAOTONG EXPRESSWAY OPERATION & MANAGEMENT CO LTD LISHUI MANAGEMENT OFFICE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610991996.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请提供一种基于跨模态多视角的车辆匹配方法和装置,以解决可见光相机图像和红外热成像相机图像中车辆匹配准确率太低,无法完成温度异常检测的问题

Benefits of technology

多视角统一适配:通过BEV透视变换将不同倾斜角度下的车辆像素坐标统一到鸟瞰图坐标系,消除了不同安装位置和不同倾斜角度带来的透视畸变影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122867078A_ABST
    Figure CN122867078A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of computer vision and multimodal information processing technology, providing a vehicle matching method and apparatus based on cross-modal multi-view perspectives. In this invention, visible light and infrared images are first acquired. Vehicle detection is performed on the visible light image to identify the target vehicle. Then, vehicles in the visible light and infrared images are subjected to BEV perspective transformation using a pre-calibrated homography matrix to obtain normalized coordinates in the bird's-eye view space. Next, the left-right relationship between adjacent vehicles is determined based on the normalized coordinates, constructing left-right relationship sequences for vehicles in the visible light and infrared images respectively. Then, an infrared image frame that temporally matches the target vehicle is searched in the infrared image sequence based on a temporal search strategy. Finally, for each searched infrared image frame, the lane number is calculated and lane consistency filtering is performed. A weighted sequence matching method is then used to calculate the matching score, and the candidate vehicle with the highest matching score is selected as the target matching vehicle. This invention does not rely on visual features for matching, effectively eliminating the influence of multi-view perspective distortion. It also solves the problem of low vehicle matching accuracy in visible light camera images and infrared thermal imaging camera images, which prevents the detection of temperature anomalies. It is suitable for cross-modal vehicle matching scenarios such as highway checkpoints and urban traffic monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision and multimodal information processing technology, and in particular to a vehicle matching method and apparatus based on cross-modal multi-view. Background Technology

[0002] In scenarios such as highway checkpoints and urban traffic monitoring, it is often necessary to correlate image information captured by visible light cameras with temperature information captured by infrared thermal imaging cameras to achieve vehicle identification and temperature detection fusion. However, due to differences in the installation location, viewing angle, and imaging modality of different cameras, cross-camera vehicle target matching faces many technical challenges.

[0003] First, the issue of perspective distortion differences from multiple viewpoints is prominent. Although all cameras are mounted on the gantry, their observation angles differ. In images taken from a front-viewing camera, vehicles appear as top projections, while in images taken from a tilted-viewing camera, vehicles appear as a mixture of top and side projections. Furthermore, the farther the camera is from the center of the road, the greater the tilt angle, and the more severe the perspective distortion. The pixel coordinates of vehicles at different tilt angles do not exhibit a simple linear mapping relationship and cannot be directly correlated through a unified coordinate transformation.

[0004] Secondly, the visual features differ significantly. Visible light images contain RGB color and texture information, while infrared images reflect temperature distribution and grayscale information; the visual feature representations of the two modalities are completely different. Traditional matching methods based on local feature descriptors such as SIFT and SURF fail across different modalities, while deep learning-based cross-modal feature extraction methods require a large amount of labeled data and have limited generalization ability. This results in extremely low vehicle matching accuracy when using visible light cameras and infrared thermal imaging cameras to detect temperature anomalies, making it impossible to complete the detection process properly. Summary of the Invention

[0005] In view of this, this application provides a vehicle matching method and apparatus based on cross-modal multi-view to solve the problem that the accuracy of vehicle matching in visible light camera images and infrared thermal imaging camera images is too low, making it impossible to complete temperature anomaly detection.

[0006] The first aspect of this application provides a vehicle matching method based on cross-modal multi-view, the method comprising: Acquire visible light and infrared images, perform vehicle detection on the visible light images, and identify the target vehicle; The vehicles in the visible light image and infrared image are respectively subjected to BEV perspective transformation through a pre-calibrated homography matrix to obtain the normalized coordinates of each vehicle in the bird's-eye view space. The left-right relationship between adjacent vehicles is determined based on the normalized coordinates, and the left-right relationship sequence of each vehicle in the visible light and infrared images is constructed respectively. Based on a temporal search strategy, infrared image frames that match the visible light target vehicle in time are searched in an infrared image sequence, wherein the infrared image sequence is a set of multiple infrared images acquired by an infrared camera in chronological order. For each frame of infrared image searched, the lane number of the target vehicle in the visible light image and each candidate vehicle in the infrared image is calculated and lane consistency filtering is performed. Then, the matching score between each candidate vehicle in the infrared image and the target vehicle is calculated based on the left-right relationship sequence using a weighted sequence matching method, and the candidate vehicle with the highest matching score is selected as the target matching vehicle.

[0007] Optionally, the method further includes: When multiple infrared cameras are present, for the target matching vehicle identified by each infrared camera, the one with the smallest matching difference is selected as the final matching result based on a multi-level sorting strategy.

[0008] Optionally, determining the target vehicle includes: The vehicle identifier of each vehicle is determined, and the target vehicle is determined based on the vehicle identifier, wherein the vehicle identifier is the license plate number or the vehicle ID matched in the previous frame.

[0009] Optionally, calculating lane numbers and performing lane consistency filtering includes: For a target vehicle in the visible light image, the lane to which it belongs is determined based on the horizontal coordinate of the center of its detection frame and the pre-calibrated horizontal coordinate range of each lane. For candidate vehicles in the infrared image, the lane to which they belong is determined based on their lateral normalized coordinates in the BEV space using a pre-calibrated lane boundary threshold. Candidate vehicles with the same lane number as the target vehicle are retained for subsequent matching.

[0010] Optionally, determining the left-right relationship of adjacent vehicles based on normalized BEV coordinates includes: For infrared images, a vehicle part detection model is used to identify vehicle parts including the hood, cab, cargo box, and tires, and parts of the same vehicle are aggregated by IoU intersection judgment. Based on different combinations of parts, a differentiated comparison strategy is used to determine the left-right relationship of adjacent vehicles.

[0011] Optionally, after calculating the matching score between each candidate vehicle and the target vehicle in the infrared image, the method further includes: When the highest matching score is lower than the lowest matching score threshold, it is determined whether there are parallel vehicles around the target vehicle. If so, the target vehicle is swapped with the parallel vehicle, and the step of calculating the matching score by weighted sequence matching is performed again based on the swapped vehicle.

[0012] Optionally, the time-series search strategy includes: The search is performed based on the timestamp status of the visible light image. When the camera timestamp on the visible light image is less than the server timestamp, it is determined to be a normal image, and N frames of infrared images are searched forward in time. When the camera timestamp is greater than or equal to the server timestamp, it is determined to be an abnormal image. The forward or backward search is dynamically determined based on the time difference with the previous visible light image and the number of vehicles.

[0013] A second aspect of this application provides a vehicle matching device based on cross-modal multi-view, the device comprising: The target vehicle determination unit is used to acquire visible light images and infrared images, perform vehicle detection on the visible light images, and determine the target vehicle. The BEV transformation unit is used to perform BEV perspective transformation on the vehicles in the visible light image and infrared image respectively through a pre-calibrated homography matrix to obtain the normalized coordinates of each vehicle in the bird's-eye view space. The left-right relationship sequence determination unit is used to determine the left-right relationship between adjacent vehicles based on the normalized coordinates, and to construct the left-right relationship sequence of each vehicle in the visible light and infrared images respectively. An infrared image frame matching unit is used to search for infrared image frames that match the visible light target vehicle in time in an infrared image sequence based on a time-series search strategy. The infrared image sequence is a set of multiple infrared images acquired by an infrared camera in chronological order. The vehicle matching unit is used to calculate the lane number of the target vehicle in the visible light image and each candidate vehicle in the infrared image for each frame of infrared image searched, and to perform lane consistency filtering. Then, it uses a weighted sequence matching method to calculate the matching score between each candidate vehicle in the infrared image and the target vehicle based on the left-right relationship sequence, and selects the candidate vehicle with the highest matching score as the target matching vehicle.

[0014] Optionally, the device further includes: The multi-camera matching unit is used to select the vehicle with the smallest matching difference as the final matching vehicle based on a multi-level sorting strategy when multiple infrared cameras are present.

[0015] Optionally, the device further includes: The vehicle swapping unit is used to calculate the matching score between each candidate vehicle and the target vehicle in the infrared image. When the highest matching score is lower than the lowest matching score threshold, it determines whether there are parallel vehicles around the target vehicle. When it is determined, the target vehicle is swapped with the parallel vehicles, and the step of calculating the matching score by weighted sequence matching method is performed again based on the swapped vehicles.

[0016] Compared with the prior art, the present invention has the following beneficial effects: Multi-view unified adaptation: By using BEV perspective transformation, the vehicle pixel coordinates under different tilt angles are unified to the bird's-eye view coordinate system, eliminating the perspective distortion caused by different installation positions and different tilt angles.

[0017] Cross-modal robustness: This invention does not rely on visual features for matching, but uses the relative positional relationship between vehicles as the matching basis, making it naturally applicable to matching between different modalities such as visible light and infrared.

[0018] Time deviation robustness: The left-right positional relationship between vehicles remains stable within a short time window. Even if there is a certain time synchronization deviation between the two cameras, the left-right relationship sequence can still remain consistent.

[0019] Adaptation to multi-vehicle side-by-side scenarios: By using a strategy for determining and swapping side-by-side vehicles, the ambiguity issues arising from relying solely on longitudinal position matching are effectively resolved.

[0020] Local detection tolerance: Using a weighted average matching score, failure or omission of some vehicles in the detection does not affect the overall matching result. Attached Figure Description

[0021] Figure 1 A flowchart illustrating the method provided in this application embodiment; Figure 2 This is a schematic diagram of the BEV perspective transformation effect provided in the embodiments of this application; Figure 3 This is a schematic diagram illustrating the construction of the left-right relationship sequence provided in an embodiment of this application; Figure 4 This is a structural diagram of the device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0025] This application provides a vehicle matching method and apparatus based on cross-modal multi-view to solve the problem that the accuracy of vehicle matching in visible light camera images and infrared thermal imaging camera images is too low, making it impossible to complete temperature anomaly detection.

[0026] The technical solutions of this application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0027] like Figure 1 The diagram shown is a flowchart of a vehicle matching method based on cross-modal multi-view provided in this application. The process may include the following steps: Step S101: Acquire visible light image and infrared image, perform vehicle detection on the visible light image, and determine the target vehicle.

[0028] In this embodiment, images collected by a visible light camera and an infrared camera are received in real time first. Inference is performed on the visible light image using the YOLO vehicle detection model to obtain all vehicle detection boxes. Meanwhile, the target vehicle in the visible light image can be determined according to an externally input target vehicle identifier, for example, a license plate number provided by an upstream ETC system, a vehicle position selected by a user clicking through a graphical user interface, or a vehicle ID transmitted from a previous frame matching result.

[0029] For example, if the input is a license plate number, the license plate in the current image is recognized through OCR, and the corresponding vehicle detection box is obtained through matching as the target vehicle. If the input is a click coordinate, it is determined which detection box the coordinate falls within. If the input is a vehicle ID from the previous frame, target tracking is performed using Kalman filtering or IoU matching to associate the same vehicle in the current frame.

[0030] To give a specific example, assume that the upstream ETC system provides the license plate number "Zhe A 12345" of the target vehicle. In this embodiment, all license plates in the current image are recognized through OCR, string matching is performed between the recognition result and "Zhe A 12345" to find the vehicle detection box where the corresponding license plate is located, and this detection box is determined as the target vehicle.

[0031] Step S102: subject the vehicles in the visible light image and the infrared image to BEV perspective transformation through a pre-calibrated homography matrix respectively, to obtain normalized coordinates of each vehicle in the bird's-eye view space.

[0032] In this embodiment, BEV perspective transformation is performed on the visible light image and the infrared image respectively. Each camera obtains its own homography matrix H through pre-calibration, which converts pixel points in the image mapped to the bird's-eye view coordinate system: . represents homogeneous coordinates in the bird's-eye view coordinate system; represents homogeneous coordinates in the image coordinate system. After mapping, the horizontal coordinate is normalized: , wherein: represents the normalized horizontal coordinate; represents the horizontal coordinate of the vehicle in the BEV space; represents the width of the current BEV image. Through homography matrix transformation and normalization processing, vehicles at different installation positions and different tilt angles can be uniformly mapped to the standardized bird's-eye view space, so as to reduce the influence of perspective distortion, and the effect Figure 2 is shown.

[0033] Step S103: determine the left-right relationship between adjacent vehicles based on the normalized coordinates, and construct left-right relationship sequences of each vehicle in the visible light image and the infrared image respectively.

[0034] In this embodiment, for infrared images, the YOLO vehicle part detection model is used to identify parts such as the hood, cab, passenger compartment, and tires. Parts belonging to the same vehicle are aggregated using the Intersection over Union (IoU) to obtain the vehicle's corner information. For visible light images, the coordinates of the lower left and lower right corners of the overall vehicle detection bounding box are used directly.

[0035] Then, the vehicles are arranged in descending order of their normalized Y-coordinates (from nearest to farthest), and the left-right relationships of adjacent vehicles are determined sequentially. A differentiated strategy is used based on the combination of parts in the determination: if adjacent vehicles both have a dual-part structure, the x-coordinates of their corner points are compared; if both have a single tire, the corner point combination with the smaller difference is selected; if one vehicle has a single tire, the offset is determined by comparing the distance between the tire's center point and the corner point of the other vehicle. Finally, visible light left-right relationship sequences and infrared left-right relationship sequences are constructed, with the following results: Figure 3 As shown.

[0036] For adjacent vehicles, a threshold can be preset. When the absolute value of the horizontal position difference in the BEV space does not exceed the threshold, the two adjacent vehicles are determined to be in overlapping positions.

[0037] Step S104: Based on the temporal search strategy, search for infrared image frames in the infrared image sequence that match the visible light target vehicle in time.

[0038] In this embodiment, the camera timestamp of the current visible light image is used as a reference to search for a time-matching infrared image frame in the infrared image sequence. If the visible light camera timestamp is less than the server timestamp, it is determined to be a normal image, and N frames are searched forward. For example, assuming the camera timestamp of the current visible light image is t=1000ms and the server's current time is t=1020ms, since 1000<1020, the timestamp is determined to be normal, and only 5 infrared images are searched in the direction of increasing time (forward).

[0039] If the timestamp is greater than or equal to the server timestamp, it is considered an anomaly. Based on the time difference between the current visible light image and the previous frame, as well as the number of vehicles in the previous visible light image, it is dynamically determined whether to search forward or backward first.

[0040] It should be noted that during the search process, the trend of the vehicle's Y-coordinate in the infrared image can be checked frame by frame. If the Y-coordinate of the possible candidate position of the target vehicle in three consecutive infrared images continues to increase, and the difference between the Y-coordinate of the target vehicle in the visible light image and the Y-coordinate of the target vehicle continues to widen, it indicates that the vehicle is moving away from the camera, and the search can be terminated in advance to avoid invalid calculations.

[0041] Step S105: for each searched frame of infrared image, calculate the lane numbers of the target vehicle in the visible light image and each candidate vehicle in the infrared image, perform lane consistency filtering, then calculate the matching score between each candidate vehicle in the infrared image and the target vehicle according to the left-right relationship sequence by a weighted sequence matching method, and select the candidate vehicle with the highest matching score as the target matching vehicle.

[0042] In this embodiment, lane consistency check is required for each candidate vehicle, requiring that the lane numbers of the visible light vehicle and the corresponding infrared vehicle must be the same. First, the lane number of the visible light target vehicle and the lane number of the candidate vehicle in each searched frame of infrared image are calculated. The specific process is as follows: Since both have been mapped to the BEV space, the judgment can be made by using the pre-calibrated lane boundary threshold based on the lateral normalized coordinate x in the BEV space. For example, if the lateral coordinates of the lane boundary pre-calibrated in the BEV space are b1 and b2, then when x <b1, it belongs to lane 1; when b1≤x <b2, it belongs to lane 2; when x≥b2, it belongs to lane 3.

[0043] When performing lane consistency check, to avoid interference from long-distance candidate vehicles, the candidate search range is limited as follows: , wherein: represents the index of an infrared candidate vehicle, represents the index of a visible light target vehicle. By limiting the sequence index difference, the matching efficiency can be improved and the probability of mismatching can be reduced.

[0044] Then the weighted sequence matching method is used to calculate the matching score of each candidate vehicle, and the calculation formula is as follows: , wherein: represents the -th infrared candidate vehicle's matching score; represents the -th relational term's corresponding weight; represents a relation comparison function, which is used to calculate the matching degree between the visible light relation term and the infrared relation term; represents the -th left-right relationship in the visible light sequence; represents the left-right relationship at the position offset from the candidate vehicle in the infrared vehicle sequence; represents the comparison position index in the relation sequence; m represents the number of relations compared forward; n represents the number of relations compared backward, wherein the comparison function The definition is as follows: when the left and right relationships are completely consistent, the value is 1.0; when the left and right relationships are inconsistent but the position overlap flag is true, the value is 0.5; when the left and right relationships are completely mismatched, the value is 0. For example, if the relationship between the target vehicle and the following vehicle in the visible light sequence is "left", and the corresponding relationship in the infrared sequence is also "left", then C(·) = 1.0; if both are "overlapping", then C(·) = 0.5; if they are inconsistent, then C(·) = 0.

[0045] The candidate vehicle with the highest matching score is selected as the target matching vehicle for that frame using the above formula.

[0046] This concludes the process. Figure 1 The process is shown below.

[0047] In the embodiments of this application, visible light images and infrared images are first acquired. Vehicle detection is performed on the visible light images to identify the target vehicle. Then, the vehicles in the visible light and infrared images are subjected to BEV perspective transformation through a pre-calibrated homography matrix to obtain normalized coordinates in the bird's-eye view space. Next, the left-right relationship between adjacent vehicles is determined based on the normalized coordinates, and visible light left-right relationship sequences and infrared left-right relationship sequences are constructed respectively. Then, an infrared image frame that temporally matches the target vehicle is searched in the infrared image sequence based on a temporal search strategy. Finally, for each searched infrared image frame, the lane number is calculated and lane consistency filtering is performed. A weighted sequence matching method is then used to calculate the matching score, and the candidate vehicle with the highest matching score is selected as the target matching vehicle. This invention does not rely on visual features for matching, effectively eliminates the influence of multi-view perspective distortion, and solves the problem of low vehicle matching accuracy in visible light camera images and infrared thermal imaging camera images, which makes it impossible to complete temperature anomaly detection.

[0048] In another embodiment, the method further includes: When multiple infrared cameras are present, for the target matching vehicle identified by each infrared camera, the one with the smallest matching difference is selected as the final matching result based on a multi-level sorting strategy.

[0049] In this embodiment, when multiple infrared cameras are present, each camera independently performs the above steps to obtain its own target matching vehicle. Then, a multi-level sorting strategy is used to select the final matching result. Taking a three-camera system as an example, assume the results from the three cameras are as follows:

[0050] The sorting keys are compared sequentially: if the side-by-side consistency of the left and middle cameras in the three-way cameras is consistent, proceed to the next level; if the Y coordinates all meet the preset range, proceed to the next level; select the target matching vehicle determined by the middle camera with the smallest matching difference as the final matching result.

[0051] In another embodiment, after calculating the matching score between each candidate vehicle and the target vehicle in the infrared image, the method further includes: When the highest matching score is lower than the lowest matching score threshold, it is determined whether there are parallel vehicles around the target vehicle. If so, the target vehicle is swapped with the parallel vehicle, and the step of calculating the matching score by weighted sequence matching is performed again based on the swapped vehicle.

[0052] Since there may be vehicles side by side in the target vehicle in the visible light image, and the relative positions of the side by side vehicles may be inconsistent due to differences in the frame rate or time of each camera, the highest matching score in step S105 will be 0.

[0053] Therefore, in this embodiment, a minimum matching score threshold, such as 0.1, is preset. When the highest matching score in a certain scenario is 0, it is determined whether there are parallel vehicles around the target vehicle in the visible light spectrum, that is, whether there are two vehicles in different lanes with overlapping detection boxes in the Y direction. If so, it is determined that parallel vehicles exist. At this time, a swapping strategy is executed: the identities of the target vehicle and the parallel vehicle on the right are swapped, and weighted sequence matching is re-executed. If the highest score after the swap exceeds the minimum matching score threshold, such as increasing to 0.8, the swapped matching result is accepted, and the infrared vehicle corresponding to the original parallel vehicle is taken as the target matching vehicle.

[0054] This application also provides a vehicle matching device based on cross-modal multi-view, such as... Figure 4 As shown, the device includes: The target vehicle determination unit 401 is used to acquire visible light images and infrared images, perform vehicle detection on the visible light images, and determine the target vehicle. BEV transformation unit 402 is used to perform BEV perspective transformation on the vehicles in the visible light image and infrared image respectively through a pre-calibrated homography matrix to obtain the normalized coordinates of each vehicle in the bird's-eye view space. The left-right relationship sequence determination unit 403 is used to determine the left-right relationship between adjacent vehicles based on the normalized coordinates, and to construct the left-right relationship sequence of each vehicle in the visible light and infrared images respectively. The infrared image frame matching unit 404 is used to search for infrared image frames that match the visible light target vehicle in time in the infrared image sequence based on a time-series search strategy. The vehicle matching unit 405 is used to calculate the lane number of the target vehicle in the visible light image and each candidate vehicle in the infrared image for each frame of infrared image searched, and to perform lane consistency filtering. Then, through a weighted sequence matching method, it calculates the matching score between each candidate vehicle in the infrared image and the target vehicle according to the left-right relationship sequence, and selects the candidate vehicle with the highest matching score as the target matching vehicle.

[0055] In another embodiment, the device further includes: The multi-camera matching unit is used to select the vehicle with the smallest matching difference as the final matching vehicle based on a multi-level sorting strategy when multiple infrared cameras are present.

[0056] In another embodiment, the device further includes: The vehicle swapping unit is used to calculate the matching score between each candidate vehicle and the target vehicle in the infrared image. When the highest matching score is lower than the lowest matching score threshold, it determines whether there are parallel vehicles around the target vehicle. When it is determined, the target vehicle is swapped with the parallel vehicles, and the step of calculating the matching score by weighted sequence matching method is performed again based on the swapped vehicles.

[0057] The above embodiments of the present invention provide a vehicle matching method based on cross-modal multi-view, and a vehicle matching device based on cross-modal multi-view based on the method. The above method and device can solve the problem that the accuracy of vehicle matching in visible light camera images and infrared thermal imaging camera images is too low, and the temperature anomaly detection cannot be completed.

[0058] This embodiment also discloses a computer device, such as... Figure 5 As shown, the computer device includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to implement any of the above-described cross-modal multi-view vehicle matching methods.

[0059] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A vehicle matching method based on cross-modal multi-view, characterized in that, The method includes: Acquire visible light and infrared images, perform vehicle detection on the visible light images, and identify the target vehicle; The vehicles in the visible light image and infrared image are respectively subjected to BEV perspective transformation through a pre-calibrated homography matrix to obtain the normalized coordinates of each vehicle in the bird's-eye view space. The left-right relationship between adjacent vehicles is determined based on the normalized coordinates, and the left-right relationship sequence of each vehicle in the visible light and infrared images is constructed respectively. Based on a temporal search strategy, infrared image frames that match the visible light target vehicle in time are searched in an infrared image sequence, wherein the infrared image sequence is a set of multiple infrared images acquired by an infrared camera in chronological order. For each frame of infrared image searched, the lane number of the target vehicle in the visible light image and each candidate vehicle in the infrared image is calculated and lane consistency filtering is performed. Then, the matching score between each candidate vehicle in the infrared image and the target vehicle is calculated based on the left-right relationship sequence using a weighted sequence matching method, and the candidate vehicle with the highest matching score is selected as the target matching vehicle.

2. The method according to claim 1, characterized in that, The method further includes: When multiple infrared cameras are present, for the target matching vehicle identified by each infrared camera, the one with the smallest matching difference is selected as the final matching result based on a multi-level sorting strategy.

3. The method according to claim 1, characterized in that, The identified target vehicle includes: The vehicle identifier of each vehicle is determined, and the target vehicle is determined based on the vehicle identifier, wherein the vehicle identifier is the license plate number or the vehicle ID matched in the previous frame.

4. The method according to claim 1, characterized in that, Calculating lane numbers and performing lane consistency filtering includes: For a target vehicle in the visible light image, the lane to which it belongs is determined based on the horizontal coordinate of the center of its detection frame and the pre-calibrated horizontal coordinate range of each lane. For candidate vehicles in the infrared image, the lane to which they belong is determined based on their lateral normalized coordinates in the BEV space using a pre-calibrated lane boundary threshold. Candidate vehicles with the same lane number as the target vehicle are retained for subsequent matching.

5. The method according to claim 1, characterized in that, The determination of the left-right relationship between adjacent vehicles based on normalized BEV coordinates includes: For infrared images, a vehicle part detection model is used to identify vehicle parts including the hood, cab, cargo box, and tires, and parts of the same vehicle are aggregated by IoU intersection judgment. Based on different combinations of parts, a differentiated comparison strategy is used to determine the left-right relationship of adjacent vehicles.

6. The method according to claim 1, characterized in that, After calculating the matching score between each candidate vehicle and the target vehicle in the infrared image, the method further includes: When the highest matching score is lower than the lowest matching score threshold, it is determined whether there are parallel vehicles around the target vehicle. If so, the target vehicle is swapped with the parallel vehicle, and the step of calculating the matching score by weighted sequence matching is performed again based on the swapped vehicle.

7. The method according to claim 1, characterized in that, The time-series search strategy includes: The search is performed based on the timestamp status of the visible light image. When the camera timestamp on the visible light image is less than the server timestamp, it is determined to be a normal image, and N frames of infrared images are searched forward in time. When the camera timestamp is greater than or equal to the server timestamp, it is determined to be an abnormal image. The forward or backward search is dynamically determined based on the time difference with the previous visible light image and the number of vehicles.

8. A vehicle matching device based on cross-modal multi-view, characterized in that, The device includes: The target vehicle determination unit is used to acquire visible light images and infrared images, perform vehicle detection on the visible light images, and determine the target vehicle. The BEV transformation unit is used to perform BEV perspective transformation on the vehicles in the visible light image and infrared image respectively through a pre-calibrated homography matrix to obtain the normalized coordinates of each vehicle in the bird's-eye view space. The left-right relationship sequence determination unit is used to determine the left-right relationship between adjacent vehicles based on the normalized coordinates, and to construct the left-right relationship sequence of each vehicle in the visible light and infrared images respectively. An infrared image frame matching unit is used to search for infrared image frames that match the visible light target vehicle in time in an infrared image sequence based on a time-series search strategy. The infrared image sequence is a set of multiple infrared images acquired by an infrared camera in chronological order. The vehicle matching unit is used to calculate the lane number of the target vehicle in the visible light image and each candidate vehicle in the infrared image for each frame of infrared image searched, and to perform lane consistency filtering. Then, it uses a weighted sequence matching method to calculate the matching score between each candidate vehicle in the infrared image and the target vehicle based on the left-right relationship sequence, and selects the candidate vehicle with the highest matching score as the target matching vehicle.

9. The apparatus according to claim 8, characterized in that, The device further includes: The multi-camera matching unit is used to select the vehicle with the smallest matching difference as the final matching vehicle based on a multi-level sorting strategy when multiple infrared cameras are present.

10. The apparatus according to claim 8, characterized in that, The device further includes: The vehicle swapping unit is used to calculate the matching score between each candidate vehicle and the target vehicle in the infrared image. When the highest matching score is lower than the lowest matching score threshold, it determines whether there are parallel vehicles around the target vehicle. When it is determined, the target vehicle is swapped with the parallel vehicles, and the step of calculating the matching score by weighted sequence matching method is performed again based on the swapped vehicles.