A vehicle type recognition method and device based on multi-view cooperative perception and a medium

CN122676451APending Publication Date: 2026-09-01SHENZHEN GENVICT TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610861733.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0003]其中,单目视觉方案通过安装于门架前向的相机拍摄车头正面及车牌,依据车头特征推断车型类别,但由于其视野局限于车头单一视角,对于车长、轴数、货箱形态、车顶特征等关键分类参数无法直接获取,导致对半挂车、特种作业车等复杂车型的识别准确率较低

Benefits of technology

[0040]本发明提供了一种基于多目协同感知的车型识别方法、设备及介质,通过构建前向长焦相机、左侧斜向相机、右侧斜向相机、俯拍广角相机及后向追拍相机组成的多目采集架构,响应于触发信号同步获取多视角图像序列,并分别提取各视角对应的特征向量。通过对上述特征向量进行时空关联与特征融合,系统能够将分散于不同视角的车辆信息整合为统一的车型判断依据,弥补了单目相机因视角受限导致的车长、轴数、货箱形态等关键分类特征缺失的不足。基于融合后的车型概率分布及置信度值确定识别结果,实现了对通行车辆车型的高精度判定,有效解决了现有方案因信息不全导致的复杂车型识别准确率低的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676451A_ABST
    Figure CN122676451A_ABST
Patent Text Reader

Abstract

The application discloses a vehicle type recognition method and device based on multi-view cooperative perception and a medium, relates to the technical field of vehicle type recognition, and comprises the following steps: in response to receiving a vehicle trigger signal, controlling a front long-focus camera, a left oblique camera, a right oblique camera, a downward wide-angle camera and a rearward tracking camera to collect a multi-view image sequence; performing feature extraction on the multi-view image sequence to obtain a front feature vector, a left feature vector, a right feature vector, an overhead feature vector and a rear feature vector corresponding to each camera; performing space-time correlation and feature fusion on the feature vectors to obtain a fused vehicle type probability distribution and a confidence value; and determining a vehicle type recognition result according to the confidence value. The application can realize high-precision recognition of the vehicle type of a passing vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle model recognition technology, and in particular to a vehicle model recognition method, device and medium based on multi-view collaborative perception. Background Technology

[0002] In the field of intelligent traffic management on highways, the automatic and accurate identification of vehicle types is a core element in ensuring the fairness of differentiated toll collection. Current mainstream vehicle type recognition solutions primarily employ techniques such as monocular vision, inductive loop detectors combined with weighing, or laser contour scanning.

[0003] Among these solutions, the monocular vision method uses a camera mounted on the front of the gantry to capture images of the vehicle's front and license plate, inferring the vehicle type based on its frontal features. However, because its field of view is limited to a single viewpoint of the front of the vehicle, it cannot directly obtain key classification parameters such as vehicle length, number of axles, cargo box shape, and roof features, resulting in low accuracy in recognizing complex vehicle types such as semi-trailers and special-purpose vehicles. Furthermore, the inductive loop combined with weighing solution measures vehicle length using loops buried in the road surface and combines this with axle load data from a weighbridge to assist in vehicle type identification. However, this type of equipment suffers from short mean time between failures (MTBF) and high maintenance costs due to direct exposure to vehicle weight, and it cannot acquire three-dimensional vehicle contour information. Finally, the laser contour scanning solution uses a laser rangefinder to scan the vehicle contour, acquiring geometric information such as vehicle length, width, height, and wheelbase. However, the equipment is expensive, and its measurement accuracy is significantly affected by adverse weather conditions such as rain and fog.

[0004] It is evident that none of the above solutions can reliably and accurately identify vehicle types in complex traffic scenarios, resulting in a high system misjudgment rate and making it difficult to fully guarantee the fairness and accuracy of toll collection. Summary of the Invention

[0005] This invention provides a method, device, and medium for vehicle model recognition based on multi-view collaborative perception. The technical problem it aims to solve is: how to provide an effective solution for high-precision recognition of vehicle models in transit.

[0006] In a first aspect, the present invention provides a vehicle model recognition method based on multi-view collaborative perception, comprising:

[0007] In response to receiving a vehicle trigger signal, the system controls the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to acquire a multi-view image sequence.

[0008] Feature extraction is performed on the multi-view image sequence to obtain the forward feature vector corresponding to the forward telephoto camera, the left feature vector corresponding to the left oblique camera, the right feature vector corresponding to the right oblique camera, the top-view feature vector corresponding to the top-view wide-angle camera, and the rear feature vector corresponding to the rear-view tracking camera.

[0009] Spatiotemporal correlation and feature fusion are performed on the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector to obtain the fused vehicle probability distribution and confidence value.

[0010] The vehicle model identification result is determined based on the confidence level value.

[0011] The process of spatiotemporally associating and fusing the forward feature vector, the left-side feature vector, the right-side feature vector, the top-view feature vector, and the backward feature vector to obtain the fused vehicle model probability distribution and confidence value includes:

[0012] Determine whether the frontal view of the forward-facing telephoto camera is in a state of obstruction failure;

[0013] If the frontal view is not in an obstructed state, then based on the preset camera calibration parameters, the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector are spatiotemporally matched to generate a unified vehicle feature object; the unified vehicle feature object is input into the feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value;

[0014] If the frontal view is in an occlusion failure state, the forward feature vector is removed from the feature vector. Based on the preset camera calibration parameters, the remaining left-side feature vector, right-side feature vector, top-view feature vector, and rear-side feature vector are spatiotemporally matched to generate a unified vehicle feature object. The unified vehicle feature object is then input into the feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value.

[0015] Optionally, determining whether the frontal view of the forward-facing telephoto camera is in an obstruction-free state includes:

[0016] Obtain the top profile of the vehicle, and perform connected component analysis or convex hull analysis on the top profile of the vehicle to determine whether there are abrupt width changes or concave regions in the profile.

[0017] If the aforementioned width abrupt change or contour concavity area exists, it is determined that the frontal view of the forward telephoto camera is in an occlusion failure state, and the current scene is identified as an abnormal mode where a large vehicle occludes a small vehicle.

[0018] If there is no abrupt change in width or concave contour area, it is determined that the frontal view of the forward telephoto camera is not in an obstructed state.

[0019] Optionally, after identifying the abnormal pattern of a large vehicle obscuring a small vehicle in the current scene, the method further includes:

[0020] Based on the left-side feature vector, the right-side feature vector, and the top-view feature vector, a multi-target tracking algorithm is used to track the motion trajectories of multiple independent objects across frames.

[0021] The velocity and acceleration of the motion trajectories of the multiple independent objects are analyzed to determine whether there is a velocity difference or an acceleration difference in the motion trajectories of the independent objects.

[0022] If the speed difference or acceleration difference exists, then the multiple independent objects are determined to belong to different vehicles;

[0023] When the obscured vehicle in the different vehicles drives out of the obscured area and reappears in the field of view of the forward telephoto camera, a new forward feature vector of the obscured vehicle is obtained based on the image re-acquired by the forward telephoto camera.

[0024] The new forward feature vector is spatiotemporally matched with the historical feature vector and the trajectory continuity is verified; the historical feature vector is the left feature vector, right feature vector or top-view feature vector corresponding to the occluded vehicle before or during the occlusion.

[0025] If the spatiotemporal matching and trajectory continuity verification pass, the feature vector of the occluded vehicle is updated based on the new forward feature vector.

[0026] Optionally, the control of the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to acquire a multi-view image sequence includes:

[0027] In response to receiving the vehicle trigger signal, a PPS signal is sent through the SOC to drive the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to perform synchronous exposure at the same time, thereby obtaining a multi-view image sequence with inter-frame jitter of less than or equal to 500 microseconds.

[0028] Optionally, the method further includes:

[0029] Before acquiring the multi-view image sequence, the coordinate system of the overhead wide-angle camera is used as the origin of the world coordinate system to jointly calibrate the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to obtain camera calibration parameters; wherein the reprojection error of the joint calibration is less than or equal to 0.5 pixels.

[0030] Optionally, the method further includes:

[0031] Based on the forward and top views in the multi-view image sequence, the vehicle's three-dimensional contour is calculated using a binocular triangulation algorithm with preset camera calibration parameters to obtain the vehicle's height, width, and length estimates. These estimates are then used as auxiliary features and input into the feature fusion model for feature fusion.

[0032] Optionally, determining the vehicle model recognition result based on the confidence value includes:

[0033] Determine whether the confidence value is greater than or equal to a first preset threshold;

[0034] If the confidence value is greater than or equal to the first preset threshold, the vehicle model recognition result is output directly;

[0035] Determine whether the confidence value is less than the first preset threshold and greater than or equal to the second preset threshold;

[0036] If the confidence value is less than the first preset threshold and greater than or equal to the second preset threshold, the ETC database is retrieved to verify the vehicle identification result, and the verified result is output.

[0037] If the confidence level is less than the second preset threshold, a manual review process is triggered.

[0038] Secondly, the present invention also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0039] Thirdly, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.

[0040] This invention provides a vehicle model recognition method, device, and medium based on multi-view collaborative perception. It constructs a multi-view acquisition architecture consisting of a forward-facing telephoto camera, a left-side oblique camera, a right-side oblique camera, a top-down wide-angle camera, and a rear-facing tracking camera. Responding to a trigger signal, it synchronously acquires multi-view image sequences and extracts feature vectors corresponding to each viewpoint. By performing spatiotemporal correlation and feature fusion on these feature vectors, the system can integrate vehicle information scattered across different viewpoints into a unified basis for vehicle model judgment, compensating for the shortcomings of monocular cameras that lack key classification features such as vehicle length, number of axles, and cargo box shape due to limited field of view. Based on the fused vehicle model probability distribution and confidence value, the recognition result is determined, achieving high-precision determination of the vehicle model of passing vehicles and effectively solving the problem of low accuracy in complex vehicle model recognition caused by incomplete information in existing solutions. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 A flowchart illustrating a vehicle model recognition method based on multi-view collaborative perception provided in an embodiment of the present invention;

[0043] Figure 2 This is a schematic diagram of the overall hardware architecture of a multi-view camera system provided in an embodiment of the present invention;

[0044] Figure 3 This is a schematic diagram of the timing of multi-view collaborative occlusion processing provided in an embodiment of the present invention;

[0045] Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0046] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0048] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0049] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0050] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0051] Please see Figure 1 This invention provides a vehicle model recognition method based on multi-view collaborative perception. By constructing a five-view collaborative perception matrix of "front-left-right-top-rear" and combining the spatiotemporal correlation and feature fusion of multi-view features, it achieves high-precision and robust vehicle model recognition for passing vehicles. Specifically, the method includes the following steps:

[0052] S1, in response to receiving a vehicle trigger signal, controls the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to acquire a multi-view image sequence.

[0053] In specific implementation, in response to receiving a vehicle trigger signal, the system controls a forward-facing telephoto camera, a left-side oblique camera, a right-side oblique camera, a top-down wide-angle camera, and a rear-facing tracking camera to acquire a multi-view image sequence. The vehicle trigger signal originates from an inductive loop, millimeter-wave radar, or the pre-detection function of the forward-facing telephoto camera; this invention is not specifically limited to these sources. The trigger signal is activated when the vehicle enters a preparatory area 15 to 30 meters in front of the gate.

[0054] Furthermore, the forward-facing telephoto camera is mounted directly in front of the highway gantry crossbeam, with a focal length greater than or equal to 200mm, a resolution of 8 megapixels (4K), and a frame rate of 60 fps, used to clearly capture license plates and vehicle front profile features at a distance of 15 m to 30 m. Further, the left-side oblique camera is mounted on the left end of the gantry crossbeam, deflected downwards and to the right, with a yaw angle of +30 degrees and a pitch angle of -30 degrees; the right-side oblique camera is mounted on the right end of the gantry crossbeam, deflected downwards and to the left, with a yaw angle of -30 degrees and a pitch angle of -30 degrees. Both the left-side and right-side oblique cameras have a variable focal length range of 16 mm to 35 mm, a resolution of 5 megapixels, and a frame rate of 30 fps, used to capture the left and right sides of the vehicle and extract lateral features such as vehicle length, number of axles, and cargo box shape. Furthermore, the overhead wide-angle camera is mounted in the center of the gantry crossbeam, facing directly downwards. It is a fisheye lens with a focal length of 8 mm or less, a resolution of 8 megapixels, a frame rate of 25 fps, and a field of view of 120 degrees x 120 degrees or more. It is used to acquire roof features (including hazardous materials markings and roof shape), vehicle width information, and lane assignment determination from an overhead perspective. Furthermore, the rear-view camera is mounted behind the gantry crossbeam, with a pitch angle of -20 degrees, a focal length of 50 mm to 85 mm, a resolution of 5 megapixels, and a frame rate of 30 fps. It is used to capture the rear license plate and taillight shape after the vehicle has passed the gantry, enabling secondary verification.

[0055] Furthermore, the five cameras are uniformly controlled at the hardware level through a System-on-a-Chip (SOC). The SOC drives each camera to perform synchronized exposure at the same time by sending PPS (pulses per second) signals, ensuring that inter-frame jitter is less than or equal to 500 microseconds. The ISP (Image Signal Processor) integrated within the SOC supports five MIPI CSI-2 inputs and incorporates algorithms such as HDR (High Dynamic Range), 3D-NR (3D Noise Reduction), and Dehazing (DCP, Dark Channel Prior) to preprocess the raw images from each input to eliminate the effects of adverse weather conditions such as backlighting, nighttime, or rain and fog. Furthermore, the forward-facing telephoto camera, the left-side oblique camera, the right-side oblique camera, and the overhead wide-angle camera are connected to the SOC via MIPI CSI-2 interfaces, while the rear-facing tracking camera is connected to the SOC via a USB 3.0 interface.

[0056] For example, in some preferred embodiments, the control of the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to acquire a multi-view image sequence includes: in response to receiving the vehicle trigger signal, sending a PPS signal through the SOC to drive the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to perform synchronous exposure at the same time, thereby obtaining a multi-view image sequence with inter-frame jitter of less than or equal to 500 microseconds.

[0057] In specific implementation, in response to receiving the vehicle trigger signal, the SOC (System-on-a-Chip) sends a PPS (Pulses Per Second) signal to drive the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to perform synchronous exposure at the same time, resulting in a multi-view image sequence with inter-frame jitter less than or equal to 500 microseconds. The SOC integrates an image signal processing unit, which sends a VSYNC signal through its GPIO pin as a unified exposure trigger signal to drive the five cameras to achieve hardware-level frame synchronization. Furthermore, the PPS signal provides a precise time reference for each camera, ensuring that they complete exposure under the same time reference. In the specific hardware implementation, the SOC connects to each camera via MIPI CSI-2 interfaces (CAM1 and CAM4) and USB 3.0 interfaces (CAM2, CAM3, and CAM5), and simultaneously links with the fill light control system via an RS-485 bus to trigger strobe lighting in backlit or nighttime scenes (white light strobe ≥30000 lux or infrared 940 nm covert fill light, strobe response ≤5 ms). This high-precision hardware-synchronized exposure mechanism is the foundation for subsequent accurate spatiotemporal matching and correlation of multi-view features in the time dimension.

[0058] For example, please see Figure 2This document demonstrates the overall hardware architecture of the multi-camera system in this embodiment. The core of this architecture is an integrated computing product that combines the aforementioned SOC (System-on-a-Chip) and MEC (Mobile Edge Computing) on ​​a single motherboard. Furthermore, this integrated computing product presents itself as a single device, integrating external interfaces including a video data stream interface (MIPI CSI-2 / USB 3.0), a network communication interface (Gigabit Ethernet, 5G / 4G backup transmission), a control signal interface (RS-485), and a power supply interface (PoE / DC). Further, the forward-facing telephoto camera (CAM1) and the overhead wide-angle camera (CAM4) are connected to the integrated computing product via the MIPI CSI-2 interface; the left-side oblique camera (CAM2), the right-side oblique camera (CAM3), and the rear-facing tracking camera (CAM5) are connected via the USB 3.0 interface. Additionally, the integrated computing product connects to a supplementary lighting control system via an RS-485 interface and provides power to each camera via a PoE / DC power supply unit. Furthermore, the integrated computing power product interacts with the toll station control center and cloud data platform via a gigabit Ethernet interface. This embodiment... Figure 2 The hardware architecture shown achieves physical-level integration from front-end sensing to back-end computing, transmission and power supply, ensuring high-speed data interaction and collaborative work of five cameras on the same physical device.

[0059] This embodiment uses a SOC to send a PPS signal to drive five cameras to perform synchronized exposure at the same time, achieving hardware-level frame synchronization with inter-frame jitter of less than or equal to 500 microseconds. Compared with traditional software timestamp synchronization schemes, this hardware synchronization method reduces the capture time difference between cameras to the sub-millisecond level, eliminating motion blur and feature misalignment between viewpoints caused by high-speed vehicle travel (e.g., above 100 km / h). Because the five images are precisely aligned in the temporal dimension, subsequent spatiotemporal matching based on preset camera calibration parameters can accurately map feature vectors from different viewpoints onto the same vehicle, thus ensuring the accuracy of feature fusion. Furthermore, the coordinated control with the supplementary lighting system ensures that each camera can obtain high-quality images that meet analysis requirements under complex lighting conditions such as backlighting and nighttime, further enhancing the system's all-weather operating capability.

[0060] In some preferred embodiments, the method further includes: before acquiring the multi-view image sequence, using the coordinate system of the overhead wide-angle camera as the origin of the world coordinate system, jointly calibrating the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to obtain camera calibration parameters; wherein the reprojection error of the joint calibration is less than or equal to 0.5 pixels.

[0061] In specific implementation, before acquiring the multi-view image sequence, the coordinate system of the overhead wide-angle camera is used as the origin of the world coordinate system to jointly calibrate the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to obtain camera calibration parameters. Specifically, in a preferred embodiment, a 3×5 checkerboard calibration board or multiple frames of road marking data are used to calculate the intrinsic parameter matrix (including focal length, principal point, and distortion coefficient) and extrinsic parameter matrix (including rotation matrix and translation vector) of each camera using the Zhang Zhengyou calibration method.

[0062] Furthermore, the coordinate system of the overhead wide-angle camera is used as the world coordinate system, and the intrinsic and extrinsic parameters of other cameras are calculated relative to this coordinate system. During calibration, the system arranges multiple calibration points in the area below the gantry, covering the entire field of view, and minimizes the reprojection error through an optimization algorithm. Furthermore, the reprojection error of the joint calibration is less than or equal to 0.5 pixels, ensuring calibration accuracy. These calibration parameters are used for spatiotemporal matching and 3D contour calculation in subsequent steps. Furthermore, after calibration, the system supports online self-calibration based on fixed reference objects (such as road shoulder boundary markers), which can be automatically performed every 24 hours to eliminate calibration parameter drift caused by temperature changes or minor displacements of the installation structure.

[0063] This embodiment provides a precise geometric benchmark for subsequent spatiotemporal matching and 3D contour calculation by performing high-precision multi-view camera joint calibration before data acquisition. By setting the coordinate system of the overhead wide-angle camera as the origin of the world coordinate system and calculating the intrinsic and extrinsic parameters of each camera with a reprojection error of less than or equal to 0.5 pixels, the system can accurately map the pixel coordinates of camera images from different viewpoints and positions to a unified 3D spatial coordinate system. This allows for accurate spatial alignment and association of feature vectors extracted from different viewpoints, a prerequisite for achieving accurate vehicle 3D contour measurement, occlusion region identification, and feature fusion.

[0064] S2, perform feature extraction on the multi-view image sequence to obtain the forward feature vector corresponding to the forward telephoto camera, the left feature vector corresponding to the left oblique camera, the right feature vector corresponding to the right oblique camera, the top-view feature vector corresponding to the top-view wide-angle camera, and the rear feature vector corresponding to the rear-view tracking camera.

[0065] In specific implementation, feature extraction is performed on the multi-view image sequence to obtain the forward feature vector corresponding to the forward telephoto camera, the left feature vector corresponding to the left oblique camera, the right feature vector corresponding to the right oblique camera, the top-view feature vector corresponding to the top-view wide-angle camera, and the rear feature vector corresponding to the rear-view tracking camera.

[0066] For example, in a preferred embodiment, the YOLOv8-L object detection model is used to perform parallel inference on each frame of image, outputting the bounding box (BBox), mask, and 512-dimensional feature vector corresponding to each image. Further, the feature vector contains key visual features of the vehicle from the corresponding viewpoint, specifically including: license plate characters, front face recognition code, and aspect ratio of the vehicle's front profile in the forward feature vector; vehicle length (converted to meters via pixel estimation), number of axles (identified via wheel set), and cargo box or body shape in the left and right feature vectors; vehicle roof area (width multiplied by length) and top features such as hazardous materials signs or antennas in the top-view feature vector; and taillight shape, rear license plate, and license plate frame type (blue or yellow) in the rear feature vector, which are not specifically limited in this invention. These features together constitute a multi-dimensional description of the vehicle, providing a rich data foundation for subsequent vehicle type recognition.

[0067] S3, perform spatiotemporal correlation and feature fusion on the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector and the backward feature vector to obtain the fused vehicle probability distribution and confidence value.

[0068] In specific implementation, the forward feature vector, left feature vector, right feature vector, top-view feature vector, and backward feature vector are spatiotemporally correlated and fused to obtain the fused vehicle probability distribution and confidence value. Specifically, the system performs the following operations based on preset camera calibration parameters: First, the image pixel coordinates corresponding to the forward feature vector, left feature vector, right feature vector, top-view feature vector, and backward feature vector are mapped from their respective camera coordinate systems to a unified world coordinate system with the top-view wide-angle camera as the origin. This mapping process is achieved through a mathematical transformation matrix, eliminating spatial deviations introduced by different camera installation positions and angles. Furthermore, in the time dimension, the system uses the timestamps of each camera image acquisition time, combined with the vehicle's motion trajectory when passing through the gantry, to perform spatiotemporal matching to determine which feature vectors belong to the same passing vehicle. Further, the multi-view feature vectors that successfully match spatiotemporally and are determined to belong to the same vehicle are encapsulated to generate an independent "unified vehicle feature object". This object is a data structure containing feature vectors of the vehicle from the respective perspectives of the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera. This "unified vehicle feature object" is then input into a feature fusion model for weighted fusion.

[0069] Furthermore, the feature fusion model adopts a Cross-Attention Transformer architecture. It dynamically calculates the attention weights of features from each perspective, performs weighted fusion of the five features, and outputs the fused vehicle probability distribution and the corresponding confidence value C. The Cross-Attention Transformer is a deep learning model based on an attention mechanism, capable of establishing long-distance dependencies between features from different perspectives, thereby achieving optimal feature fusion.

[0070] Further, in some preferred embodiments, the step of performing spatiotemporal correlation and feature fusion on the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector to obtain the fused vehicle model probability distribution and confidence value includes: determining whether the frontal view of the forward telephoto camera is in an occlusion failure state; if the frontal view is not in an occlusion failure state, then based on preset camera calibration parameters, performing spatiotemporal matching on the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector to generate... A unified vehicle feature object is generated; the unified vehicle feature object is input into a feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value; if the frontal view is in an occlusion failure state, the forward feature vector is removed from the feature vector, and based on the preset camera calibration parameters, the remaining left feature vector, right feature vector, top-view feature vector, and rear feature vector are spatiotemporally matched to generate a unified vehicle feature object; the unified vehicle feature object is input into the feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value.

[0071] In practice, the system first determines whether the frontal view of the forward-facing telephoto camera is in an occlusion failure state. When a large truck obstructs a small vehicle behind it, the forward-facing camera may not be able to effectively capture the license plate and front-facing features of the small vehicle. The system makes this determination in several ways: for example, by analyzing the vehicle's top profile presented by the top-view feature vector to determine if there are any abrupt width changes or concave areas in the profile; or by detecting target loss in the forward-facing camera through continuous frame tracking, while the side and top-view cameras still have multiple independent trajectories. If the determination result is occlusion failure, the system identifies the current scene as an abnormal pattern where a large truck obstructs a small vehicle.

[0072] Furthermore, if the frontal view is not in an obstructed state, then based on preset camera calibration parameters, spatiotemporal matching is performed on the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the rear feature vector to generate a unified vehicle feature object. Specifically, based on preset camera calibration parameters, the image pixel coordinates corresponding to the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector are mapped from their respective camera coordinate systems to a unified world coordinate system with the top-view wide-angle camera as the origin, obtaining the three-dimensional coordinates of each feature vector in physical space. On this basis, using the timestamps accompanying the synchronous acquisition by each camera, combined with the spatial motion trajectory and continuity constraints of the vehicle passing through the gantry, spatiotemporal matching is performed on feature vectors extracted from different perspectives and timestamps. Feature vectors with matching spatial positions and time sequences conforming to the vehicle's motion patterns are determined to belong to the same passing vehicle. All feature vectors determined to belong to the same vehicle are encapsulated in a data structure to generate a unified vehicle feature object. This unified vehicle feature object includes the forward feature vector, left feature vector, right feature vector, top-view feature vector, and backward feature vector corresponding to the vehicle, and serves as the input data for the subsequent feature fusion model.

[0073] Furthermore, the unified vehicle feature object is input into a feature fusion model (e.g., Cross-AttentionTransformer) for weighted fusion to obtain the fused vehicle model probability distribution and confidence value.

[0074] Furthermore, if the frontal view is in an occluded state, the forward feature vector is removed from the feature vector. Based on the preset camera calibration parameters, spatiotemporal matching is performed on the remaining left-side feature vector, right-side feature vector, top-view feature vector, and rear-view feature vector to generate a unified vehicle feature object. This object only contains unoccluded side, top, and rear features. Subsequently, the unified vehicle feature object is input into the feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value.

[0075] It should be noted that, in this embodiment of the invention, the feature fusion model is specifically constructed and trained in the following manner: the model includes a feature embedding layer, multiple cascaded cross-attention encoder layers, and a classification head; wherein, the feature embedding layer maps the feature vectors of each viewpoint to feature tokens of the same dimension, and adds viewpoint type encoding and temporal position encoding; the cross-attention encoder layer calculates the attention weight between the classification token and the feature tokens of each viewpoint through a multi-head cross-attention mechanism, and performs weighted aggregation of the features of each viewpoint; the classification head outputs the vehicle model probability distribution and confidence value through a linear layer and a Softmax layer. Further, during training, a training sample set covering different vehicle models, lighting conditions, and occlusion states is obtained. For each training sample, the extracted multi-view feature vector is input into the model, the cross-entropy loss function is used to calculate the loss value between the prediction result and the true label, and the AdamW optimizer is used to update the network parameters in reverse until the model converges. Through the above training, the model can adaptively enhance the weights of the remaining effective viewpoint features when some viewpoint features are missing or invalid, and robustly output the vehicle model recognition result.

[0076] This embodiment solves the problem of lost forward camera information caused by vehicle occlusion in existing technologies by determining whether the frontal view is in an occluded state, selectively removing features from the occluded view and retaining features from the valid view for fusion. Its technical principle lies in utilizing the vertical downward view advantage of the top-view camera, reliably determining occlusion by analyzing the geometric features of the vehicle's top profile (amplitude changes, contour indentations). Therefore, when the frontal view fails, the system can automatically activate side and top-view cameras for compensation. Measurements show that this improves the vehicle type recognition accuracy in occluded scenarios from 55% in existing solutions to over 93%, fundamentally solving the serious problem of missed charges caused by large trucks obscuring smaller vehicles.

[0077] Furthermore, to more intuitively describe the collaborative working process of each camera and the timeline logic of feature fusion decision in the above "large vehicle occluding small vehicle" scenario, please refer to [link to relevant documentation]. Figure 3 . Figure 3The sequence diagram illustrates the multi-camera collaborative occlusion processing timeline. This timeline is based on the timeline of a vehicle passing through the preparatory area (Z1), main recognition area (Z2), and verification area (Z3) of the gantry along its travel direction. Further, during the t0 to t1 phase of this timeline, the vehicle enters the Z1 region, and the forward-facing telephoto camera (CAM1) performs a forward pre-detection and initially identifies the license plate. Further, during the t1 to t2 phase, the vehicle enters the Z2 region. If a large vehicle occludes a smaller vehicle, CAM1 is in a disabled state (marked as △ occlusion / backlighting in the diagram, disabled), and cannot acquire the smaller vehicle's license plate and front-end features. At this time, the continuous side imaging from the left oblique camera (CAM2) and the right oblique camera (CAM3) (axle number / vehicle length / cargo box continuously visible, no occlusion), and the continuous overhead shooting from the overhead wide-angle camera (CAM4) (roof features, vehicle width, lane affiliation fully visible), become the primary data sources for determining the existence of the occluded vehicle and extracting its key features. System Activation Figure 3 The compensation fusion decision shown is as follows: CAM1 fails → CAM2 / CAM3 / CAM4 features are activated for compensation to generate a unified vehicle feature object. Further, during the t2 to t3 phase when the vehicle passes the gantry and enters the Z3 region, CAM1 resumes recognition, and simultaneously the rear-view camera (CAM5) starts working, capturing the license plate and rear features for verification. At this time, the system merges the information from the five cameras (CAM1 and CAM2, CAM3, CAM4, CAM5), performs spatiotemporal matching and feature fusion, and outputs the final high-confidence vehicle model recognition result. Figure 3 The timing logic shown ensures the continuity and accuracy of vehicle model recognition and vehicle identity tracking in complex occlusion scenarios.

[0078] In some preferred embodiments, determining whether the frontal view of the forward-facing telephoto camera is in an occlusion failure state includes: acquiring the vehicle's top profile, performing connected component analysis or convex hull analysis on the vehicle's top profile, and determining whether there are width abrupt changes or profile concave regions; if the width abrupt changes or profile concave regions exist, then the frontal view of the forward-facing telephoto camera is determined to be in an occlusion failure state, and the current scene is identified as an abnormal pattern where a large vehicle occludes a small vehicle; if the width abrupt changes or profile concave regions do not exist, then the frontal view of the forward-facing telephoto camera is determined not to be in an occlusion failure state.

[0079] In specific implementation, the first step is to acquire the vehicle's roof outline. In a preferred embodiment, the system extracts the vehicle's roof outline information from the top-view feature vector. Since the top-view wide-angle camera is mounted vertically downwards, its field of view is greater than or equal to 120 degrees multiplied by 120 degrees, enabling clear and complete acquisition of the two-dimensional planar shape of the vehicle roof, including its width, length, and the presence of features such as hazardous materials markings and antennas. The vehicle's roof outline is presented in the form of image pixel coordinates or three-dimensional spatial coordinates.

[0080] Furthermore, connected component analysis or convex hull analysis is performed on the vehicle's top profile to determine whether there are abrupt width changes or concave regions. When two independent vehicles (e.g., a large vehicle and a small vehicle) are driving closely behind each other, even if the large vehicle obscures the front of the small vehicle, their top profiles are still spatially separated, forming an extremely long and irregular top profile with abrupt width changes or concave regions. Connected component analysis can identify these discontinuous regions, while convex hull analysis can detect geometric abrupt changes in the profile. Connected component analysis is an image processing technique used to detect connected regions in an image with the same gray value or similar features; convex hull analysis is used to detect the minimum convex polygon boundary of a point set, thereby identifying concave parts of the profile.

[0081] Furthermore, if the aforementioned abrupt width change or contour depression area exists, it is determined that the frontal view of the forward-facing telephoto camera is in an occlusion failure state, and the current scene is identified as an abnormal pattern where a large vehicle occludes a small vehicle. This indicates that the frontal camera cannot obtain valid information about the occluded vehicle, and an occlusion compensation mechanism needs to be activated.

[0082] Furthermore, if there are no abrupt width changes or concave contour areas, it is determined that the frontal view of the forward-facing telephoto camera is not in an occlusion failure state. This indicates that the current vehicle top contour is complete and continuous, the front camera is functioning normally, and no compensation mechanism is required.

[0083] This embodiment provides a precise and robust technical means for determining occlusion failure states by analyzing the vehicle's top profile based on top-view feature vectors. Unlike traditional solutions that rely solely on the brightness of the forward-facing camera image or target loss to determine occlusion, this embodiment analyzes the geometric features of the vehicle's roof profile (abrupt width changes, profile concavity) to accurately identify the abnormal pattern in space the instant a larger vehicle occludes a smaller vehicle. Furthermore, due to the vertical downward viewing angle of the top-view camera, the top profile information it acquires is unaffected by frontal occlusion, thus ensuring high reliability of the determination results. This allows the system to trigger feature removal and compensation logic instantly and reliably without waiting for the occluded vehicle to reappear, thereby guaranteeing the continuity and real-time nature of the vehicle model recognition process.

[0084] In some preferred embodiments, after identifying the abnormal pattern of a large vehicle obscuring a small vehicle in the current scene, the method further includes: using a multi-target tracking algorithm to track the motion trajectories of multiple independent objects across frames based on the left-side feature vector, the right-side feature vector, and the top-view feature vector; performing velocity and acceleration analysis on the motion trajectories of the multiple independent objects to determine whether there is a velocity difference or acceleration difference in the motion trajectories of each independent object; if there is a velocity difference or acceleration difference, then determining that the multiple independent objects belong to different vehicles; when the obscured vehicle among the different vehicles drives out of the obscuration area and reappears in the field of view of the forward telephoto camera, obtaining a new forward feature vector of the obscured vehicle based on the image re-acquired by the forward telephoto camera; performing spatiotemporal matching and trajectory continuity verification between the new forward feature vector and the historical feature vector; the historical feature vector is the left-side feature vector, right-side feature vector, or top-view feature vector corresponding to the obscured vehicle before or during the obscuration; if the spatiotemporal matching and trajectory continuity verification pass, updating the feature vector of the obscured vehicle based on the new forward feature vector.

[0085] In specific implementation, firstly, based on the left-side feature vector, the right-side feature vector, and the top-view feature vector, a multi-object tracking algorithm is used to track the motion trajectories of multiple independent objects across frames. In a preferred embodiment, advanced multi-object tracking algorithms such as DeepSORT or ByteTrack are used to continuously track each independent contour segmented from the lateral and top-view features, forming their respective motion trajectories. DeepSORT (Deep Simple Online and Realtime Tracking) is an online multi-object tracking algorithm that combines deep learning feature extraction. Its core idea is to combine object detection (YOLOv8-L) with a feature extraction network, extracting visual features (e.g., appearance and texture) and motion features (e.g., position and velocity) for each detected contour in each time frame, and predicting the trajectory using Kalman filtering. ByteTrack is a simple and efficient multi-object tracking algorithm based on detection results, improving tracking performance by associating all detection boxes (including low-confidence detection boxes). During tracking, the system assigns a unique tracking ID to each contour and performs feature matching and position prediction between consecutive frames to maintain the persistence of the ID.

[0086] Furthermore, the system analyzes the velocity and acceleration of the motion trajectories of the multiple independent objects to determine whether there are velocity or acceleration differences in their trajectories. Even when two independent vehicles are closely following each other, their motion trajectories remain independent. For example, when the vehicle in front brakes, the vehicle behind will have a brief delay before starting to decelerate; this slight difference in acceleration is strong evidence of independence. Further, the system calculates the instantaneous velocity and acceleration of each trajectory and compares them. If a velocity or acceleration difference is detected, the system determines that the multiple independent objects belong to different vehicles. When two independent vehicles are closely following each other, although there is physical obstruction between them, their motion trajectories remain independent in the time dimension. Specifically, there are deviations in the instantaneous velocity, acceleration, and steering fine-tuning details between the vehicle in front (larger vehicle) and the vehicle behind (smaller vehicle), especially when the vehicle in front brakes, the vehicle behind will have a brief delay before starting to decelerate. Therefore, by comparing the velocity and acceleration differences between adjacent trajectories, the system can reliably determine whether they are independent vehicle entities. This motion analysis mechanism, combined with spatial dimension contour analysis, constitutes a doubly robust vehicle separation strategy.

[0087] Furthermore, when the obscured vehicle leaves the obscured area and reappears in the field of view of the forward telephoto camera, a new forward feature vector of the obscured vehicle is obtained based on the image re-captured by the forward telephoto camera. At this time, the forward camera recaptures the front face and license plate information of the vehicle.

[0088] Furthermore, the new forward feature vector is spatiotemporally matched and its trajectory continuity is verified with historical feature vectors. The historical feature vectors are the left-side, right-side, or top-view feature vectors corresponding to the occluded vehicle before or during occlusion. The verification process includes comparing the newly appearing vehicle trajectory with the trajectory of the following vehicle previously tracked by the lateral and top-view cameras in terms of position, speed, and direction. If the matching degree meets a preset threshold, such as a position deviation of less than 0.5 m and a speed deviation of less than 5%, the spatiotemporal matching and trajectory continuity verification are considered successful. The feature vector of the occluded vehicle is updated based on the new forward feature vector, thereby fully recovering the vehicle's full-view features.

[0089] This embodiment achieves continuous tracking and final confirmation of a vehicle's identity throughout the entire process of vehicle occlusion by comprehensively utilizing spatial dimension segmentation (contour analysis) and temporal dimension segmentation (motion trajectory analysis). The technical principle lies in the fact that multi-target tracking algorithms such as DeepSORT utilize both visual and motion features, combined with Kalman filtering for trajectory prediction, effectively handling situations where the target is temporarily occluded. When the frontal view fails, continuous observation from side and top views, along with motion analysis algorithms, separates the occluded vehicle from closely spaced vehicles while maintaining its ID, thus avoiding vehicle identity loss or misjudgment due to occlusion. Furthermore, when the occluded vehicle reappears, trajectory continuity verification accurately correlates new forward features with historical features, completing the reconstruction of the vehicle's full-view features. This multi-dimensional, multi-stage processing approach enables the system to maintain high-precision vehicle model recognition and vehicle identity consistency even when handling complex car-following scenarios.

[0090] In some preferred embodiments, the method further includes: based on the forward image and top view image in the multi-view image sequence, using preset camera calibration parameters and a binocular triangulation algorithm to calculate the three-dimensional contour of the vehicle, obtaining the vehicle's height estimate, width estimate and length estimate, and inputting the vehicle's height estimate, width estimate and length estimate as auxiliary features into the feature fusion model to participate in feature fusion.

[0091] In specific implementation, based on the forward and top-view images in the multi-view image sequence, the system uses preset camera calibration parameters and a binocular triangulation algorithm to calculate the vehicle's three-dimensional contour, obtaining estimated height, width, and length values. Specifically, the system detects the pixel coordinates of the highest point of the roof and the tire contact point in the forward image, and calculates the vehicle height using trigonometric relationships by combining the camera's mounting height (H), pitch angle (θ), focal length (f), and pixel size. The calculation formula is approximately: Vehicle height ≈ H - (h_pixel × f) / (pixel size × sin(θ)), where h_pixel is the pixel height from the bottom of the vehicle to the roof. Simultaneously, the roof boundary is extracted from the top-view image, and inverse perspective transformation is used to map the image pixels to the real-world coordinate system, thereby calculating the vehicle width and length.

[0092] Furthermore, in a preferred embodiment, the measurement accuracy is improved by averaging the results of consecutive frames (3-5 frames) and removing outliers. Further, the calculated vehicle height, width, and length are compared with the standard height, width, and length ranges of common vehicle types (sedans, SUVs, light trucks, heavy trucks) to help correct the calculation results. Finally, the estimated vehicle height, width, and length values ​​are used as auxiliary features input into the feature fusion model to participate in feature fusion, providing richer geometric dimensional information for vehicle type classification.

[0093] This embodiment achieves non-contact online measurement of vehicle 3D contours using only existing forward-facing and top-facing cameras and a binocular triangulation algorithm, without adding expensive sensors such as LiDAR. Because the system can directly calculate key geometric features such as vehicle height (accuracy ±0.2 m), width (accuracy ±0.1 m), and length (accuracy ±0.1 m), and use these as auxiliary features input to the feature fusion model, it provides a more physically meaningful geometric dimension for vehicle classification, in addition to visual appearance features. This is particularly helpful in distinguishing between vehicles with similar appearances but different sizes, such as sedans and SUVs, light trucks and heavy trucks, thereby further improving the accuracy of vehicle identification and enabling the detection of overloaded vehicles.

[0094] S4, determine the vehicle model recognition result based on the confidence value.

[0095] In practice, the vehicle model recognition result is determined based on the confidence level value. When the confidence level value C is greater than or equal to 0.95, the vehicle model and license plate information are directly output; when C is between 0.7 and 0.95, the ETC (Electronic Toll Collection) database is retrieved to verify the vehicle model recognition result, and the verified result is jointly reported with the visual recognition result; when C is less than 0.7, a manual review process is triggered, and the back-end staff retrieves five synchronous images of the vehicle through the monitoring terminal for manual judgment.

[0096] This embodiment constructs a five-view collaborative perception architecture (front-left-right-top-rear) and precisely correlates and fuses feature vectors from each viewpoint across time and space, solving the information deficiency problem caused by the single viewpoint in existing monocular solutions. Furthermore, because the system can simultaneously acquire multi-dimensional visual features of the vehicle's front, sides, top, and rear, and map these features to a unified world coordinate system using mathematical transformation matrices, spatial bias is eliminated, enabling accurate extraction of key classification information such as vehicle length, number of axles, cargo box shape, roof features, and taillight shape. Furthermore, for complex vehicle types such as semi-trailers and special-purpose vehicles, the recognition accuracy of this embodiment is improved from below 70% in existing monocular solutions to over 99%. Furthermore, through the complementarity of multi-view features, the system's robustness in adverse lighting conditions such as backlight, nighttime, and rain / fog is significantly enhanced; the recognition rate in backlight and nighttime scenes increases from 65% to over 95%, while the misclassification rate decreases to below 0.3%.

[0097] In some preferred embodiments, determining the vehicle model recognition result based on the confidence value includes: determining whether the confidence value is greater than or equal to a first preset threshold; if the confidence value is greater than or equal to the first preset threshold, directly outputting the vehicle model recognition result; determining whether the confidence value is less than the first preset threshold and greater than or equal to a second preset threshold; if the confidence value is less than the first preset threshold and greater than or equal to the second preset threshold, retrieving the ETC database to verify the vehicle model recognition result and outputting the verified result; if the confidence value is less than the second preset threshold, triggering a manual review process.

[0098] In practice, firstly, it is determined whether the confidence value is greater than or equal to a first preset threshold. In a preferred embodiment, the first preset threshold is 0.95. If the confidence value is greater than or equal to 0.95, it indicates that the fused feature information is highly consistent and reliable, and the system directly outputs the vehicle model recognition result without any additional verification.

[0099] Further, if the confidence value is less than 0.95, it is further determined whether the confidence value is greater than or equal to a second preset threshold. In a preferred embodiment, the second preset threshold is 0.7. If the confidence value is less than 0.95 but greater than or equal to 0.7, it indicates that there is a certain degree of uncertainty in the recognition result. The system retrieves vehicle model information stored in the ETC database to verify the visual recognition result. The verification process includes: comparing the vehicle model output by the visual recognition with the vehicle model recorded in the ETC database; if they match, the result is confirmed; if they do not match, a comprehensive judgment is made based on the source with higher confidence (visual or ETC). The verified result is then jointly reported with the visual recognition result.

[0100] Furthermore, if the confidence level is less than 0.7, it indicates that the identification result has significant uncertainty or contradiction, and the system triggers a manual review process. Back-end staff retrieve five simultaneous images of the vehicle through a monitoring terminal and perform manual judgment to ensure the accuracy of the final result.

[0101] This embodiment constructs a three-tiered confidence level output mechanism of "direct reporting - ETC verification - manual review" by setting a first preset threshold (0.95) and a second preset threshold (0.7). The technical principle is that the confidence value C is a function of the probability distribution output by the feature fusion model, reflecting the system's certainty regarding the current identification result. High confidence (C≥0.95) means that the five features are highly consistent and have a high matching degree with known vehicle models, and can be directly adopted; medium confidence (0.7≤C<0.95) means that there is some uncertainty, requiring the introduction of third-party ETC information for joint verification; low confidence (C<0.7) means that the features have significant contradictions or deficiencies, requiring manual intervention. This mechanism balances system processing efficiency and accuracy, avoiding misjudgments and omissions caused by blindly relying on system identification results, and also preventing efficiency bottlenecks caused by excessive use of manual review. With hundreds of billions of transactions, this mechanism can ensure that the annual uncharged fees at a single point are reduced by about 9 million yuan, while keeping the misclassification rate below 0.3%.

[0102] Please see Figure 4 , Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.

[0103] The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.

[0104] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, it enables the processor 502 to execute a vehicle model recognition method based on multi-view cooperative perception.

[0105] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0106] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a vehicle model recognition method based on multi-view cooperative perception.

[0107] The network interface 505 is used for network communication with other devices. Those skilled in the art will understand that the above structure is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. A specific computer device 500 may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0108] The processor 502 is used to run a computer program 5032 stored in a memory to implement the steps of a vehicle model recognition method based on multi-view collaborative perception provided in any of the above method embodiments.

[0109] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0110] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0111] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program. When executed by a processor, the computer program causes the processor to perform the steps of a vehicle model recognition method based on multi-view collaborative perception provided in any of the above method embodiments.

[0112] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code. The computer-readable storage medium can be non-volatile or volatile.

[0113] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0114] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0115] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0116] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0117] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0118] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, this invention is also intended to include these modifications and variations as long as they fall within the scope of the claims of this invention and their equivalents.

[0119] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A vehicle model recognition method based on multi-view collaborative perception, characterized in that, include: In response to receiving a vehicle trigger signal, the system controls the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to acquire a multi-view image sequence. Feature extraction is performed on the multi-view image sequence to obtain the forward feature vector corresponding to the forward telephoto camera, the left feature vector corresponding to the left oblique camera, the right feature vector corresponding to the right oblique camera, the top-view feature vector corresponding to the top-view wide-angle camera, and the rear feature vector corresponding to the rear-view tracking camera. Spatiotemporal correlation and feature fusion are performed on the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector to obtain the fused vehicle probability distribution and confidence value. The vehicle model identification result is determined based on the confidence level value.

2. The method according to claim 1, characterized in that, The process of spatiotemporally associating and fusing the forward feature vector, the left-side feature vector, the right-side feature vector, the top-view feature vector, and the backward feature vector to obtain the fused vehicle model probability distribution and confidence value includes: Determine whether the frontal view of the forward-facing telephoto camera is in a state of obstruction failure; If the frontal view is not in an obstructed state, then based on the preset camera calibration parameters, the forward feature vector, the left feature vector, the right feature vector, the top-view feature vector, and the backward feature vector are spatiotemporally matched to generate a unified vehicle feature object; the unified vehicle feature object is input into the feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value; If the frontal view is in an occlusion failure state, the forward feature vector is removed from the feature vector. Based on the preset camera calibration parameters, the remaining left-side feature vector, right-side feature vector, top-view feature vector, and rear-side feature vector are spatiotemporally matched to generate a unified vehicle feature object. The unified vehicle feature object is then input into the feature fusion model for weighted fusion to obtain the fused vehicle model probability distribution and confidence value.

3. The method according to claim 2, characterized in that, The step of determining whether the frontal view of the forward-facing telephoto camera is in a state of obstruction failure includes: Obtain the top profile of the vehicle, and perform connected component analysis or convex hull analysis on the top profile of the vehicle to determine whether there are abrupt width changes or concave regions in the profile. If the aforementioned width abrupt change or contour concavity area exists, it is determined that the frontal view of the forward telephoto camera is in an occlusion failure state, and the current scene is identified as an abnormal mode where a large vehicle occludes a small vehicle. If there is no abrupt change in width or concave contour area, it is determined that the frontal view of the forward telephoto camera is not in an obstructed state.

4. The method according to claim 3, characterized in that, After identifying the abnormal pattern of a large vehicle obscuring a small vehicle in the current scene, the method further includes: Based on the left-side feature vector, the right-side feature vector, and the top-view feature vector, a multi-target tracking algorithm is used to track the motion trajectories of multiple independent objects across frames. The velocity and acceleration of the motion trajectories of the multiple independent objects are analyzed to determine whether there is a velocity difference or an acceleration difference in the motion trajectories of the independent objects. If the speed difference or acceleration difference exists, then the multiple independent objects are determined to belong to different vehicles; When the obscured vehicle in the different vehicles drives out of the obscured area and reappears in the field of view of the forward telephoto camera, a new forward feature vector of the obscured vehicle is obtained based on the image re-acquired by the forward telephoto camera. The new forward feature vector is spatiotemporally matched with the historical feature vector and the trajectory continuity is verified; the historical feature vector is the left feature vector, right feature vector or top-view feature vector corresponding to the occluded vehicle before or during the occlusion. If the spatiotemporal matching and trajectory continuity verification pass, the feature vector of the occluded vehicle is updated based on the new forward feature vector.

5. The method according to claim 1, characterized in that, The control of the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to acquire multi-view image sequences includes: In response to receiving the vehicle trigger signal, a PPS signal is sent through the SOC to drive the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to perform synchronous exposure at the same time, thereby obtaining a multi-view image sequence with inter-frame jitter of less than or equal to 500 microseconds.

6. The method according to claim 2, characterized in that, The method further includes: Before acquiring the multi-view image sequence, the coordinate system of the overhead wide-angle camera is used as the origin of the world coordinate system to jointly calibrate the forward telephoto camera, the left oblique camera, the right oblique camera, the overhead wide-angle camera, and the rear tracking camera to obtain camera calibration parameters; wherein the reprojection error of the joint calibration is less than or equal to 0.5 pixels.

7. The method according to claim 1, characterized in that, The method further includes: Based on the forward and top views in the multi-view image sequence, the vehicle's three-dimensional contour is calculated using a binocular triangulation algorithm with preset camera calibration parameters to obtain the vehicle's height, width, and length estimates. These estimates are then used as auxiliary features and input into the feature fusion model for feature fusion.

8. The method according to claim 1, characterized in that, The step of determining the vehicle model recognition result based on the confidence value includes: Determine whether the confidence value is greater than or equal to a first preset threshold; If the confidence value is greater than or equal to the first preset threshold, the vehicle model recognition result is output directly; Determine whether the confidence value is less than the first preset threshold and greater than or equal to the second preset threshold; If the confidence value is less than the first preset threshold and greater than or equal to the second preset threshold, the ETC database is retrieved to verify the vehicle identification result, and the verified result is output. If the confidence level is less than the second preset threshold, a manual review process is triggered.

9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method as described in any one of claims 1-8.