Toll vehicle type recognition method and system based on highway traffic images

By performing multi-scale brightness compensation and contour sharpening on the side and top images of the vehicle in the highway toll system, combined with dual-channel feature fusion and pre-training models, the accuracy and efficiency of the single-view model recognition method is solved, and multi-dimensional feature matching is achieved, which improves the robustness of vehicle model recognition and the traffic efficiency of toll stations.

CN119992483BActive Publication Date: 2025-07-18GUIZHOU HUILIANTONG ELECTRONIC COMMERCE SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510468638.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the existing highway toll system, the single-view vehicle image recognition method is susceptible to light conditions and shooting angles, resulting in low vehicle model recognition accuracy, difficulty in feature extraction, and lack of multi-dimensional feature fusion, which affects the robustness and efficiency of vehicle model classification.

Method used

By acquiring the vehicle side and top contour images shot continuously from multiple frames, multi-scale brightness compensation and contour sharpening processing are performed, side and top feature enhancement images are generated, and dual-channel feature fusion is performed. The pre-trained model matching model is used for multi-dimensional feature matching, and similarity comparison is performed in combination with the pre-stored model feature library to determine the model category.

Benefits of technology

It improves the robustness and accuracy of vehicle model identification, improves the model classification capabilities in complex lighting environments, and enhances the traffic efficiency of toll stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992483B_ABST
    Figure CN119992483B_ABST
Patent Text Reader

Abstract

The present invention provides a toll vehicle type recognition method and system based on highway passing images. The method includes: obtaining an image set of passing vehicles, performing multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image, performing contour sharpening processing on the vehicle top profile image to generate a top feature enhanced image. Performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a vehicle multi-dimensional fusion feature map and inputting it into a pre-trained vehicle type matching model to obtain a vehicle type feature matching result output by the vehicle type matching model. Comparing the similarity between the vehicle type feature matching result and a pre-stored vehicle type feature library to determine the toll vehicle type category corresponding to the passing vehicle and generate a vehicle type recognition identifier. The present invention can improve the robustness of vehicle type recognition in complex lighting environments, and at the same time, by integrating the dual-view feature enhancement and multi-dimensional feature matching mechanisms, improve the accuracy of vehicle type classification and the toll passing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and machine learning, and particularly to a toll vehicle type recognition method and system based on highway passing images. Background Art

[0002] In highway toll systems, vehicle type recognition is a core link in determining tolls and improving the efficiency of toll stations. Current mainstream technologies mainly classify vehicle types by collecting single-view images of vehicles (such as only side or top views). However, such methods have significant defects in practical applications: on the one hand, single-view vehicle images are susceptible to the influence of lighting conditions (such as backlighting, shadows), shooting angle deviations, or environmental interferences (such as rainy or foggy weather), resulting in blurred vehicle contours and difficulties in extracting key structural features (such as wheelbase, tire distribution). For example, side images may not clearly show the tire spacing in low-light environments, and top images may cause deformation of the carriage structure due to tilted shooting angles, which will directly affect the accuracy of vehicle type matching. On the other hand, traditional image processing methods usually only perform simple enhancement for a single view (such as global brightness adjustment), lacking the collaborative analysis and fusion of multi-dimensional features (such as the correlation between the side wheelbase and the top carriage structure), resulting in a single feature dimension for vehicle type classification and limited discriminative ability. For example, it may be impossible to distinguish vehicle types with similar carriage heights but different wheelbases relying only on side images, and it is difficult to identify vehicles with the same number of tires but different carriage structures relying only on top images.

[0003] The above technical defects lead to problems such as high feature mis-extraction rates (such as wheelbase measurement deviations due to blurred images), low classification fault tolerance (such as misjudgments caused by lighting changes), and poor adaptability to complex scenarios (such as insufficient classification basis due to non-fusion of multi-view features) in existing vehicle type recognition systems, seriously restricting toll accuracy and passing efficiency. Therefore, there is an urgent need for a vehicle type recognition method to improve the robustness and accuracy of vehicle type classification in complex environments. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a toll vehicle type recognition method and system based on highway passing images. The technical solutions of the embodiments of the present invention are realized as follows:

[0005] On the one hand, embodiments of the present invention provide a toll vehicle type recognition method based on highway passing images, the method comprising:

[0006] Obtaining an image set of a passing vehicle, the image set including multiple frames of continuously captured vehicle side contour images and vehicle top contour images;

[0007] Perform multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image, and perform contour sharpening processing on the vehicle top profile image to generate a top feature enhanced image;

[0008] Perform dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a vehicle multi-dimensional fusion feature map;

[0009] Input the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model to obtain the vehicle model feature matching result output by the vehicle model matching model, where the vehicle model feature matching result includes vehicle wheelbase features, carriage structure features, and tire distribution features;

[0010] Compare the similarity between the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark.

[0011] On the other hand, an embodiment of the present invention provides a computer system, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, the steps in the above method are implemented.

[0012] Advantages of the present invention:

[0013] By obtaining a passing vehicle image set including multi-frame continuously captured vehicle side profile images and vehicle top profile images, the present invention performs multi-scale brightness compensation processing on the side images to generate side feature enhanced images, and at the same time performs contour sharpening processing on the top images to generate top feature enhanced images, effectively solving the problem of blurred features of vehicle images caused by uneven illumination or shooting angles. Specifically, the side feature enhanced image can highlight key structural information such as the vehicle wheelbase and tire distribution, while the top feature enhanced image strengthens the top view features such as the carriage structure and cargo form. The vehicle multi-dimensional fusion feature map generated through dual-channel feature fusion can comprehensively utilize complementary feature information in different dimensions of the vehicle, avoiding the limitations of single-view features. Further, by extracting vehicle wheelbase features, carriage structure features, and tire distribution features through a pre-trained vehicle model matching model, these features are highly correlated with the key discriminant indicators of vehicle model classification. Combining with the pre-stored vehicle model feature library for multi-dimensional feature similarity comparison can accurately match the vehicle model category. The finally generated vehicle model identification mark realizes automatic and multi-dimensional feature analysis and classification of passing vehicles, significantly improving the robustness of vehicle model recognition in complex lighting environments. At the same time, by integrating the dual-view feature enhancement and multi-dimensional feature matching mechanisms, the accuracy of vehicle model classification and the passing efficiency of toll stations are effectively improved.

[0014] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solution of the present invention. Description of the Drawings

[0015] The drawings herein are incorporated into the specification and form a part of this specification. These drawings illustrate embodiments consistent with the present invention and, together with the specification, are used to explain the technical solution of the present invention.

[0016] Figure 1 It is a schematic diagram of the implementation process of a toll vehicle type recognition method based on highway passing images provided by an embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of the hardware entity of a computer system provided by an embodiment of the present invention. Detailed Embodiments

[0018] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0019] An embodiment of the present invention provides a toll vehicle type recognition method based on highway passing images, and this method can be executed by a processor of a computer system. Among them, the computer system may refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, desktop computers, etc. For example, servers deployed in the background of the highway toll system, computer devices deployed at the edge side of highway toll stations, and so on.

[0020] Figure 1 It is a schematic diagram of the implementation process of a toll vehicle type recognition method based on highway passing images provided by an embodiment of the present invention. As Figure 1 shown, this method includes the following steps:

[0021] Step S100: Obtain an image set of passing vehicles, where the image set includes multiple frames of continuously captured side profile images and top profile images of the vehicle.

[0022] In this step, the acquisition of the image set is the basic data source for vehicle model recognition. The vehicle side profile image refers to the static or dynamic image of the overall external profile of the vehicle captured by a camera perpendicular to the vehicle's traveling direction, which can reflect key features such as the number of doors, window layout, and body length; the vehicle top profile image is the image of the vehicle top structure collected from a top-down perspective, used to present three-dimensional space information such as the roof shape, cargo box height, and antenna position. Multiple consecutive frames of shooting emphasizes that the image acquisition device needs to dynamically capture the moving vehicle at fixed time intervals (such as 30 frames per second) to ensure the complete appearance features of the vehicle in motion are captured at different time nodes. For example, in a highway toll booth, when a vehicle passes at a speed of 20 kilometers per hour, the high-speed camera sets deployed on both sides of the lane synchronously trigger the shooting mechanism. The left camera set captures a sequence of side-view images including the cab height and tire size, while the top camera set records the details of the top structure such as the cargo container connection and the roof ventilation device. For special scenarios, such as night passage or rainy and foggy weather, the system can enhance the image clarity through the combination of an infrared supplementary light module and a polarization filter to ensure the continuity of the vehicle contour edges in each frame image. In the data preprocessing stage, it is necessary to perform frame alignment operations on the image set to eliminate the perspective shift caused by vehicle movement and achieve spatio-temporal synchronous matching of multi-angle images through timestamp marking.

[0023] Step S200: Perform multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image, and perform contour sharpening processing on the vehicle top profile image to generate a top feature enhanced image.

[0024] This step involves differential enhancement processing for images from different perspectives. Multi-scale brightness compensation processing refers to optimizing the image brightness distribution using a hierarchical adjustment strategy, specifically including the combined application of global histogram equalization and local contrast-limited adaptive histogram equalization (CLAHE). For example, for the overexposure problem of the window area caused by backlighting in the side profile image, first adjust the overall brightness to the standard range on a global scale, and then perform CLAHE processing on detail areas such as wheels and door handles on a local scale to enhance the dark texture while suppressing highlight overflow. The contour sharpening process strengthens the vehicle top edge features by combining the Laplacian operator and non-linear filtering. For example, for the blurred boundary at the connection between the cargo box and the vehicle head in the top image, a frequency domain high-pass filter is used to extract high-frequency components and superimpose them on the original image, making structural features such as welding seams and waterproof rubber strips more prominent. During the implementation process, the sharpening kernel size needs to be dynamically adjusted according to the image resolution: for high-resolution images (such as 4096×2160 pixels), a 5×5 sharpening template is used to cover a wider neighborhood information; for low-resolution images, a 3×3 template is used to avoid noise amplification. The processed side feature-enhanced image can clearly present micro-features such as tire tread depth and fender curvature, while the top feature-enhanced image can accurately separate the superimposed area of the roof solar panel and the luggage rack.

[0025] Step S300: Perform two-channel feature fusion on the side feature-enhanced image and the top feature-enhanced image to generate a vehicle multi-dimensional fusion feature map.

[0026] Dual-channel feature fusion integrates heterogeneous image data from the spatial dimension and the semantic dimension. In the spatial dimension, a feature registration algorithm based on affine transformation is adopted to project the top image into the coordinate system of the side image, ensuring the spatial consistency of the cargo box length and the roof height. For example, by extracting the spatial correspondence between the vertex of the rearview mirror in the side image and the midline of the skylight in the top image, a three-dimensional projection matrix is established to achieve perspective alignment. In the semantic dimension, a dual-branch feature extraction network is constructed using depthwise separable convolution: the side feature branch focuses on learning linear features such as wheelbase and suspension height, while the top feature branch focuses on regional features such as cargo container segmentation and cooling device layout. After the output feature maps of the two branches are concatenated along the channel dimension and cross-channel information interaction is performed using a 1×1 convolution kernel, a multi-dimensional fusion feature map containing geometric structures and surface textures is formed. During specific implementation, the fusion weights are dynamically adjusted according to the differences between heavy trucks and small passenger cars: for the multi-axle structure of trucks, the contribution of the side feature channels is increased to accurately capture the towing connection points; for the streamlined roof of passenger cars, the top feature channels are enhanced to distinguish the composite structure of the skylight and the solar panel. The finally generated fusion feature map will integrate the tire contact marks from the side view and the cargo box dividing line from the top view to form a joint representation with three-dimensional spatial semantics.

[0027] Step S400: Input the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model to obtain the vehicle model feature matching result output by the vehicle model matching model, where the vehicle model feature matching result includes vehicle wheelbase features, carriage structure features, and tire distribution features.

[0028] The pre-trained vehicle model matching model adopts a multi-task learning architecture based on Transformer, and its core consists of a feature encoder and an attribute decoder. The feature encoder captures the correlation between the wheelbase and the carriage structure through the multi-head self-attention mechanism, for example, identifying the proportional relationship between the distance from the front axle to the cab and the number of compartments in the cargo box; the attribute decoder outputs the wheelbase feature (the quantization value is the pixel distance between the centers of adjacent tires converted to the actual number of meters), the carriage structure feature (classification labels such as flatbed, barn-type, refrigerated), and the tire distribution feature (heat map marking the position of single and double tires and the number of tandem axles) in parallel through a fully connected layer. During model training, a multi-stage transfer learning strategy is adopted: first, the basic feature extraction layer is pre-trained on a public dataset (such as COCO-Vehicles), and then fine-tuned on the labeled data in the highway toll collection scenario. The loss function is designed as a weighted combination of cross-entropy loss and mean squared error loss to balance the optimization requirements of classification and regression tasks. During the actual inference process, for the occluded areas (such as tires covered with mud) in the fused feature map, the model infers the missing information through the context attention mechanism, for example, estimating the number of tires based on the suspension height and the cargo load in the cargo box. The output wheelbase feature is accurate to the centimeter level, the carriage structure feature supports multi-layer label nesting (such as "refrigerated + double side doors"), and the tire distribution feature can distinguish the tread pattern differences between the steering axle and the load-bearing axle.

[0029] Step S500: Compare the similarity between the vehicle model feature matching result and the pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification.

[0030] The pre-stored vehicle model feature library is a structured database, and each record contains the wheelbase range of the standard vehicle model, the carriage type code, and the tire configuration template. The similarity comparison adopts a multi-feature weighted Euclidean distance algorithm: the weight of the wheelbase feature is set to 0.5, the weight of the carriage structure feature is 0.3, and the weight of the tire distribution feature is 0.2. For example, when the wheelbase of a vehicle in the matching result is 3.2 meters (error tolerance ±0.1 meter), the carriage structure is "barn-type double-layer", and the tire distribution is "single front and double rear", the system traverses all records in the feature library, calculates its comprehensive similarity score with each candidate vehicle model, and selects the vehicle model with the highest score and exceeding the threshold (such as 0.85) as the recognition result. For the multi-candidate situation with similar scores (such as the difference in the comprehensive scores of two types of trucks is less than 0.05), start the manual review process and record the conflicting feature items for model iteration and optimization. The finally generated vehicle model identification follows the JT / T 489 standard coding rule, including a 7-digit alphanumeric combination, where the first two digits represent the vehicle category (such as "H" represents trucks), the middle three digits identify the specific model, and the last two digits are the version variant code. This identification will be written into the transaction flow data of the toll collection system and associated with the rate calculation module to complete automatic deduction of fees.

[0031] As an implementation manner, step S200 of performing multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image may include the following steps:

[0032] Step S210: Extract a continuous brightness distribution curve in the vehicle side profile image, where the continuous brightness distribution curve includes a first brightness distribution in the vehicle bottom area and a second brightness distribution in the vehicle top area.

[0033] During this process, the first brightness distribution in the vehicle bottom area refers to the set of pixel brightness values within the vertical linear range from the vehicle tire grounding point to the chassis area, which reflects the influence of ground reflection, shadow occlusion, and metal component reflection characteristics on imaging; the second brightness distribution in the vehicle top area covers the set of brightness data from the upper edge of the window to the highest point of the roof, which is jointly affected by the natural light illumination angle, the diffuse reflection coefficient of the vehicle paint material, and the environmental light occlusion effect. For example, in a strong backlight scenario, the bottom area in the vehicle side profile image may show a low brightness distribution (average value lower than 50, in the 8-bit grayscale range) due to the light absorption characteristics of the ground asphalt, while the top area shows a high brightness concentration (average value higher than 200) due to direct sunlight on the roof, forming a steep brightness gradient between the two. To accurately quantify this difference, the brightness values need to be sampled at each pixel interval along the vertical central axis of the vehicle, and a cubic spline interpolation algorithm is used to generate a smooth continuous curve, where the horizontal axis of the curve represents the normalized height coordinate from the bottom to the top (0.0 to 1.0), and the vertical axis is the normalized brightness value at the corresponding position. Specifically, in implementation, for high-sided vehicles such as container trucks, the first brightness distribution in the bottom area needs to separately distinguish the tire and container support frame areas to avoid the interference of metal bracket reflection on the curve shape; for small passenger cars, the second brightness distribution needs to focus on analyzing the brightness difference between the skylight glass and the roof steel plate to ensure that the curve can accurately represent the transition characteristics of the light-transmitting material and the opaque material.

[0034] Step S220: Generate a dynamic brightness compensation coefficient according to the gradient difference between the first brightness distribution and the second brightness distribution.

[0035] When generating the dynamic brightness compensation coefficient, comprehensively evaluate the spatial correlation between the first brightness distribution and the second brightness distribution. The gradient difference is defined as the integral of the brightness difference between the two distribution curves at the same height coordinate, which reflects the overall light and dark contrast of the upper and lower regions in the side image of the vehicle. During specific calculation, first align the two distribution curves according to the height coordinate, then calculate the difference point by point and accumulate to obtain the total gradient difference value. For example, when it is detected that the average value of the first brightness distribution in the bottom area of the vehicle reaches 180 due to the direct illumination of tunnel lights, while the average value of the second brightness distribution in the top area is only 30 due to the absence of external light sources, the gradient difference value increases significantly (such as the cumulative difference exceeds the preset threshold of 15,000). At this time, a high-strength dynamic brightness compensation coefficient needs to be generated to balance the light and dark differences. This coefficient is generated by mapping through a piecewise linear function. When the gradient difference is relatively low (such as less than 5,000), a fixed coefficient of 1.0 is used to maintain the original image brightness. When the difference is medium (5,000 to 10,000), the compensation intensity increases at a slope of 0.002 per unit. When the difference is relatively high (greater than 10,000), the maximum coefficient is limited to 3.0 to prevent overexposure. For the mixed traffic scenario of vans and flatbed trailers, the piecewise interval needs to be dynamically adjusted according to the vehicle height: for vehicles with a height exceeding 3 meters, an extended threshold range is adopted (such as raising the high-difference threshold to 20,000) to meet the compensation requirements for the large low-brightness area at the top of the container.

[0036] Step S230: Locally enhance the low-brightness areas in the side contour image of the vehicle according to the dynamic brightness compensation coefficient to generate an initial compensation image.

[0037] The implementation of local enhancement depends on the synergistic effect of the dynamic brightness compensation coefficient and the region segmentation result. The low-brightness area refers to the set of pixels whose gray value is lower than 70% of the global average value, and its boundary is determined by the adaptive threshold segmentation algorithm combined with the morphological closing operation. For example, for the side image of a vehicle taken at night, large low-brightness areas (gray value 30 - 50) often form at the connection between the tire and the chassis. At this time, perform gamma correction (Gamma = 0.4) on this area according to the dynamic brightness compensation coefficient of 2.5, which significantly improves the texture visibility and at the same time avoids the loss of details in high-brightness areas (such as the reflective license plate). During the enhancement process, a local window sliding mechanism is adopted, and the window size is dynamically set according to the image resolution: for high-definition images of 4096×2160 pixels, a 51×51 pixel window is used to ensure local consistency; for standard-definition images of 720×480 pixels, a 15×15 pixel window is used to prevent noise amplification. When generating the initial compensation image, metadata needs to be recorded synchronously, including the compensation intensity distribution map of each area and the original brightness histogram transformation parameters, for subsequent multi-scale analysis in the following steps.

[0038] Step S240: Perform multi-scale filtering processing on the initial compensation image to extract edge texture features at different resolutions.

[0039] Multi-scale filtering processing realizes multi-resolution analysis by constructing a Gaussian pyramid. Specifically, the initial compensated image is successively downsampled by 1 / 2 to generate a three-level pyramid (the original image, 1 / 2 scale, 1 / 4 scale), and then the Laplace operator, Sobel edge detection operator, and Gabor filter bank are applied at each level. For example, at the original image scale, a 5×5 Laplace kernel is used to extract fine edges such as door gaps and tire treads; at the 1 / 2 scale, the Sobel horizontal operator is used to detect the continuity of the welding line of the carriage side panel; at the 1 / 4 scale, Gabor filters (wavelength 16 pixels, directions 0°, 45°, 90°, 135°) are applied to extract the periodic texture of the container corrugated board. During this process, the filtering results at each scale are upsampled to the original image size by bilinear interpolation and superimposed to generate a multi-channel edge texture feature map. For the unique concave-convex insulation layer texture of the refrigerated truck, the Gabor filter output at the 1 / 4 scale can effectively enhance the periodic features of 0.5 - 1.2 mm / pixel, while the smooth side panel of the ordinary truck is mainly characterized by the Laplace response at the original image scale.

[0040] Step S250: Adjust the filtering parameters according to the density distribution of the edge texture features to generate the side feature enhanced image.

[0041] The analysis of the density distribution can be achieved, for example, by calculating the number of edge pixels per unit area, which determines the contribution weights of the filters at each scale. For example, when the edge density in the door area is detected to be as high as 120 pixels / cm² (corresponding to the Laplace filtering result at the original image scale), the weight of the Gabor filter at the 1 / 4 scale is reduced to 0.3 to avoid the moiré effect caused by the superposition of high-frequency textures; conversely, for areas where the edge density in the container flat area is lower than 20 pixels / cm², the weight of the Sobel filter at the 1 / 2 scale is increased to 0.7 to enhance the long straight line features. After parameter adjustment, a weighted fusion algorithm is used to combine the multi-scale feature map with the initial compensated image: the details at the original image scale are emphasized in the edge-dense areas (fusion coefficient 0.6), and the contour information at the 1 / 2 scale is enhanced in the flat areas (fusion coefficient 0.8). The finally generated side feature enhanced image is globally optimized by Contrast Limited Adaptive Histogram Equalization (CLAHE), where the Tile grid size is set to 64×64, the contrast limit threshold is 2.0, and the histogram bins are 256. This image can clearly present key features such as the wear degree of the tire tread, the geometric shape of the door handle, and the micro-cracks of the reflective strip, providing high-fidelity input data for subsequent vehicle type recognition.

[0042] As an implementation manner, in step S300, performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a vehicle multi-dimensional fusion feature map, including the following steps:

[0043] Step S310: Perform spatial coordinate transformation on the side feature enhanced image to obtain a set of key contour points in the side image coordinate system.

[0044] In this step, the implementation of spatial coordinate transformation aims to establish a mapping relationship between the vehicle side features and the three-dimensional physical space. The side image coordinate system refers to a two-dimensional pixel coordinate system with the upper left corner of the image as the origin, the horizontal right direction as the positive X-axis direction, and the vertical downward direction as the positive Y-axis direction. The set of key contour points is a set of coordinates of the significant structural feature points of the vehicle side extracted by an edge detection algorithm, including points with geometric invariance such as the vertex of the car door hinge, the center point of the tire in contact with the ground, and the outer edge endpoint of the rearview mirror. For example, for the side feature enhanced image of a van, the Canny edge detector combined with the Hough transform line detection is used to accurately identify the intersection point of the vertical edge line of the cargo container and the horizontal support beam of the chassis as a key contour point. After the coordinate value is optimized by the sub-pixel interpolation algorithm, the accuracy can reach 0.1 pixel. For the side of a bus with a complex curved surface structure, it is necessary to introduce a curvature continuity constraint condition to screen out the curvature extreme points of the arc-shaped window frame of the car door and the corner points of the roof luggage rack and add them to the set of key contour points. During the implementation process, it is necessary to dynamically adjust the edge detection threshold according to the vehicle type: the flatbed trailer has a high proportion of straight edges, and a high threshold (such as the Canny high threshold of 120) is used to suppress the noise points generated by rivet holes; for the streamlined coupe, the threshold is reduced to 60 to retain the weak edge features in the arc transition area.

[0045] Step S320: Perform spatial coordinate transformation on the top feature enhanced image to obtain a set of key contour points in the top image coordinate system.

[0046] The definition method of the top image coordinate system is the same as that of the side image coordinate system, but the set of its key contour points focuses on the geometric features of the vehicle top structure, including the four corner vertices of the sunroof, the installation points of the air conditioner outdoor unit, and the endpoints of the transverse stiffeners of the container top cover, etc. Due to the occlusion of the roof curvature and installation accessories in the top view, a multi-stage feature extraction strategy needs to be adopted: First, fill the gap noise between the solar panel and the roof steel plate through morphological closing operation, then use the Harris corner detector to locate the intersection points of the welding seams of the container top cover, and combine the region growing algorithm to extend to the complete contour line. For example, in the enhanced image of the refrigerated truck top features, after sub-pixel correction of the intersection points of the ventilation grille of the refrigeration unit, the coordinate error does not exceed 0.3 pixels; in the top image of the container truck, the center coordinates of the container corner fitting lock holes can be accurately extracted through the circular Hough transform. For images with perspective distortion (such as the shooting angle with a pitch angle exceeding 15°), it is necessary to first perform projection distortion correction, convert the top image coordinate system to the orthographic projection plane, and then perform key point extraction to ensure that the measurement error of the container length and width is less than 1% of the actual physical size.

[0047] Step S330: Perform three-dimensional space mapping on the set of key contour points in the side image coordinate system and the set of key contour points in the top image coordinate system to generate a vehicle stereo contour model.

[0048] The core of three-dimensional space mapping lies in establishing the spatial correspondence relationship between the side and top key contour points, and it is necessary to calculate the external parameter matrix of the binocular vision system based on a calibration board or a reference object with known dimensions. Specifically, the spatial positions of the tire ground contact center point in the side image coordinate system and the container front edge center point in the top image coordinate system are fitted by the least squares method, and combined with the camera focal length, pixel size, and baseline distance parameters to construct three-dimensional point cloud data. For example, when the coordinate of the left rear tire ground contact center point in the side image is (x1, y1) and the coordinate of the corresponding container left front corner point in the top image is (x2, y2), calculate its three-dimensional coordinates (X, Y, Z) through the principle of triangulation, and iteratively optimize the projection error of all matching point pairs until the root mean square error is less than 0.5 pixels. The generated vehicle stereo contour model is composed of triangular mesh patches, and the vertex density of the mesh is adaptively adjusted according to the curvature change: in high-curvature regions (such as the rearview mirror surface), dense vertices with a 5-mm spacing are used, and in flat regions (such as the container side), sparse sampling is performed at a 20-mm spacing. For special vehicles (such as oil tankers), it is necessary to additionally introduce a cylindrical parametric model constraint to fit the point cloud of the tank part into a parametric surface with a constant radius and continuously changing axis direction to ensure that the calculation accuracy of the tank volume reaches more than 98%.

[0049] Step S340: Extract the vehicle surface curvature features according to the curvature distribution of each contour point in the vehicle stereo contour model.

[0050] The calculation of the curvature distribution is based on the local surface differential geometric properties of each vertex in the triangular mesh model, and the combination of Gaussian curvature and mean curvature is used to characterize the concave and convex characteristics of the vehicle surface. For each vertex, the principal curvatures k1 and k2 are calculated through the change rate of the normal vectors of its adjacent triangular patches, and then the Gaussian curvature K = k1 * k2 and the mean curvature H = (k1 + k2) / 2 are obtained. For example, the Gaussian curvature of the flat area of the side panel of the van cargo container approaches zero, and the mean curvature is also close to zero; while the arc-shaped fairing on the top of the refrigerated truck shows negative Gaussian curvature (K < 0) and non-zero mean curvature (H ≈ 0.03 / mm). To extract the curvature features, multiple-level thresholds need to be set: the area where the absolute value of the Gaussian curvature is greater than 0.05 / mm² is marked as the high-curvature feature area (such as the depression of the door handle), the area between 0.01 and 0.05 / mm² is defined as the medium-curvature transition area (such as the tread pattern of the tire), and the area less than 0.01 / mm² is regarded as the low-curvature flat area (such as the window glass). During the implementation process, for the tiny wrinkles of the metal stamping parts (the curvature fluctuation range is 0.02 - 0.08 / mm²), the curvature gradient histogram is used to statistically analyze their distribution rules and filter out the outliers caused by point cloud noise (the points with gradient mutation exceeding 3σ).

[0051] Step S350: Superimpose and fuse the vehicle surface curvature features and the edge texture features of the side feature enhanced image to generate the vehicle multi-dimensional fusion feature map.

[0052] The superimposition and fusion process adopts a hybrid strategy combining channel-weighted stitching and spatial attention mechanism. First, the vehicle surface curvature features are converted into a two-dimensional feature map with the same resolution as the side feature enhanced image, and the channel value of each pixel point corresponds to the three-dimensional curvature attribute (the linear combination of Gaussian curvature and mean curvature) at its projection position. Subsequently, the spatial attention module is used to calculate the significance weights of each pixel in the edge texture feature map: for the edge pixels corresponding to the high-curvature area (such as the door handle), a fusion weight of 0.7 is assigned to strengthen the three-dimensional geometric characteristics; for the low-curvature area (such as the side panel of the cargo container), it is reduced to 0.3, focusing on retaining the two-dimensional texture details. For example, after the ripple texture of the insulation layer (edge density 60 pixels / cm²) in the side feature enhanced image of the refrigerated truck and the Gaussian curvature feature (K = -0.04 / mm²) of the top fairing are fused, the generated feature map can simultaneously present the ripple period (two-dimensional attribute) and the fairing arc (three-dimensional attribute). The final multi-dimensional fusion feature map is uniformly scaled to the standard size (such as 1024 × 512 pixels) through the bicubic interpolation algorithm and is subjected to normalization processing (the channel values are scaled to the [0, 1] interval), providing input data with both geometric structure and surface details for the subsequent vehicle model matching model.

[0053] As an implementation manner, in step S400, inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model to obtain a vehicle model feature matching result output by the vehicle model matching model, including the following steps:

[0054] Step S410: Divide the vehicle multi-dimensional fusion feature map into multiple local feature regions, and each local feature region corresponds to one feature type among the vehicle wheelbase feature, the carriage structure feature, or the tire distribution feature.

[0055] The division of the vehicle multi-dimensional fusion feature map needs to perform spatial region annotation based on the prior knowledge of the vehicle physical structure. The local feature region refers to a specific range delimited by a rectangular box or a semantic segmentation mask in the feature map. Among them, the vehicle wheelbase feature region covers the horizontal projection range from the center point of the front axle of the vehicle to the center point of the rear axle. The carriage structure feature region includes the vertical section and the top arc structure at the connection between the cab and the cargo box. The tire distribution feature region focuses on the spatial distribution of the grounding points of each tire and the tread pattern. For example, for the fusion feature map of a container truck, the wheelbase feature region is delimited along the longitudinal beam of the frame as a rectangular region with a width of 50 pixels and a length of 80% of the total length of the feature map, accurately covering the wheelbase at the connection between the tractor and the trailer. The carriage structure feature region uses a polygon mask to mark the hexahedron contour of the refrigerated container and the position of the ventilation openings. The tire distribution feature region detects the range of 50×50 pixels around the grounding points of each tire through a sliding window to ensure that the measurement accuracy of the tire spacing of the dual-tire assembly structure reaches ±2 pixels. During the division process, the region size needs to be dynamically adjusted according to the vehicle type: the wheelbase feature region of a flatbed trailer needs to be extended to include multiple sets of assembled axles, while the carriage structure feature region of a small passenger car needs to be shrunk to the joint between the roof skylight and the trunk lid.

[0056] Step S420: Perform feature vectorization processing on each local feature region to generate a set of local feature vectors.

[0057] Feature vectorization is achieved by combining depthwise separable convolution and global average pooling, which converts a two-dimensional feature map into a numerical vector of a fixed dimension. For the vehicle wheelbase feature region, a 3×3 convolution kernel is used to extract the feature responses in the horizontal direction, capturing the relative distance between adjacent axles and the change in suspension height; for the carriage structure feature region, a 5×5 dilated convolution is used to expand the receptive field, effectively identifying the spatial relationship between the cargo compartment partition and the arc of the roof; for the tire distribution feature region, a fusion descriptor of Histogram of Oriented Gradients (HOG) and Local Binary Pattern (LBP) is used to quantify the directionality of the tread pattern and the degree of shoulder wear. For example, after global average pooling, the convolution output of the wheelbase feature region of a refrigerated truck generates a 128-dimensional vector, and the peak component corresponds to the 3.2-meter feature of the distance between the third and fourth axles; in the structure feature vector of a container cargo container, the values of the 56-72nd dimensions are significantly higher than other parts, representing the grid density of the corner fitting reinforcement plate of the cargo container; while the distribution feature vector of dual tires mounted side by side forms a 247-dimensional hybrid descriptor through the combination of the 9-direction gradient histogram of HOG and the 59 texture patterns of LBP.

[0058] Step S430: Input the set of local feature vectors into the multi-level convolutional network in the vehicle model matching model to obtain the feature response maps output by each convolutional layer.

[0059] The multi-level convolutional network consists of three parallel branches, which process the wheelbase, carriage, and tire feature vectors respectively. The wheelbase branch uses a Temporal Convolutional Network (TCN) to capture the sequential dependencies of multi-axle vehicles along the wheelbase direction; the carriage branch uses a Residual Dense Network (RDN) to enhance the representational ability of complex structures through cross-layer feature reuse; the tire branch deploys a Graph Convolutional Network (GCN) to construct tire positions as graph nodes and model the spatial topological relationship. For example, when inputting the wheelbase feature vector of an 8-axle heavy-duty tractor, in the feature response map output by the fifth layer of the TCN, the third channel shows significant activation (response value 0.92) at the position of the fourth axle, indicating that this is the steering axle feature; when the RDN processes the structure features of a refrigerated cargo container, the skip connections between the dense blocks increase the response intensity of the ventilation opening area in the 12th layer feature map by 40%; the GCN constructs edges between each tire node and its adjacent tires, and after three-layer graph convolution iteration, the feature similarity of the dual-tire mounting nodes reaches 0.87, significantly higher than 0.32 of the single-tire nodes.

[0060] Step S440: Determine the wheelbase matching degree corresponding to the vehicle wheelbase feature, the carriage matching degree corresponding to the carriage structure feature, and the tire matching degree corresponding to the tire distribution feature according to the activation region distribution of the feature response map.

[0061] The analysis of the activation region distribution adopts a strategy combining spatial weighted pooling and saliency detection. The wheelbase matching degree is calculated by the degree of coincidence between the intervals of the activation peaks at each axis position in the output feature map of the TCN and the actual calibration value. For example, if it is detected that there are three peaks with a standard deviation exceeding 2.5σ at 3.2 meters, 7.8 meters, and 11.4 meters in the feature response map, then it is matched with the three-axis spacing template in the vehicle type library by Dynamic Time Warping (DTW) to obtain a wheelbase matching degree of 0.89; the carriage matching degree calculates the cosine similarity of the high-dimensional features output by the RDN. For example, the included angle between the ventilation port layout of the refrigerated container and the 512-dimensional vector of the template is 18°, corresponding to a similarity of 0.95; the tire matching degree needs to comprehensively consider the consistency of the tire tread direction (the chi-square distance of the HOG histogram) and the distribution topology matching degree (the Euclidean distance of the GCN node embedding). For example, the cumulative value of the difference in the tire tread direction of a certain vehicle is 125, and the topological distance is 0.34. After normalization, a tire matching degree of 0.76 is obtained.

[0062] Step S450: Perform weighted fusion on the wheelbase matching degree, the carriage matching degree, and the tire matching degree to generate the vehicle type feature matching result.

[0063] The weighted fusion adopts adaptive weight allocation based on feature importance. Among them, the initial value of the wheelbase weight is 0.5, the carriage is 0.3, and the tire is 0.2, but it is dynamically adjusted according to the vehicle type: when it is detected that the structure of the cargo box is complex (such as the multi-layer partition of a refrigerated truck), the carriage weight is increased to 0.4; for multi-axis vehicles (such as 5-axis tank trucks), the wheelbase weight is increased to 0.6. The fusion formula is: comprehensive matching degree = wheelbase matching degree × W1 + carriage matching degree × W2 + tire matching degree × W3. For example, for a container truck, the wheelbase matching degree is 0.92 (W1 = 0.5), the carriage matching degree is 0.88 (W2 = 0.3), and the tire matching degree is 0.81 (W3 = 0.2), and the comprehensive value is 0.5×0.92 + 0.3×0.88 + 0.2×0.81 = 0.892; if it is an 8-axis flatbed trailer, due to the increase in the number of axes, the wheelbase weight is increased to 0.6, then W1 = 0.6, W2 = 0.25, and W3 = 0.15 during calculation. The calculation result is mapped to the [0, 1] interval by the Sigmoid function. The final vehicle type feature matching result includes each sub-item matching degree and the comprehensive value, and marks the confidence interval (such as 0.892 ± 0.03).

[0064] Based on this, as an implementation, step S500 of comparing the similarity between the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark includes the following steps:

[0065] Step S510: Extract the vehicle model template data from the vehicle model feature library, where the vehicle model template data includes a template wheelbase range, a template carriage structure template, and a template tire distribution template.

[0066] The vehicle model template data is stored in a structured form. Each record includes a template wheelbase range (such as multiple intervals like 3.2 - 3.5 meters, 7.6 - 7.9 meters, etc.), a template carriage structure template (a 512-dimensional feature vector), and a template tire distribution template (a topological graph including the number of tires, tire position coordinates, and tire pattern codes). For example, the template of the refrigerated truck numbered HT-45C in the vehicle model library includes three wheelbase ranges [3.18, 3.22], [7.75, 7.85], [11.3, 11.5] meters. The carriage structure template corresponds to a wavy cargo container feature vector with a ventilation port spacing of 600 mm. The tire distribution template defines 6 groups of twin-tire assembly nodes and a standard value of the tire pattern direction angle of 32°. The construction of the template data requires taking the average value of the feature data of 200 vehicles of the same model through actual measurement and calculating the allowable fluctuation range of ±3σ to ensure coverage of manufacturing tolerances and measurement errors.

[0067] Step S520: Calculate the first similarity between the wheelbase matching degree and the template wheelbase range, calculate the second similarity between the carriage matching degree and the template carriage structure template, and calculate the third similarity between the tire matching degree and the template tire distribution template.

[0068] The first similarity is calculated by combining the interval overlap ratio and Gaussian weighting: If the measured wheelbase falls within the template range, the basic similarity is 1.0; if it deviates, it is calculated according to the formula exp(-(d / σ) 2 ), where d is the deviation amount and σ is the template standard deviation. For example, the measured wheelbase of 3.25 meters exceeds the upper limit of the third section of the HT-45C template, which is 11.5 meters, by 0.05 meters. Assuming σ = 0.1 meter, the similarity of the third section is exp(-(0.05 / 0.1) 2 ) = 0.778; the second similarity measures the direction consistency between the measured carriage feature vector and the template vector through cosine similarity. When the angle between the measured vector and the template is 15°, the similarity is cos(15°) = 0.966; the third similarity needs to comprehensively consider the tire number matching degree (such as getting 1.0 when the measured 12 tires vs the template 12 tires), the topological difference of the tire positions (the average Euclidean distance of the node coordinates is 0.2 meters, and after Sigmoid conversion, it is 0.82), and the tire pattern direction difference (the histogram intersection ratio is 0.75), and is obtained as 0.86 after weighted averaging.

[0069] Step S530: Determine the comprehensive matching degree between the passing vehicle and each vehicle type template data according to the weighted sum of the first similarity, the second similarity, and the third similarity.

[0070] The weighting coefficients are dynamically adjusted according to vehicle type classification: for freight vehicles, the wheelbase is 0.5, the carriage is 0.3, and the tires are 0.2; for passenger vehicles, more emphasis is placed on the comfort characteristics of the carriage, and the weights are set as the wheelbase 0.4, the carriage 0.4, and the tires 0.2. For example, for a vehicle suspected to be HT-45C, the three similarities are 0.92 (wheelbase), 0.94 (carriage), and 0.88 (tires) respectively. The comprehensive matching degree under the freight weight is 0.5×0.92 + 0.3×0.94 + 0.2×0.88 = 0.918; if this vehicle is a passenger version, the calculation according to the passenger weight is 0.4×0.92 + 0.4×0.94 + 0.2×0.88 = 0.924. After the calculation results are sorted in descending order, a candidate list is generated, such as HT-45C (0.918), HT-46B (0.895), GT-32D (0.872), and the sub-item matching degrees are marked for manual review reference.

[0071] Step S540: Use the vehicle type number corresponding to the vehicle type template data with the highest comprehensive matching degree as the toll vehicle type category, and generate a vehicle type identification mark including the vehicle type number.

[0072] The conversion of the vehicle type number shall follow the JT / T 489-2019 standard, where the first two letters represent the major vehicle type (such as HT represents a refrigerated truck), the middle two digits are the load capacity level (45 represents 45 tons), and the last letter is the version number (C is the third version). The vehicle type identification mark is encapsulated in JSON format, including fields: {"vehicle type number":"HT-45C", "matching degree":0.918, "timestamp":"2023-08-20T14:23:05Z", "wheelbase details":[{"position":"First wheelbase","measured value":3.21,"template range":[3.18,3.22]},... ]}. This mark is written into the toll system transaction record after being Base64 encoded, and at the same time, it triggers the rate calculation module to call the rate of 0.45 yuan / ton·kilometer corresponding to HT-45C. For vehicles with a comprehensive matching degree lower than 0.85 (such as 0.832), the system automatically triggers the high-definition reshooting process and pushes the queue to be reviewed to the toll collector's console.

[0073] As an implementation method, the method further includes a pre-training step of the vehicle type matching model, including the following steps:

[0074] Step S10: Obtain a historical passing vehicle training set, where the historical passing vehicle training set includes multiple groups of vehicle image samples and corresponding labeled vehicle type feature data.

[0075] The construction of the historical passing vehicle training set needs to cover diverse data under different lighting conditions, shooting angles, and vehicle models. Vehicle image samples refer to the original image sequences collected by the highway toll lane monitoring system. Each group of samples contains 10 consecutive frames of vehicle side-view images and 5 frames of top-view images, and the frame rate is set to 30fps to ensure that the motion blur is controlled within a manageable range. The labeled vehicle model feature data adopts a structured tag format, including vehicle wheelbase features (measured values of the three-axis spacing accurate to the millimeter level), carriage structure features (classification codes based on the JT / T 489 standard, such as "HT-45C" representing a 45-ton refrigerated truck), and tire distribution features (pixel coordinates of each tire contact point and tire tread type code). For example, for a certain batch of training data collected in 2019, the vehicle image samples include high-light interference images in the window area caused by water film reflection in rainy weather. In its labeled data, the wheelbase feature is labeled as "3.21±0.03 m, 7.85±0.05 m, 11.42±0.07 m", the carriage structure feature is labeled as "double-layer partition refrigerated container", and the tire distribution feature is labeled as "12R22.5 tubeless tire, twin tire spacing 200±5 mm". The training set needs to eliminate duplicate vehicle data through a spatio-temporal deduplication algorithm and be divided into a training subset, a validation subset, and a test subset according to a ratio of 8:1:1 to ensure the generalization ability of the model.

[0076] Step S20: Perform image segmentation processing on each group of vehicle image samples to extract vehicle side training images and vehicle top training images.

[0077] The image segmentation processing adopts an instance segmentation algorithm based on Mask R-CNN and combines vehicle three-dimensional geometric constraint conditions to separate the target area. The segmentation of vehicle side training images needs to generate the minimum bounding rectangle along the vehicle longitudinal axis and cut off the guardrails, green belts, and other vehicle interference areas in the background; for vehicle top training images, through top-down projection transformation, the trapezoidal distortion area in the original image is corrected to a rectangular area and cropped within the range of 10 pixels expanded outside the roof outline boundary. For example, for an articulated truck with a trailer, the segmentation algorithm first identifies the joint point between the tractor cab and the trailer container, generates a side mask along the extension direction of the trailer longitudinal beam to ensure that the container side panel and the tractor rear wheels are both included in the side training image; during the top image segmentation, a fish-eye correction model is used to eliminate the bending of the container top cover edge caused by the wide-angle lens, and the inclined view is converted to a front projection view through a perspective transformation matrix to restore the length-width ratio of the container to the actual value of 1:2.5. For low-contrast areas in night images (such as the tire contour of a black vehicle body), an auxiliary edge enhancement module needs to be enabled, combined with Canny edge detection and morphological closing operations to fill the broken boundaries to ensure the integrity of the closed area of the segmentation mask.

[0078] Step S30: Perform noise suppression and distortion correction on the vehicle side training images to generate a side training sample set.

[0079] Noise suppression adopts the cascaded processing of the Non-Local Means (NLM) algorithm and Block-Matching and 3D filtering (BM3D) to hierarchically eliminate high-frequency noise (such as raindrop noise) and low-frequency noise (such as sensor thermal noise). For example, for the side training images collected in rainy and foggy weather, first apply the NLM algorithm (search window 21×21 pixels, similarity window 7×7 pixels, filtering parameter h = 15) to smooth the raindrop texture, and then use BM3D (block size 8×8, hard threshold shrinkage) to remove the remaining salt-and-pepper noise, improving the SSIM (structural similarity index) of the tire tread pattern from 0.65 to 0.89. Distortion correction needs to construct a radial distortion model based on the camera calibration parameters and use the k1, k2, k3 coefficients in the Brown-Conrady model to perform reverse compensation for the stretching of the image edges. Specifically, in implementation, for a surveillance camera with a focal length of 8 mm, when the straightness deviation of the longitudinal beam of the vehicle frame in the side image exceeds 3 pixels, use the correction parameters of k1 = -0.25 and k2 = 0.06 to control the straightness error within 0.5 pixels. The corrected side training sample set also needs to be subjected to brightness normalization processing, adjusting the mean of the image grayscale histogram to 128±10 and expanding the standard deviation to 45±5 to enhance the discernibility of the details in the dark areas.

[0080] Step S40: Perform shadow elimination and contrast enhancement on the vehicle top training images to generate a top training sample set.

[0081] Shadow elimination can be achieved, for example, through the inversion of the physically based rendering equation. Specifically, a Light Estimation Network (LEN) is used to predict the direction and intensity of the main light source in the scene. Subsequently, a shadow mask generation algorithm is employed to separate the projected shadow areas on the roof surface (such as the strip-like shadows formed by the gaps between containers). For example, for the top image taken at noon, LEN outputs a light source azimuth angle of 120°, an elevation angle of 60°, and an intensity of 85000 lux. Based on this, the theoretical projection range of the shadow is calculated, and the brightness of the shadow area is increased to be the same as that of the non-shadow area through the Poisson image editing algorithm. Contrast enhancement uses an improved Retinex algorithm to decompose the image into a reflection component and an illumination component: the reflection component extracts the surface texture through multi-scale Gaussian filtering (σ = 15, 80, 200 pixels), and the illumination component compresses the dynamic range through gamma correction (γ = 0.6). After implementation, the contrast between the rust area on the container top and the waterproof strip is increased from the original value of 1.8 to 4.2, and the edge sharpness is increased by 30%. The processed top training sample set needs to be verified for geometric alignment. The feature point matching algorithm (such as SIFT) is used to check the projection consistency between the corrected image and the 3D point cloud model to ensure that the coordinate error of the container corner points does not exceed ±2 pixels.

[0082] Step S50: Input the side training sample set and the top training sample set into the initial convolutional neural network for multi-stage training, and adjust the network weight parameters until the error between the output features of the initial convolutional neural network and the labeled vehicle model feature data is lower than the preset threshold, thereby generating the vehicle model matching model.

[0083] The initial convolutional neural network can, for example, adopt a two-stream architecture design. The side and top image branches respectively use ResNeXt-101 and DenseNet-201 as the backbone networks, and fuse features through a cross-modal attention mechanism. The multi-stage training includes three stages: In the first stage, the parameters of the top branch are frozen, and only the wheelbase regression layer and the carriage classification layer of the side branch are trained. The learning rate is set to 1e-4, the batch size is 32, and the smooth L1 loss function is used to optimize the wheelbase prediction error; In the second stage, the top branch is unfrozen, and the tire distribution prediction module of the two-stream network is jointly trained. The learning rate is reduced to 5e-5, the batch size is 16, and the loss function is changed to a mixed loss (wheelbase MSE loss + carriage cross-entropy loss + tire contrast loss); In the third stage, all parameters are enabled for end-to-end fine-tuning, and a curriculum learning strategy is introduced. The training data is gradually increased from easy to difficult according to the sample difficulty, and the learning rate adopts a cosine annealing schedule (initial value 3e-5, period 20 epochs). During the training process, the comprehensive error (wheelbase MAE + carriage F1 score + tire IoU) of the validation subset needs to be evaluated every epoch. When the error decrease amplitude is less than 0.1% for 5 consecutive epochs, the early stopping mechanism is triggered. For example, in a certain training, the average absolute error (MAE) of the wheelbase of the initial model on the test subset is 38 mm, and it drops to 12 mm after 120 epochs of training; The macro-average F1 score of the carriage structure classification is improved from 0.76 to 0.93; The IoU (Intersection over Union) of the tire distribution is optimized from 0.65 to 0.88. Finally, the model is saved as a vehicle type matching model after meeting the preset thresholds (wheelbase MAE ≤ 15 mm, F1 ≥ 0.9, IoU ≥ 0.85). When the model is deployed, the floating-point model is quantized to the INT8 format through the TensorRT engine, and the inference latency is reduced from 210 ms to 45 ms, meeting the real-time processing requirements of the toll station.

[0084] As an implementation manner, in step S20, the vehicle image samples in each group are subjected to image segmentation processing to extract the vehicle side training image and the vehicle top training image, including the following steps:

[0085] Step S21: Perform background separation processing on the vehicle image samples to obtain a vehicle main body region mask.

[0086] The background separation process can be implemented, for example, through a deep learning-based semantic segmentation network. The vehicle body area mask refers to a set of vehicle pixel positions represented in the form of a binary image, where foreground pixels (value 1) represent the vehicle entity area, and background pixels (value 0) represent irrelevant areas such as roads, guardrails, or the sky. Specifically, when implementing, the DeepLabv3+ network architecture is adopted, and its backbone network is Xception-65. After initializing with the model weights pre-trained on the highway vehicle dataset, the input image extracts multi-scale features through the spatial pyramid pooling module and outputs a mask image with the same resolution as the input. For example, for an image sample containing a container truck, the network can accurately separate the container side panel from the background of the rear green belt. Even if there is a light blue coating on the container surface that is similar in color to the sky, the vehicle body boundary can still be recognized through high-level semantic features, ensuring that the mask edge error does not exceed 3 pixels. For the segmentation of the headlight highlight area in night images, an edge-sensitive weight needs to be added to the loss function to enhance the ability to distinguish the taillight contour from the reflective road surface.

[0087] Step S22: Divide the vehicle side area and the vehicle top area according to the geometric center position of the vehicle body area mask.

[0088] The calculation of the geometric center position is based on, for example, the minimum bounding rectangle of the vehicle body area mask, and the center point coordinates are the pixel coordinates of the intersection of the rectangle diagonals. The vehicle side area is defined as a rectangular range extending vertically downward from the center point to the vehicle bottom grounding line, covering the tires, chassis, and door structures; the vehicle top area is the area extending vertically upward from the center point to the highest point of the vehicle roof, including components such as the container top cover, air conditioner outdoor unit, and antenna. For example, for a refrigerated truck with a height of 4.2 meters, the geometric center of the mask is located at the 380th pixel of the image vertical coordinate (the total height of the image is 720 pixels). The side area is delimited as the range from the 380th to the 720th pixel of the vertical coordinate, and the top area is the range from the 0th to the 380th pixel of the vertical coordinate. For special vehicle types such as articulated trailers, the division ratio needs to be dynamically adjusted according to the trailer connection point: the tractor part is divided according to the standard ratio, and the trailer part is divided again based on the geometric center of its independent mask to ensure complete coverage of the side and top areas of the front and rear sections of the container.

[0089] Step S23: Perform edge tracing on the vehicle side area to extract a continuous and closed side contour line.

[0090] Edge tracking can, for example, adopt an improved boundary tracking algorithm. Starting from the starting point at the lower left corner of the side area mask of the vehicle, it traverses the pixel neighborhood in the counterclockwise direction, records the sequence of boundary points that meet the 8-connectivity condition, and finally generates a closed contour line composed of polygonal broken lines. For example, when processing the side area of a refrigerated truck, the algorithm starts tracking from the grounding point of the left rear tire, rises along the vertical line of the container side panel to the arc transition area of the roof, then extends horizontally to the right front fender, and finally closes downward to the starting point, forming a polygonal contour line with 256 vertices. To ensure the smoothness of the contour line, the Douglas-Peucker algorithm is used to compress the initial tracking result, reducing the number of redundant points to 30% of the original data while keeping the contour shape error less than 1.5 pixels. For the inner edges generated by concave structures such as door handles and reflective strips, concave points are identified by calculating the curvature change rate, and control points are inserted into the contour line to retain the detailed features.

[0091] Step S24: Determine the cropping boundary of the vehicle side training image according to the curvature change points of the side contour line.

[0092] The curvature change points refer to the local extreme points on the contour line where the curvature value exceeds the preset threshold, reflecting the geometric mutation positions of the vehicle side structure. The curvature is calculated using the discrete differential geometry method. For each point Pi in the vertex sequence of the contour line, the ratio of the chord length to the arc length formed by its k previous and next points (k = 5) is calculated. When the ratio exceeds the threshold of 0.15, it is marked as a curvature change point. For example, in the side contour line of a container truck, the curvature ratio at the connection between the front edge of the container and the cab reaches 0.22 and is identified as a key cropping boundary point; the periodic concave and convex structures of the corrugated board of the refrigerated truck container generate evenly spaced curvature change points, and the spacing distance corresponds to the corrugation spacing (about 200 mm). The cropping boundary is determined by the horizontal projection of the curvature change points: a straight line is extended vertically to the image boundary to form a rectangular area containing the entire vehicle side. If the spacing between adjacent curvature change points is too small (such as less than 50 pixels), a B-spline curve is used for fitting to generate a smooth transition boundary to avoid the jagged cropping affecting subsequent feature extraction.

[0093] Step S25: Segment the vehicle image sample according to the cropping boundary to generate the vehicle side training image and the vehicle top training image.

[0094] The splitting operation can be achieved, for example, by combining image cropping and mask bitwise operations: the vehicle side training image is the logical AND result of the pixel set within the side cropping boundary of the original image and the vehicle body region mask; the vehicle top training image is the intersection of the top region cropping result and the mask. For example, for a van image sample, the side cropping boundary is a rectangular region from 120 pixels from the left to 880 pixels from the right and from 200 to 680 pixels in the vertical coordinate, generating a side training image with a resolution of 760×480 pixels, where only the door, tire, and container side panel regions are retained; the top training image is cropped to a 320×1280 pixel region from 0 to 200 pixels in the vertical coordinate, focusing on the container top cover and deflector structure. For vehicles with abnormal heights (such as engineering forklifts), a dynamic scaling strategy is adopted: if the height of the cropped image exceeds the standard size (such as 512 pixels), it is scaled down proportionally to the target size and the edge blank areas are filled, while maintaining the aspect ratio and ensuring the integrity of key structures.

[0095] As an implementation, in step S30, performing noise suppression and distortion correction on the vehicle side training image to generate a side training sample set includes the following steps:

[0096] Step S31: Detect high-frequency noise regions in the vehicle side training image to generate a noise distribution map.

[0097] High-frequency noise regions refer to interference signals in the image with a spatial frequency higher than the vehicle structure features, such as those caused by sensor thermal noise, raindrops, or dust adhering to the lens. The noise distribution map is generated by wavelet transform decomposition: the image is decomposed by 3-layer Haar wavelet transform, and the sum of the absolute values of the high-frequency subbands (HL, LH, HH) is extracted as the noise intensity index. For example, in a vehicle side training image collected on a rainy day, the response value of raindrop noise in the HH subband (diagonal high frequency) exceeds 120 (8-bit grayscale range), forming a scattered noise region; while the sensor noise is evenly distributed in the HL subband (horizontal high frequency), and the intensity value remains between 30-50. The noise distribution map is represented in the form of a heat map, and the closer the color is to red, the higher the noise intensity. For periodic noise (such as LED stroboscopic stripes), additional Fourier transform needs to be performed to identify specific frequency components and mark the corresponding frequency domain coordinate regions in the distribution map.

[0098] Step S32: Adjust the window size of the adaptive filter according to the density value of the noise distribution map and smooth the high-frequency noise region.

[0099] The window size of the adaptive filter is negatively correlated with the noise density value. For example, in areas where the density value is higher than the threshold (such as 150 noise points per square centimeter), a 3×3 window is used to retain details; in areas with medium density values (50 - 150 points per square centimeter), a 5×5 window is used to balance denoising and blurring; in low-density areas (below 50 points), a 7×7 window is applied to achieve strong smoothing. For example, for a large rusty area on the side panel of a container (noise density of 280 points per square centimeter), bilateral filtering with a 3×3 window is used to suppress isolated noise points while retaining the rust texture; while for the evenly painted area of the car door (density of 40 points), 7×7 Gaussian filtering is used to eliminate sensor noise. During the implementation process, the local noise density needs to be calculated dynamically: the sum of pixel values of the noise distribution map in this area is statistically calculated using a sliding window (default 32×32 pixels), and then normalized by the area to obtain the density value. For boundary areas (such as the contact area between the tire and the ground), an asymmetric window is used to avoid cross-region mixing. The horizontal size of the window remains 5 pixels, and the vertical size is reduced to 3 pixels to adapt to the edge trend.

[0100] Step S33: Extract the perspective distortion parameters in the vehicle side training image, where the perspective distortion parameters include the horizontal tilt angle and the vertical scaling ratio.

[0101] For example, the perspective distortion parameters are calculated with the assistance of a calibration board or estimated based on the geometric constraints of the vehicle structure. The horizontal tilt angle refers to the angle between the longitudinal axis of the vehicle and the horizontal axis of the image, which is obtained by calculating the slope after detecting the straight line of the vehicle frame longitudinal beam through the Hough transform; the vertical scaling ratio is the ratio of the actual vehicle height to the pixel height of the vehicle in the image, and it needs to be calibrated in combination with known calibration objects (such as reflective strips of standard size) or laser ranging data. For example, when the straight line equation of the vehicle longitudinal beam edge is detected as y = 0.82x + 120, the horizontal tilt angle is arctan(0.82) = 39.3°; if the actual measured vehicle height is 3.5 meters and the height in the image is 700 pixels, then the vertical scaling ratio is 3.5 / 700 = 0.005 meters per pixel. For scenarios without calibration objects, the vanishing point estimation method is adopted: the horizontal tilt angle is determined by detecting the intersection points of the images of two parallel lane lines, and the scaling ratio is deduced based on the prior knowledge of the vehicle width (such as the standard truck width of 2.55 meters).

[0102] Step S34: Perform affine transformation correction on the vehicle side training image according to the horizontal tilt angle, and adjust the image ratio of the vehicle side training image according to the vertical scaling ratio to generate the side training sample set.

[0103] For example, the affine transformation correction is achieved by combining a rotation matrix and a translation matrix. That is, with the center of the image as the origin, the image is rotated by the opposite of the horizontal tilt angle (e.g., -39.3°) around the origin to align the longitudinal axis of the vehicle with the horizontal axis of the image. For the vertical scaling adjustment, the bilinear interpolation algorithm is applied to scale the image vertically to the actual ratio. For example, if 700 pixels in the vertical direction of the original image correspond to 3.5 meters and the scaling ratio is 0.005 meters / pixel, and it is required to convert to the standard ratio of 0.004 meters / pixel (corresponding to a height of 875 pixels), then the vertical stretching coefficient is 875 / 700 = 1.25. The corrected image needs to be subjected to edge padding processing, and the mirror replication method is used to expand the blank area generated by the rotation to ensure the integrity of the vehicle body without shearing. The finally generated side training sample set includes the geometrically corrected images and the corresponding distortion parameter metadata, which are used for spatial feature alignment and scale normalization during subsequent model training.

[0104] As an implementation manner, in step S50, the side training sample set and the top training sample set are input into an initial convolutional neural network for multi-stage training, and the network weight parameters are adjusted until the error between the output features of the initial convolutional neural network and the labeled vehicle model feature data is lower than a preset threshold, generating the vehicle model matching model, including the following steps:

[0105] Step S51: In the first stage of training, the side training sample set is input into the first feature extraction layer of the initial convolutional neural network to obtain a side primary feature vector.

[0106] The first feature extraction layer can adopt the ResNet-50 architecture as the backbone network. Its input end receives a three-channel side training image with a resolution of 512×512 pixels. After passing through the initial convolutional layer (7×7 kernel, stride 2) and the max pooling layer (3×3 kernel, stride 2), a feature map of 256×256×64 is output. Subsequently, high-level semantic features are gradually extracted through four residual blocks (each block contains 3 bottleneck structures), and finally a side primary feature vector with a dimension of 2048 is generated in the global average pooling layer. For example, for a side training image of a van, in the feature map output by the first residual block, the 32nd channel shows significant activation (response value 0.92) at the position of the cargo door hinge, and the 128th channel is sensitive to the high-frequency details of the tire ground pattern; after passing through the fourth residual block, the 512th dimension in the feature vector corresponds to the continuity measurement value of the cargo container side plate weld seam, and the 1024th dimension encodes the relative height difference between the door handle and the vehicle body plane. At this stage, the network weight parameters are initialized in a transfer learning manner, loading the model parameters pre-trained on the ImageNet dataset, and freezing the weights of the first three residual blocks to accelerate convergence.

[0107] Step S52: In the second training stage, input the top training sample set into the second feature extraction layer of the initial convolutional neural network to obtain the top primary feature vector.

[0108] The second feature extraction layer can be designed with the DenseNet-121 architecture. Its input is a single-channel top training image of 256×256 pixels (after grayscale conversion). It gradually aggregates multi-scale features through an initial convolutional layer (7×7 kernel, stride 2) and a sequence of dense blocks (each block contains 6 dense connection layers), and finally compresses the feature map size to 8×8×1024 through a transition layer (1×1 convolution and 2×2 average pooling), and then outputs a top primary feature vector with a dimension of 1024 through global max pooling. For example, when processing the top image of a refrigerated truck, in the feature map output by the first dense block, the 56th channel is sensitive to the lateral reinforcement spacing of the container top cover, and the 89th channel detects the rectangular contour of the refrigeration unit housing; after the transition layer, the 256th dimension in the feature vector reflects the material difference between the solar panel and the roof steel plate, and the 768th dimension encodes the radius change rate of the arc-shaped surface of the fairing. In this stage, all parameters of the second feature extraction layer are unfrozen, and an adaptive learning rate strategy is adopted (initial value 1e-4, decaying to 0.5 times every 10 epochs), focusing on optimizing the fine-grained feature extraction ability of the roof structure.

[0109] Step S53: In the third training stage, perform cross-attention fusion on the side primary feature vector and the top primary feature vector to generate a fused feature vector.

[0110] The cross-attention fusion module is constructed based on the multi-head attention mechanism. Its core is to achieve cross-modal information complementarity by calculating the dynamic weight distribution between the side and top features. Specifically, the side primary feature vector is used as the query vector (Query), and the top primary feature vector is used as the key-value vector (Key-Value). They are respectively mapped to a 64-dimensional subspace through linear transformation, and then the dot product similarity matrix between the query and the key is calculated, and the attention weight distribution is generated through Softmax normalization. For example, when the dot product score between the 512th dimension (container weld seam) in the side feature and the 256th dimension (lateral reinforcement) in the top feature reaches 0.87, it indicates a strong spatial correlation between the two, and this weight will enhance the fusion contribution of the two features; conversely, if the similarity between the tire feature (1024th dimension on the side) and the fairing feature on the roof (768th dimension on the top) is only 0.12, then its fusion ratio is reduced to avoid noise interference. The number of attention heads is set to 8, and each head independently calculates the weight distribution of the 64-dimensional subspace. Finally, the multi-head outputs are concatenated into a 512-dimensional fused feature vector. During this process, the interactive learning of the side-top features enables the model to capture three-dimensional geometric constraint relationships such as the container height and the roof arc, improving the representational ability of the features.

[0111] Step S54: In the fourth stage of training, input the fused feature vector into the fully-connected classification layer, and calculate the loss function value between the predicted vehicle model feature data and the labeled vehicle model feature data.

[0112] The fully-connected classification layer is composed of three parallel sub-networks, for example: the wheelbase regression sub-network (2-layer 512-dimensional fully-connected + linear output), the carriage structure classification sub-network (3-layer 1024-dimensional fully-connected + Softmax output), and the tire distribution matching sub-network (graph convolutional layer + contrastive loss calculation). The loss function adopts the form of multi-task weighted sum: the smooth L1 loss is used to measure the deviation between the predicted wheelbase sequence and the labeled value in the wheelbase part, and the weight is set to 0.5; the cross-entropy loss is used to optimize the classification accuracy in the carriage structure part, with a weight of 0.3; the tire distribution part is optimized based on the triplet loss to optimize the similarity of the tire position embedding vector, with a weight of 0.2. For example, when the predicted wheelbase is [3.21, 7.83, 11.41] meters and the labeled value is [3.19, 7.85, 11.43] meters, the smooth L1 loss is calculated as 0.015; the probability that the predicted carriage category is "refrigerated container" is 0.92, and the label is a one-hot vector [0, 1, 0,...], and the cross-entropy loss is 0.083; the triplet loss of the tire distribution is calculated through the embedding distance difference between the positive sample (tire position of the same vehicle model) and the negative sample (tire position of different vehicle models). If the positive sample distance is 0.2, the negative sample distance is 1.3, and the margin α = 0.5, then the loss value is max(0.2 - 1.3 + 0.5, 0) = 0. The total loss function value is 0.5×0.015 + 0.3×0.083 + 0.2×0 = 0.034.

[0113] Step S55: According to the loss function value, backpropagate to adjust the weight parameters of the first feature extraction layer, the second feature extraction layer, and the fully-connected classification layer until the loss function value converges to the preset threshold.

[0114] Adam optimizer can be used for backpropagation, with the initial learning rate set to 3e-5, β1 = 0.9, β2 = 0.999, and the weight decay coefficient 1e-4. During the training process, 32 groups of samples (16 pairs of side-top images) are input in each batch. The gradients of each layer are calculated through the chain rule of differentiation, and the parameters are updated according to the learning rate. For example, at the 150th epoch of training, the gradient of the weight matrix W1 of the fully-connected classification layer is ∂L / ∂W1 = 0.0023, and the update amount based on the momentum estimation of Adam is , and the parameter update step size When the loss of the validation set drops by less than 0.1% for 10 consecutive epochs or the total number of training epochs reaches 300, the training is terminated. The average absolute error of the wheelbase of the final model on the test set drops to 8 mm, the classification accuracy of the carriage is 97.2%, and the F1 score of the tire distribution matching is 0.91, meeting the preset thresholds (wheelbase MAE ≤ 10 mm, classification accuracy ≥ 95%, F1 ≥ 0.9), and a deployable vehicle model matching model is generated.

[0115] As an implementation manner, in step S53, cross-attention fusion is performed on the side primary feature vector and the top primary feature vector to generate a fused feature vector, including the following steps:

[0116] Step S531: Calculate the correlation matrix between the side primary feature vector and the top primary feature vector, and the correlation matrix reflects the association strength between different feature dimensions.

[0117] The correlation matrix can be calculated, for example, through a bilinear attention mechanism. Given the side feature vector Vs ∈ R 2048 and the top feature vector Vt ∈ R 1024 , first project Vs into the query space Q = Wq·Vs (Wq ∈ R 512×2048 ), project Vt into the key space K = Wk·Vt (Wk ∈ R 512×1024 ), and then calculate the correlation matrix C = Q·K T ∈ R 512×512 , where each element C_ij represents the association strength between the i-th feature dimension on the side and the j-th feature dimension on the top. For example, when the C value between the 312th dimension (encoding the height of the door handle) in the side feature and the 198th dimension (encoding the position of the skylight) in the top feature is 0.94, it indicates a strong spatial correspondence; while the C value between the 1024th dimension (tire ground contact mark) on the side and the 56th dimension (material of the container top cover) on the top is only 0.12, reflecting a low correlation. After matrix calculation, row-wise Softmax normalization is applied to make the sum of weights in each row equal to 1.

[0118] Step S532: Generate a side feature weight distribution and a top feature weight distribution according to the correlation matrix.

[0119] The side feature weight distribution is obtained, for example, by summing and normalizing the columns of the correlation matrix, that is, Ws = Softmax(sum(C, axis = 1)) ∈ R 512 , reflecting the importance of each side feature dimension in the global context; the top feature weight distribution is obtained by summing and normalizing the rows, Wt = Softmax(sum(C, axis = 0)) ∈ R 512, representing the contribution degree of the top feature dimension. For example, if the column sum of the 512th dimension (container weld seam) on the side reaches 38.7 (the maximum value) in the correlation matrix, then Ws

[512] = 0.15; the row sum of the 256th dimension (horizontal stiffener) on the top is 29.3, and Wt

[256] = 0.12. This process strengthens the feature fusion of the key structures of the vehicle by focusing on the high-response areas.

[0120] Step S533: Weight and sum the side primary feature vectors according to the side feature weight distribution to generate a side weighted feature vector.

[0121] Side weighted feature vector , where Ws i is the extended weight mapped to the original 2048-dimensional space (the 512-dimensional Ws is extended to 2048 dimensions by linear interpolation). For example, in the side features, the dimensions related to the door handle (the 312th - 320th dimensions) have higher Ws weights (0.08 - 0.12), and after weighting, their feature values are increased from the original 0.75 to 0.89; while the feature values of the low-weight dimensions (such as the background noise in the 1500th - 1600th dimensions) are decreased from 0.32 to 0.05. After weighting, Vsw ∈ R 2048 Retain important features while suppressing redundant information.

[0122] Step S534: Weight and sum the top primary feature vectors according to the top feature weight distribution to generate a top weighted feature vector.

[0123] Top weighted feature vector , where Wt j is extended from 512 dimensions to 1024 dimensions by interpolation. For example, in the top features, the dimensions related to the fairing curvature (the 768th - 780th dimensions) have Wt weights of 0.10 - 0.15, and after weighting, the feature values are enhanced from 0.68 to 0.82; the feature values of the container top cover edge features (the 100th - 120th dimensions) have lower weights (0.03 - 0.05) and are decreased from 0.91 to 0.45. Vtw ∈ R 1024 Highlight the key structures of the vehicle roof and weaken the secondary areas.

[0124] Step S535: Concatenate the side weighted feature vector and the top weighted feature vector to generate the fused feature vector.

[0125] The concatenation operation is performed along the feature dimension, and the fused feature vector Vf = Concat(Vsw, Vtw) ∈ R 3072For example, the value of the 2048th dimension (door handle feature) in Vsw is 0.89, and the value of the 1024th dimension (fairing feature) in Vtw is 0.82. After splicing, Vf

[2048] = 0.89 and Vf

[3072] = 0.82. To reduce the dimension, subsequent compression is performed through a fully connected layer (3072→512 dimensions), and LayerNorm normalization and ReLU activation are applied. Finally, a fused feature vector suitable for multi-task learning is generated.

[0126] In an alternative embodiment, after step S500 of determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification, the method may further include: real-time detecting whether the comprehensive matching degree is lower than a preset matching threshold. When it is lower, extracting the unmatched vehicle wheelbase feature, carriage structure feature, and tire distribution feature from the vehicle type feature matching result; generating a temporary wheelbase code according to the unmatched vehicle wheelbase feature, and constructing a new vehicle type feature data packet in combination with the carriage structure feature and tire distribution feature; performing artificial verification interface matching between the new vehicle type feature data packet and each vehicle type template data in the vehicle type feature library to obtain verified feature data passed by manual review; adding the verified feature data to the vehicle type feature library and updating the template wheelbase range, template carriage structure template, and template tire distribution template; dynamically correcting the vehicle type identification of subsequent passing vehicles according to the updated vehicle type feature library.

[0127] This embodiment is applicable to the learning of vehicle type characteristics and the update of the database. When the comprehensive matching degree is detected to be lower than the preset matching threshold (such as 0.85) in real time, the system automatically triggers the process of extracting unmatched characteristics. The unmatched vehicle wheelbase characteristics refer to the numerical segments in the predicted wheelbase sequence that have no intersection with the template wheelbase ranges of all vehicle type templates. For example, if the measured wheelbase values of a vehicle are [3.25, 7.91, 11.62] meters, and the nearest template wheelbase range in the vehicle type library is [3.18 - 3.22, 7.75 - 7.85, 11.3 - 11.5] meters, then the third wheelbase value of 11.62 meters is marked as an unmatched characteristic segment. The temporary wheelbase encoding is generated using a segmented hashing algorithm: the unmatched wheelbase values are discretized at 10 - centimeter intervals (e.g., 11.62 meters is mapped to 116.2 discrete units), and the first 8 bits of the MD5 hash operation on each discrete unit are taken as the encoding identifier (e.g., "3D5F8A2B"). The new vehicle type characteristic data packet consists of the temporary wheelbase encoding, the carriage structure feature vector (512 - dimensional), and the tire distribution topology map (including tire position coordinates and adjacency relationships). The data packet format is a JSON - LD structured document, accompanied by a timestamp and a spatial position label. The manual verification interface is deployed on the toll system management platform. The reviewer can compare the new vehicle type characteristic data packet with the historical abnormal records. If it is confirmed to be a new type of vehicle (such as the first passage of a certain type of electric heavy truck), the verification - passed option is selected and vehicle type parameters (such as load - carrying level, power type) are supplemented. The updated vehicle type characteristic library appends the new vehicle type template data to the storage node by means of incremental writing. The template wheelbase range is extended to [3.18 - 3.22, 7.75 - 7.85, 11.3 - 11.6] meters, the template carriage structure template adds the battery compartment layout feature vector unique to electric trucks, and the template tire distribution template adds the tread encoding rule for wide - body single tires. When subsequent passing vehicles conduct similarity comparison, the updated vehicle type characteristic library is dynamically loaded. For example, when a vehicle of the same type of electric truck passes again, its vehicle type identification is updated from "UNKNOWN" to "ET - 38D", and the new electricity rate of 0.52 yuan / ton·kilometer is triggered for the rate calculation module to call.

[0128] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating the vehicle type identification identifier in step S500, the method may further include: generating an error distribution heat map according to the wheelbase matching degree, the carriage matching degree, and the tire matching degree, where the error distribution heat map reflects the recognition deviation positions of each local feature region; extracting the coordinates of the abnormal regions in the error distribution heat map whose deviation values exceed the tolerance threshold, and mapping them to the corresponding positions of the vehicle multi-dimensional fusion feature map; adjusting the convolution kernel size and the stride parameter of the multi-level convolutional network in the vehicle type matching model according to the abnormal region coordinates; performing incremental training on the vehicle type matching model using the adjusted convolution kernel parameters, and updating the activation weights of the feature response map of the multi-level convolutional network; and applying the updated vehicle type matching model to calculate the vehicle type feature matching results of the next batch of passing vehicles.

[0129] This embodiment is a process of dynamic optimization and incremental learning of model parameters. The error distribution heat map characterizes the recognition deviation intensity of each local feature region in the form of a three-dimensional tensor, where the heat value is calculated from the weighted residuals of the wheelbase matching degree, the carriage matching degree, and the tire matching degree. For example, in the front axle area of the container (coordinates x: 120 - 180, y: 300 - 400 pixels), the wheelbase matching degree residual is detected to be 0.23 (threshold 0.15), the carriage structure matching degree residual is 0.18 (threshold 0.1), and the tire matching degree residual is 0.32 (threshold 0.2), then the heat value of this area is marked as red (RGB: 255, 0, 0). The coordinates of the abnormal regions are extracted through connected component analysis. When the heat value of a continuous 10×10 pixel area exceeds the tolerance threshold, the coordinates of its minimum bounding rectangle are recorded (such as the upper left corner point (125, 305) and the lower right corner point (175, 395)). When mapping to the vehicle multi-dimensional fusion feature map, a spatial projection conversion algorithm is used to convert the two-dimensional coordinates into the channel index of the feature map. For example, the coordinates (150, 350) correspond to the 35th activation region of the 24th channel of the feature map. The convolution kernel size adjustment strategy is dynamically set according to the area of the abnormal region: for regions with an area less than 200 pixel² (such as abnormal tire contact points), the kernel size of the corresponding convolutional layer is expanded from 3×3 to 5×5 to enhance the local feature capture ability; for regions with an area exceeding 500 pixel² (such as large-scale mismatch on the side of the container), the stride parameter is adjusted from 2 to 1 to reduce the downsampling loss. The incremental training uses an online learning framework, extracts 5% of the samples from the real-time data stream as the training set, inputs 16 groups of samples in each batch, sets the training cycle to 10 epochs, and the learning rate decays to 0.1 times the initial value. The distribution of the activation weights of the feature response map of the updated multi-level convolutional network changes significantly: for example, the average activation value of the original model in the front axle area of the container is 0.75, and it is increased to 0.88 after adjustment; the standard deviation of the feature response intensity in the tire contact point area is reduced from 0.15 to 0.09.

[0130] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating the vehicle type identification identifier in step S500, the method may further include: judging whether the vehicle type identification identifier contains a preset unknown vehicle type identifier, and if so, activating an image re-acquisition instruction; controlling the shooting device to adjust the focal length and angle according to the image re-acquisition instruction, and re-acquiring the high-resolution side contour image and top contour image of the passing vehicle; performing multi-frame de-blurring processing on the re-acquired images to generate a set of re-acquired images with enhanced clarity; inputting the set of re-acquired images into the vehicle type matching model for secondary feature matching, and updating the vehicle wheelbase feature and carriage structure feature in the vehicle type feature matching result; re-performing the similarity comparison according to the updated vehicle type feature matching result to overwrite the original vehicle type identification identifier.

[0131] This embodiment relates to the re-identification of unknown vehicle types and result coverage. The preset unknown vehicle type identifier adopts a coding rule of adding a 6-bit random character after the prefix "UNKNOWN-" (such as "UNKNOWN-3A5F9D"). When the vehicle type identification identifier contains this prefix, the re-acquisition instruction system activates a multi-modal triggering mechanism: first, calculate the optimal shooting angle according to the real-time position of the vehicle (positioned by the inductive loop coordinates), adjust the pitch angle of the top camera from 30° to 45° to cover the full height of the container, and at the same time switch the focal length of the side-view camera from 50mm to 200mm telephoto mode to capture the details of the door rivets. The resolution of the re-acquired high-resolution image is 8192×4320 pixels, and the frame rate is increased to 60fps to support multi-frame de-blurring processing. The multi-frame de-blurring adopts an algorithm combining non-uniform motion estimation and iterative deconvolution: for the lateral blur caused by vehicle vibration, estimate the displacement vector of each frame image by the optical flow method (such as the inter-frame displacement Δx = 2.3 pixels, Δy = 0.8 pixels), and perform Wiener filtering restoration after constructing the point spread function matrix. In the enhanced set of re-acquired images, the MTF (modulation transfer function) value of the tire tread pattern is increased from 0.35 to 0.62, and the edge sharpness of the door handle is increased by 40%. During secondary feature matching, the vehicle type matching model enables the high-precision mode: the output dimension of the wheelbase regression sub-network is expanded from 3 to 5 to increase the detection of the support axis of multi-axle vehicles; the number of categories of the carriage structure classifier is increased from 120 to 200 to accommodate new special vehicles. In the updated vehicle type feature matching result, the vehicle wheelbase feature of a certain electric truck is corrected from [3.25, 7.91, 11.62] meters to [3.24, 7.89, 11.60] meters, and the cosine similarity between the carriage structure feature vector and the template is increased from 0.72 to 0.91. The similarity comparison module executes a forced coverage protocol: when the comprehensive matching degree of the secondary matching exceeds the original result and the difference is greater than 0.1, the original identifier "UNKNOWN-3A5F9D" is replaced by "ET-38D", and the version number (such as "Ver2.1") and the correction timestamp are recorded in the transaction log.

[0132] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating the vehicle type identification identifier in step S500, the method further includes: counting the identification frequencies of each toll vehicle type category within a preset time period to generate a vehicle type frequency distribution table; sorting the storage priorities of the vehicle type template data in the vehicle type feature library according to the vehicle type frequency distribution table; migrating the template wheelbase range, template carriage structure template, and template tire distribution template of the high-frequency vehicle types to the cache area; establishing a fast retrieval queue based on the cache area, and preferentially comparing the vehicle type feature matching results of subsequent passing vehicles with the vehicle type template data in the fast retrieval queue; when the comprehensive matching degree does not meet the standard in the fast retrieval queue, switching to the full amount of data in the vehicle type feature library for secondary retrieval.

[0133] This embodiment realizes the dynamic optimization and retrieval of the vehicle type feature library, counts the recognition frequencies of each toll vehicle type category within a preset time period, and generates a vehicle type frequency distribution table. The preset time period is set as a dynamically adjustable time window. For example, different statistical granularities are used during the morning rush hour (07:00 - 09:00) and the night time period (22:00 - 06:00) (the former is sliced by 15 minutes and the latter is sliced by 1 hour). The vehicle type frequency distribution table is stored in a hash map structure, where the key is the vehicle type number (such as "HT-45C"), and the value is a composite object containing the recognition times, timestamp sequence, and average matching degree. For example, in the morning rush hour statistics, the frequency of the logistics truck "LT-32B" is 8.7 vehicles per minute, and the average matching degree is 0.93, while the frequency of the passenger car "CT-18D" is 2.3 vehicles per minute, and the average matching degree is 0.87. When sorting the storage priorities of the vehicle type template data in the vehicle type feature library according to the vehicle type frequency distribution table, a heat weighted algorithm is used: priority score = frequency × 0.6 + average matching degree × 0.4. For example, the score of "LT-32B" = 8.7 × 0.6 + 0.93 × 0.4 = 5.93, ranking first in priority. When migrating the template wheelbase range, template carriage structure template, and template tire distribution template of high-frequency vehicle types to the cache area, the cache adopts the LRU (Least Recently Used) replacement strategy, and the capacity is set to 20% of the total number of vehicle types (such as the first 200 out of 1000). The migration process is realized through memory mapping technology, loading the template data on the disk into the DDR4-3200 memory module, and reducing the access latency from 15ms to 0.2ms. When establishing a fast retrieval queue based on the cache area, the queue index structure is organized by a B+ tree, where the key is the vehicle type number, and the value points to the address pointer of the template data in memory. For example, the template wheelbase range [3.18, 3.22, 7.75, 7.85, 11.3, 11.5] meters, the carriage structure template (512-dimensional vector), and the tire distribution template (12-node topology map) of "LT-32B" are loaded into the memory address area 0x7FFA2B1C0000 - 0x7FFA2B1D8000. When preferentially comparing the vehicle type feature matching results of subsequent passing vehicles with the vehicle type template data in the fast retrieval queue, the SIMD instruction set is used to accelerate the vector similarity calculation. For example, the AVX-512 instruction is used to process the 512-dimensional carriage structure vector in parallel, and the single comparison time is shortened from 1.2ms to 0.15ms. When the comprehensive matching degree does not meet the standard in the fast retrieval queue (such as the highest matching degree 0.82 is lower than the threshold 0.85), switch to the full amount of data in the vehicle type feature library for secondary retrieval. The full amount retrieval enables a multi-threaded sharding query mechanism, dividing the vehicle type feature library into 8 shards, and each shard is scanned in parallel by an independent thread in the NVMe SSD storage pool. The total query time is optimized from 220ms for full amount scanning to 35ms.For example, when a new electric truck passes for the first time and there is no matching item in the quick search queue, after the system switches to full-scale search, a new "ET-38D" vehicle model template is matched in shard 3, with a comprehensive matching degree of 0.89, triggering the cache update mechanism to add it to the cache area.

[0134] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification identifier in step S500, it further includes: extracting the maximum interval point coordinates of the vehicle wheelbase feature and the edge connection point sequence of the carriage structure feature from the vehicle type feature matching result; generating a three-dimensional vehicle type contour report based on the maximum interval point coordinates and the edge connection point sequence, the report including a wheelbase ratio parameter and a carriage structure topology map; binding the three-dimensional vehicle type contour report to the vehicle type identification identifier, and adding a timestamp and lane position information; performing lossless compression encoding on the bound data packet to generate a standardized identification report; and sending the standardized identification report to the toll terminal and the cloud audit platform for parallel verification.

[0135] This embodiment relates to the encapsulation of three-dimensional vehicle model data and multi-terminal verification, extracting the coordinates of the maximum interval points of the vehicle wheelbase feature and the sequence of edge connection points of the carriage structure feature from the vehicle model feature matching results. The coordinates of the maximum interval points refer to the two axis center projection points with the largest adjacent axle spacing in the vehicle wheelbase sequence. For example, for a five-axle truck with a wheelbase sequence of [3.2, 7.8, 11.4, 15.0, 18.6] meters, the maximum interval is 3.6 meters from the fourth axle to the fifth axle, corresponding to the point coordinates (x1 = 1520, y1 = 450) and (x2 = 1840, y2 = 450) in the image coordinate system. The sequence of edge connection points extracts the contour polygon vertices of the carriage structure feature through the Alpha Shape algorithm. For example, the edge point sequence of the side panel of the refrigerated truck container contains 56 vertices and is stored in clockwise order as [(x1, y1), (x2, y2),..., (x56, y56)]. When generating a three-dimensional vehicle model contour report based on the coordinates of the maximum interval points and the sequence of edge connection points, a point cloud-based surface reconstruction technology is adopted: taking the maximum interval points as the longitudinal reference line, stretching the sequence of edge connection points vertically to generate a three-dimensional mesh model, and calculating the wheelbase ratio parameter (such as the distance between the fourth axle and the fifth axle accounting for 19.3% of the total vehicle length). The carriage structure topology map is constructed through the Delaunay triangulation algorithm, converting the sequence of edge connection points into a non-uniform rational B-spline (NURBS) surface. For example, the corrugated structure of the container side panel is modeled as a periodic surface with an amplitude of 5 mm and a wavelength of 120 mm. When binding the three-dimensional vehicle model contour report with the vehicle model identification identifier and adding the timestamp and lane position information, the bound data packet adopts the Protobuf binary encoding format, with the timestamp accurate to the millisecond level (such as "2023-08-20T14:23:05.235Z"), and the lane position information is combined and located through the RFID landmark coordinates of the toll station (such as the WGS-84 coordinates 118.3245°E, 32.4567°N of lane 3) and the relative position of the vehicle (12.5 meters from the entrance landmark). When performing lossless compression encoding on the bound data packet, the DEFLATE algorithm combined with Huffman coding is adopted, and the compression ratio is set to the highest level (Level 9). For example, the original data packet size of 2.3 MB is reduced to 480 KB after compression. When generating a standardized identification report, the report header contains fields such as the version number (such as "V1.2"), checksum (CRC-32), and data packet length, and the compressed three-dimensional vehicle model data is embedded in the body part.When sending the standardized recognition report to the toll terminal and the cloud audit platform for parallel verification, the toll terminal locally verifies the data integrity through SHA-256, for example, calculating the hash value "a1b2c3d4e5f6..." of the compressed package and comparing it with the checksum in the report header; the cloud audit platform performs redundancy verification through distributed verification nodes, for example, three nodes respectively verify the timestamp continuity, vehicle type identification compliance, and coordinate validity. After all pass, a confirmation signal is returned to the toll system to trigger the rate calculation and release instruction. If a certain verification fails (such as the vehicle type identification "ET-38D" is not registered in the cloud), the system automatically isolates the abnormal data packet and starts the manual review process, and at the same time records the abnormal event in the audit log.

[0136] Figure 2 A schematic diagram of the hardware entity of a computer system provided by an embodiment of the present invention is as Figure 2 shown. The hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002. Among them, the memory 1002 stores a computer program that can run on the processor 1001. When the processor 1001 executes the program, it implements the steps in the method of any of the above embodiments.

[0137] The memory 1002 stores a computer program that can run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001, and can also cache data to be processed or already processed by the processor 1001 and each module in the computer system 1000 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0138] When the processor 1001 executes the program, it implements the steps of the toll vehicle type recognition method based on highway passing images in any of the above items. The processor 1001 generally controls the overall operation of the computer system 1000.

[0139] An embodiment of the present invention provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the toll vehicle type recognition method based on highway passing images in any of the above embodiments.

[0140] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present invention, please refer to the description of the method embodiments of the present invention for understanding. The above processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices for implementing the functions of the above processor are also possible, and the embodiments of the present invention do not make specific limitations.

[0141] The above computer storage medium / memory may be a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read only memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0142] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above steps / processes do not mean the sequence of execution, and the execution sequence of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages or disadvantages of the embodiments. It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0143] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.

[0144] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0146] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.

[0147] Alternatively, if the above integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes.

[0148] As described above, the above are only the implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A toll vehicle type recognition method based on highway traffic images, characterized in that The method includes: Obtaining an image set of passing vehicles, where the image set includes multiple frames of continuously captured vehicle side profile images and vehicle top profile images; Performing multi-scale brightness compensation processing on the vehicle side profile images to generate side feature enhanced images, and performing contour sharpening processing on the vehicle top profile images to generate top feature enhanced images; Performing dual-channel feature fusion on the side feature enhanced images and the top feature enhanced images to generate a vehicle multi-dimensional fusion feature map; Inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model to obtain a vehicle model feature matching result output by the vehicle model matching model, where the vehicle model feature matching result includes vehicle wheelbase features, carriage structure features, and tire distribution features; Comparing the similarity between the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark.

2. The method according to claim 1, wherein The performing multi-scale brightness compensation processing on the vehicle side profile images to generate side feature enhanced images includes: Extracting a continuous brightness distribution curve in the vehicle side profile image, where the continuous brightness distribution curve includes a first brightness distribution in the vehicle bottom area and a second brightness distribution in the vehicle top area; Generating a dynamic brightness compensation coefficient according to the gradient difference between the first brightness distribution and the second brightness distribution; Locally enhancing the low brightness area in the vehicle side profile image according to the dynamic brightness compensation coefficient to generate an initial compensation image; Performing multi-scale filtering processing on the initial compensation image to extract edge texture features at different resolutions; Adjusting the filtering parameters according to the density distribution of the edge texture features to generate the side feature enhanced image.

3. The method according to claim 2, wherein The performing dual-channel feature fusion on the side feature enhanced images and the top feature enhanced images to generate a vehicle multi-dimensional fusion feature map includes: Performing spatial coordinate transformation on the side feature enhanced image to obtain a set of key contour points in the side image coordinate system; Performing spatial coordinate transformation on the top feature enhanced image to obtain a set of key contour points in the top image coordinate system; Performing three-dimensional space mapping on the set of key contour points in the side image coordinate system and the set of key contour points in the top image coordinate system to generate a vehicle three-dimensional contour model; Extracting vehicle surface curvature features according to the curvature distribution of each contour point in the vehicle three-dimensional contour model; Superimposing and fusing the vehicle surface curvature features and the edge texture features of the side feature enhanced image to generate the vehicle multi-dimensional fusion feature map.

4. The method according to claim 1, characterized in that, The inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model to obtain a vehicle model feature matching result output by the vehicle model matching model includes: Dividing the vehicle multi-dimensional fusion feature map into multiple local feature regions, and each local feature region corresponds to one feature type of vehicle wheelbase features, carriage structure features, or tire distribution features; Performing feature vectorization processing on each local feature region to generate a set of local feature vectors; Input the set of local feature vectors into the multi-level convolutional network in the vehicle model matching model to obtain the feature response maps output by each convolutional layer; Determine the wheelbase matching degree corresponding to the vehicle wheelbase feature, the carriage matching degree corresponding to the carriage structure feature, and the tire matching degree corresponding to the tire distribution feature according to the activation region distribution of the feature response maps; Perform weighted fusion on the wheelbase matching degree, the carriage matching degree, and the tire matching degree to generate the vehicle model feature matching result.

5. The method according to claim 4, wherein The comparing the similarity between the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark includes: Extract the vehicle model template data from the vehicle model feature library, where the vehicle model template data includes a template wheelbase range, a template carriage structure template, and a template tire distribution template; Calculate a first similarity between the wheelbase matching degree and the template wheelbase range, calculate a second similarity between the carriage matching degree and the template carriage structure template, and calculate a third similarity between the tire matching degree and the template tire distribution template; Determine the comprehensive matching degree between the passing vehicle and each vehicle model template data according to the weighted sum of the first similarity, the second similarity, and the third similarity; Use the vehicle model number corresponding to the vehicle model template data with the highest comprehensive matching degree as the toll vehicle model category, and generate a vehicle model identification mark including the vehicle model number.

6. The method according to claim 1, wherein The method further includes a pre-training step of the vehicle model matching model: Obtain a historical passing vehicle training set, where the historical passing vehicle training set includes multiple groups of vehicle image samples and corresponding labeled vehicle model feature data; Perform image segmentation processing on each group of vehicle image samples to extract a vehicle side training image and a vehicle top training image; Perform noise suppression and distortion correction on the vehicle side training image to generate a side training sample set; Perform shadow elimination and contrast enhancement on the vehicle top training image to generate a top training sample set; Input the side training sample set and the top training sample set into an initial convolutional neural network for multi-stage training, and adjust the network weight parameters until the error between the output features of the initial convolutional neural network and the labeled vehicle model feature data is lower than a preset threshold to generate the vehicle model matching model.

7. The method according to claim 6, wherein The performing image segmentation processing on each group of vehicle image samples to extract a vehicle side training image and a vehicle top training image includes: Perform background separation processing on the vehicle image samples to obtain a vehicle main body region mask; Divide the vehicle side region and the vehicle top region according to the geometric center position of the vehicle main body region mask; Perform edge tracking on the vehicle side region to extract a continuous closed side contour line; Determine the cropping boundary of the vehicle side training image according to the curvature change points of the side contour line; Segment the vehicle image samples according to the cropping boundary to generate the vehicle side training image and the vehicle top training image; The performing noise suppression and distortion correction on the vehicle side training image to generate a side training sample set includes: Detect the high-frequency noise area in the side training image of the vehicle and generate a noise distribution map; Adjust the window size of the adaptive filter according to the density value of the noise distribution map and smooth the high-frequency noise area; Extract the perspective distortion parameters in the side training image of the vehicle, where the perspective distortion parameters include the horizontal tilt angle and the vertical scaling ratio; Perform affine transformation correction on the side training image of the vehicle according to the horizontal tilt angle, and adjust the image ratio of the side training image according to the vertical scaling ratio to generate the side training sample set.

8. The method according to claim 6, wherein The step of inputting the side training sample set and the top training sample set into the initial convolutional neural network for multi-stage training, adjusting the network weight parameters until the error between the output features of the initial convolutional neural network and the labeled vehicle model feature data is lower than the preset threshold, and generating the vehicle model matching model, includes: In the first training stage, input the side training sample set into the first feature extraction layer of the initial convolutional neural network to obtain the side primary feature vector; In the second training stage, input the top training sample set into the second feature extraction layer of the initial convolutional neural network to obtain the top primary feature vector; In the third training stage, perform cross-attention fusion on the side primary feature vector and the top primary feature vector to generate a fused feature vector; In the fourth training stage, input the fused feature vector into the fully connected classification layer to calculate the loss function value between the predicted vehicle model feature data and the labeled vehicle model feature data; Backpropagate according to the loss function value to adjust the weight parameters of the first feature extraction layer, the second feature extraction layer, and the fully connected classification layer until the loss function value converges to the preset threshold.

9. The method according to claim 8, wherein The step of performing cross-attention fusion on the side primary feature vector and the top primary feature vector to generate a fused feature vector includes: Calculate the correlation matrix between the side primary feature vector and the top primary feature vector, where the correlation matrix reflects the association strength between different feature dimensions; Generate the side feature weight distribution and the top feature weight distribution according to the correlation matrix; Perform weighted summation on the side primary feature vector according to the side feature weight distribution to generate a side weighted feature vector; Perform weighted summation on the top primary feature vector according to the top feature weight distribution to generate a top weighted feature vector; Concatenate the side weighted feature vector and the top weighted feature vector to generate the fused feature vector.

10. A computer system, comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Vehicle re-identification method and device based on roadside perception, and electronic equipment

    CN114170516A

  • Vehicle type identification method based on deep learning fusion model

    CN116863412A