Toll vehicle type identification method and system based on expressway passing image
Feature enhancement and dual-channel feature fusion through multi-frame continuous shooting of vehicle side and top images, combined with pre-trained model matching models, the problem of vehicle model recognition in the prior art is solved, and the recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510468638.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the existing highway toll system, vehicle model identification methods are susceptible to light conditions and shooting angles, resulting in blurred vehicle profiles and difficult extraction of key structural features, which affect the accuracy of vehicle model matching.
The vehicle side contour image and top contour image are used to continuously shoot multi-frames, and the side feature enhancement image and top feature enhancement image are generated through multi-scale brightness compensation processing and contour sharpening processing, and dual-channel feature fusion is performed to generate a vehicle multi-dimensional fusion feature map, and a pre-trained vehicle model matching model is input for feature matching.
It effectively solves the problem of feature blurring caused by uneven lighting or shooting angle of vehicle images, improves the robustness and accuracy of vehicle model recognition, and significantly improves the model classification ability in complex environments.
Smart Images

Figure CN119992483A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and machine learning, and in particular to a method and system for identifying toll vehicle types based on highway traffic images. Background Art
[0002] In the highway toll collection system, vehicle model recognition is the core link in determining tolls and improving the efficiency of toll stations. The current mainstream technology mainly collects single-view images of vehicles (such as only side or top views) to classify vehicle models. However, such methods have significant defects in practical applications: on the one hand, vehicle images with a single view are easily affected by lighting conditions (such as backlighting, shadows), shooting angle deviations or environmental interference (such as rainy and foggy weather), resulting in blurred vehicle contours and difficulty in extracting key structural features (such as wheelbase and tire distribution). For example, the side image may not clearly present the tire spacing in a low-light environment, and the top image may cause deformation of the cabin structure due to the tilt of the shooting angle, which will directly affect the accuracy of vehicle model matching. On the other hand, traditional image processing methods usually only perform simple enhancements (such as global brightness adjustment) for a single view, lacking collaborative analysis and fusion of multi-dimensional features (such as the correlation between the side wheelbase and the top cabin structure), resulting in a single feature dimension for vehicle model classification and limited discrimination ability. For example, relying only on side images may not be able to distinguish between models with similar cabin heights but different wheelbases, while relying only on top images will make it difficult to identify vehicles with the same number of tires but different cabin structures.
[0003] The above technical defects have led to the existing vehicle model recognition systems generally facing problems such as high feature extraction error rate (such as wheelbase measurement deviation due to image blur), low classification error tolerance (such as misjudgment caused by illumination changes) and poor adaptability to complex scenes (such as insufficient classification basis due to the lack of fusion of multi-view features), which seriously restricts the accuracy of toll collection and traffic efficiency. Therefore, there is an urgent need for a vehicle model recognition method to improve the robustness and accuracy of vehicle model classification in complex environments. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a method and system for identifying toll vehicle types based on highway traffic images. The technical solution of the embodiment of the present invention is implemented as follows: On the one hand, an embodiment of the present invention provides a method for identifying a toll vehicle type based on a highway passing image, the method comprising: Acquire an image set of passing vehicles, wherein the image set includes multiple frames of vehicle side profile images and vehicle top profile images taken continuously; Performing multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image, and performing contour sharpening processing on the vehicle top profile image to generate a top feature enhanced image; Performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a multi-dimensional fusion feature map of the vehicle; Inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model, obtaining a vehicle model feature matching result output by the vehicle model matching model, wherein the vehicle model feature matching result includes a vehicle wheelbase feature, a vehicle compartment structure feature, and a tire distribution feature; A similarity comparison is performed based on the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark.
[0005] On the other hand, an embodiment of the present invention provides a computer system, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.
[0006] Beneficial effects of the present invention: The present invention obtains a set of passing vehicle images including multiple frames of continuously shot vehicle side profile images and vehicle top profile images, performs multi-scale brightness compensation processing on the side images to generate side feature enhanced images, and performs contour sharpening processing on the top images to generate top feature enhanced images, which effectively solves the feature blurring problem of vehicle images caused by uneven lighting or shooting angles. Specifically, the side feature enhanced image can highlight key structural information such as vehicle wheelbase and tire distribution, while the top feature enhanced image strengthens the top perspective features such as the vehicle compartment structure and cargo form. The multi-dimensional fusion feature map of the vehicle generated by dual-channel feature fusion can comprehensively utilize the complementary feature information of different dimensions of the vehicle to avoid the limitations of single perspective features. Furthermore, the vehicle wheelbase features, vehicle compartment structure features and tire distribution features are extracted through the pre-trained vehicle model matching model. These features are highly correlated with the key discrimination indicators of vehicle model classification. Combined with the pre-stored vehicle model feature library, multi-dimensional feature similarity comparison is performed to accurately match the vehicle model category. The vehicle model identification logo finally generated realizes automated, multi-dimensional feature analysis and classification of passing vehicles, significantly improving the robustness of vehicle model recognition in complex lighting environments. At the same time, by integrating dual-perspective feature enhancement and multi-dimensional feature matching mechanism, it effectively improves the accuracy of vehicle model classification and the efficiency of toll station passage.
[0007] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present invention and, together with the specification, are used to explain the technical solutions of the present invention.
[0009] Figure 1 The present invention provides a flowchart of a method for identifying toll vehicle types based on highway traffic images according to an embodiment of the present invention.
[0010] Figure 2 A hardware entity schematic diagram of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention.
[0012] The embodiment of the present invention provides a method for identifying toll vehicle types based on highway traffic images, which can be executed by a processor of a computer system. The computer system can refer to a device with data processing capabilities such as a server, a laptop, a tablet computer, and a desktop computer. For example, a server deployed in the background of a high-speed toll system, a computer device deployed at the edge of a high-speed toll station, and so on.
[0013] Figure 1 A schematic diagram of the implementation flow of a method for identifying toll vehicle types based on highway traffic images provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method comprises the following steps: Step S100: Acquire an image set of passing vehicles, wherein the image set includes multiple frames of vehicle side profile images and vehicle top profile images taken continuously.
[0014] In this step, the acquisition of image sets is the basic data source for vehicle model recognition. The side profile image of a vehicle refers to a static or dynamic image of the overall side profile of the vehicle captured by a camera perpendicular to the direction of travel of the vehicle, which can reflect key features such as the number of doors, window layout and body length; the top profile image of the vehicle is an image of the top structure of the vehicle captured from a bird's-eye view, which is used to present three-dimensional spatial information such as the shape of the roof, the height of the cargo box and the position of the antenna. Multi-frame continuous shooting emphasizes that the image acquisition device needs to dynamically capture the moving vehicle at a fixed time interval (such as 30 frames per second) to ensure that the complete appearance characteristics of the vehicle in motion are captured at different time nodes. For example, in a highway toll lane, when a vehicle passes at a speed of 20 kilometers per hour, the high-speed camera groups deployed on both sides of the lane synchronously trigger the shooting mechanism. The left camera group captures a side view sequence of images including the height of the cab and the size of the tire, while the top camera group records the top structural details such as the container connection and the roof ventilation device. For special scenarios, such as nighttime driving or rainy and foggy weather, the system can enhance the image clarity through a combination of an infrared fill light module and a polarizing filter to ensure the continuity of the vehicle contour edge in each frame. In the data preprocessing stage, the image collection needs to be frame aligned to eliminate the perspective offset caused by vehicle movement, and the spatiotemporal synchronization matching of multi-angle images is achieved through timestamp marking.
[0015] Step S200: performing multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image, and performing contour sharpening processing on the vehicle top profile image to generate a top feature enhanced image.
[0016] This step involves differentiated enhancement processing for images with different viewing angles. Multi-scale brightness compensation processing refers to the optimization of image brightness distribution using a hierarchical adjustment strategy, specifically including the combined application of global histogram equalization and local contrast limited adaptive histogram equalization (CLAHE). For example, for the overexposure problem of the window area caused by backlighting in the side profile image, the overall brightness is first adjusted to the standard range on a global scale, and then CLAHE processing is performed on detail areas such as wheels and door handles on a local scale to enhance dark textures while suppressing highlight overflow. Contour sharpening processing combines the Laplacian operator with nonlinear filtering to enhance the top edge features of the vehicle. For example, for the blurred boundary at the connection between the cargo box and the front of the vehicle in the top image, a frequency domain high-pass filter is used to extract the high-frequency component and superimpose it on the original image, making structural features such as welding seams and waterproof strips more prominent. During the implementation, the sharpening kernel size needs to be dynamically adjusted according to the image resolution: for high-resolution images (such as 4096×2160 pixels), a 5×5 sharpening template is used to cover a wider neighborhood information; for low-resolution images, a 3×3 template is used to avoid noise amplification. The processed side feature enhanced image can clearly present micro features such as tire tread depth and fender curvature, while the top feature enhanced image can accurately separate the overlapping area of the roof solar panel and the luggage rack.
[0017] Step S300: performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a multi-dimensional fusion feature map of the vehicle.
[0018] Dual-channel feature fusion is the integration of heterogeneous image data from the spatial dimension and the semantic dimension. In the spatial dimension, an affine transformation-based feature registration algorithm is used to project the top image into the coordinate system of the side image to ensure the spatial consistency of the cargo box length and the roof height. For example, by extracting the spatial correspondence between the vertex of the rearview mirror in the side image and the center line of the skylight in the top image, a three-dimensional projection matrix is established to achieve perspective alignment. In the semantic dimension, a dual-branch feature extraction network is constructed using deep separable convolution: the side feature branch focuses on learning linear features such as wheelbase and suspension height, while the top feature branch focuses on regional features such as container segmentation and cooling device layout. The output feature maps of the two branches are concatenated through channels and interact with the cross-channel information of the 1×1 convolution kernel to form a multi-dimensional fusion feature map containing geometric structure and surface texture. In specific implementation, the fusion weight is dynamically adjusted according to the differences between heavy trucks and small buses: for the multi-axis structure of trucks, the contribution of the side feature channel is increased to accurately capture the trailer connection point; for the streamlined roof of buses, the top feature channel is enhanced to distinguish the composite structure of the skylight and solar panels. The final fused feature map will integrate the tire footprint from the side view and the cargo box divider from the top view to form a joint representation with three-dimensional spatial semantics.
[0019] Step S400: inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model, and obtaining a vehicle model feature matching result output by the vehicle model matching model, wherein the vehicle model feature matching result includes vehicle wheelbase features, vehicle compartment structure features, and tire distribution features.
[0020] The pre-trained vehicle model matching model adopts a Transformer-based multi-task learning architecture, and its core consists of a feature encoder and an attribute decoder. The feature encoder captures the correlation between the wheelbase and the vehicle structure through a multi-head self-attention mechanism, such as identifying the proportional relationship between the distance between the front axle and the cab and the number of compartments in the cargo box; the attribute decoder outputs the wheelbase features (the quantized value is the pixel distance between the center points of adjacent tires converted into actual meters), vehicle structure features (classification labels such as flatbed, grid, and refrigerated) and tire distribution features (heat map marks the position of single and double tires and the number of parallel axles). When training the model, a multi-stage transfer learning strategy is adopted: first, the basic feature extraction layer is pre-trained on public data sets (such as COCO-Vehicles), and then fine-tuned on the labeled data of highway toll scenes. The loss function is designed as a weighted combination of cross entropy loss and mean square error loss to balance the optimization requirements of classification and regression tasks. During the actual reasoning process, for the occluded areas in the fusion feature map (such as tires covered by mud), the model infers the missing information through the contextual attention mechanism, such as estimating the number of tires based on the suspension height and the cargo box load. The output wheelbase features are accurate to centimeter-level accuracy, the cabin structure features support multi-layer label nesting (such as "refrigerated + double-open side doors"), and the tire distribution features can distinguish the tread pattern differences between the steering axle and the load-bearing axle.
[0021] Step S500: performing a similarity comparison based on the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark.
[0022] The pre-stored vehicle model feature library is a structured database, and each record contains the wheelbase range, car body type code and tire configuration template of the standard vehicle model. The similarity comparison adopts a multi-feature weighted Euclidean distance algorithm: the wheelbase feature weight is set to 0.5, the car body structure feature weight is 0.3, and the tire distribution feature weight is 0.2. For example, when the wheelbase of a certain vehicle is 3.2 meters (error tolerance ±0.1 meters), the car body structure is "double-layer warehouse grid", and the tire distribution is "single front and double rear", the system traverses all records in the feature library, calculates its comprehensive similarity score with each candidate model, and selects the model with the highest score that exceeds the threshold (such as 0.85) as the recognition result. For multiple candidates with similar scores (such as the comprehensive score difference between two types of trucks is less than 0.05), the manual review process is started and the conflicting feature items are recorded for model iteration and optimization. The vehicle identification mark generated finally follows the JT / T 489 standard coding rules and contains a 7-digit alphanumeric combination, of which the first two digits represent the vehicle category (such as "H" for trucks), the middle three digits identify the specific model, and the last two digits are the version variant code. This mark will be written into the transaction flow data of the toll collection system and linked to the rate calculation module to complete automatic deduction.
[0023] As an implementation manner, the step S200, performing multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image, may include the following steps: Step S210: extracting a continuous brightness distribution curve in the vehicle side profile image, wherein the continuous brightness distribution curve includes a first brightness distribution in a bottom area of the vehicle and a second brightness distribution in a top area of the vehicle.
[0024] In this process, the first brightness distribution of the bottom area of the vehicle refers to the pixel brightness value set within the vertical linear range from the point where the vehicle tire touches the ground to the chassis area, which reflects the impact of ground reflection, shadow occlusion and the reflective characteristics of metal parts on imaging; the second brightness distribution of the top area of the vehicle covers the brightness data set from the upper edge of the window to the highest point of the roof, which is affected by the natural light angle, the diffuse reflection coefficient of the paint material and the ambient light occlusion effect. For example, in a strong backlight scene, the bottom area of the vehicle side profile image may show a low brightness distribution (mean value less than 50, 8-bit grayscale range) due to the light absorption characteristics of the ground asphalt, while the top area has a high brightness concentration (mean value greater than 200) due to direct sunlight on the roof, forming a steep brightness gradient between the two. In order to accurately quantify this difference, it is necessary to sample the brightness value at each pixel interval along the vertical center axis of the vehicle, and use the cubic spline interpolation algorithm to generate a smooth continuous curve, where the horizontal axis of the curve represents the normalized height coordinate from the bottom to the top (0.0 to 1.0), and the vertical axis is the normalized brightness value of the corresponding position. During specific implementation, for high-box vehicles such as container trucks, the first brightness distribution in the bottom area needs to additionally separate the tire and container support frame areas to prevent the reflection of the metal bracket from interfering with the curve shape; and for small passenger cars, the second brightness distribution needs to focus on analyzing the brightness difference between the sunroof glass and the roof steel plate to ensure that the curve can accurately characterize the transition characteristics between translucent materials and opaque materials.
[0025] Step S220: generating a dynamic brightness compensation coefficient according to a gradient difference between the first brightness distribution and the second brightness distribution.
[0026] When the dynamic brightness compensation coefficient is generated, the spatial correlation between the first brightness distribution and the second brightness distribution is comprehensively evaluated. The gradient difference is defined as the integral of the brightness difference between the two distribution curves at the same height coordinate, which reflects the overall brightness contrast of the upper and lower areas in the side image of the vehicle. In the specific calculation, the two distribution curves are first aligned according to the height coordinate, and then the difference is calculated point by point and accumulated to obtain the total gradient difference value. For example, when it is detected that the mean value of the first brightness distribution in the bottom area of the vehicle reaches 180 due to direct tunnel light, and the mean value of the second brightness distribution in the top area is only 30 due to the absence of external light sources, the gradient difference value increases significantly (such as the cumulative difference exceeds the preset threshold of 15000). At this time, a high-intensity dynamic brightness compensation coefficient needs to be generated to balance the brightness difference. The coefficient is generated by piecewise linear function mapping, where a fixed coefficient of 1.0 is used to maintain the brightness of the original image when the gradient difference is low (such as less than 5000), and the compensation intensity is increased by a slope of 0.002 / unit when the difference is medium (5000 to 10000). When the difference is high (greater than 10000), the maximum coefficient is limited to 3.0 to prevent overexposure. For mixed traffic scenarios of vans and flatbed trailers, the segmentation interval needs to be dynamically adjusted according to the vehicle height: vehicles with a height of more than 3 meters use an extended threshold range (such as increasing the high difference threshold to 20,000) to adapt to the compensation needs of a large range of low-brightness areas on the top of the container.
[0027] Step S230: locally enhancing the low-brightness area in the vehicle side profile image according to the dynamic brightness compensation coefficient to generate an initial compensated image.
[0028] The implementation of local enhancement relies on the synergy between the dynamic brightness compensation coefficient and the regional segmentation results. The low brightness area refers to a set of pixels with grayscale values lower than 70% of the global mean, and its boundaries are determined by an adaptive threshold segmentation algorithm combined with a morphological closing operation. For example, for the side image of a vehicle taken at night, the connection between the tire and the chassis often forms a large area of low brightness (grayscale value 30-50) due to shadows. At this time, gamma correction (Gamma=0.4) is performed on this area according to the dynamic brightness compensation coefficient of 2.5, which significantly improves the texture visibility while avoiding the loss of details in high brightness areas (such as reflective license plates). The local window sliding mechanism is used in the enhancement process, and the window size is dynamically set according to the image resolution: for high-definition images of 4096×2160 pixels, a 51×51 pixel window is used to ensure local consistency; for standard-definition images of 720×480 pixels, a 15×15 pixel window is used to prevent noise amplification. The generation of the initial compensation image requires the simultaneous recording of metadata, including the compensation intensity distribution map of each region and the original brightness histogram transformation parameters, so that multi-scale analysis can be performed in subsequent steps.
[0029] Step S240: performing multi-scale filtering processing on the initial compensated image to extract edge texture features at different resolutions.
[0030] Multi-scale filtering processing realizes multi-resolution analysis by constructing a Gaussian pyramid. Specifically, the initial compensated image is downsampled by 1 / 2 in turn to generate a three-level pyramid (original image, 1 / 2 scale, 1 / 4 scale), and then the Laplacian operator, Sobel edge detection operator and Gabor filter group are applied at each level. For example, at the original image scale, a 5×5 Laplacian kernel is used to extract fine edges such as door gaps and tire patterns; at the 1 / 2 scale, the Sobel horizontal operator is used to detect the continuity of the welding line of the side panel of the car; at the 1 / 4 scale, a Gabor filter (wavelength 16 pixels, directions 0°, 45°, 90°, 135°) is applied to extract the periodic texture of the corrugated plate of the container. In this process, the filtering results at each scale are upsampled to the original image size through bilinear interpolation, and superimposed to generate a multi-channel edge texture feature map. For the concave-convex insulation layer texture unique to refrigerated trucks, the 1 / 4 scale Gabor filter output can effectively enhance the periodic features of 0.5-1.2 mm / pixel, while the smooth side panels of ordinary trucks are mainly characterized by the Laplace response of the original image scale.
[0031] Step S250: adjusting filtering parameters according to the density distribution of the edge texture features to generate the side feature enhanced image.
[0032] The analysis of density distribution can be achieved, for example, by calculating the number of edge pixels per unit area, which determines the contribution weight of each scale filter. For example, when the edge density of the door area is detected to be as high as 120 pixels / cm2 (corresponding to the result of the original image scale Laplacian filter), the weight of the 1 / 4 scale Gabor filter is reduced to 0.3 to avoid the moiré effect caused by the superposition of high-frequency textures; on the contrary, for the area where the edge density of the container flatbed area is less than 20 pixels / cm2, the weight of the 1 / 2 scale Sobel filter is increased to 0.7 to enhance the long straight line features. After the parameters are adjusted, the weighted fusion algorithm is used to combine the multi-scale feature map with the initial compensation image: the edge-dense area focuses on retaining the original image scale details (fusion coefficient 0.6), and the flat area enhances the 1 / 2 scale contour information (fusion coefficient 0.8). The resulting side feature enhanced image is globally optimized by contrast limited adaptive histogram equalization (CLAHE), where the tile grid size is set to 64×64, the contrast limit threshold is 2.0, and the histogram group is 256. The image can clearly show key features such as the degree of tire tread wear, the geometry of the door handle and micro-cracks in the reflective strips, providing high-fidelity input data for subsequent vehicle model identification.
[0033] As an implementation manner, the step S300, performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a multi-dimensional fusion feature map of the vehicle, includes the following steps: Step S310: performing spatial coordinate transformation on the side feature enhanced image to obtain a set of key contour points in the side image coordinate system.
[0034] In this step, the implementation of spatial coordinate transformation aims to establish the mapping relationship between the side features of the vehicle and the three-dimensional physical space. The side image coordinate system refers to a two-dimensional pixel coordinate system with the upper left corner of the image as the origin, the horizontal right as the positive direction of the X axis, and the vertical downward as the positive direction of the Y axis. The key contour point set is the coordinate set of the significant structural feature points on the side of the vehicle extracted by the edge detection algorithm, including the door hinge vertex, the tire ground contact center point and the outer edge end point of the rearview mirror. For example, for the side feature enhancement image of the van, the Canny edge detector is combined with the Hough transform line detection to accurately identify the intersection of the vertical edge of the container and the horizontal support beam of the chassis as the key contour point. After the coordinate value is optimized by the sub-pixel interpolation algorithm, it can reach 0.1 pixel accuracy. For the side of the bus with complex curved surface structure, it is necessary to introduce the curvature continuity constraint condition to select the curvature extreme point of the door curved window frame and the corner point of the roof luggage rack to add to the key contour point set. During the implementation process, the edge detection threshold needs to be dynamically adjusted according to the vehicle type: the flatbed trailer has a high proportion of straight edges, and a high threshold (such as Canny high threshold 120) is used to suppress the noise points generated by rivet holes; for streamlined coupes, the threshold is lowered to 60 to retain the weak edge features of the arc transition area.
[0035] Step S320: performing spatial coordinate transformation on the top feature enhanced image to obtain a set of key contour points in the top image coordinate system.
[0036] The top image coordinate system is defined in the same way as the side image coordinate system, but its key contour point set focuses on the geometric features of the top structure of the vehicle, including the vertices of the four corners of the skylight, the installation point of the air conditioner outdoor unit, and the end points of the transverse reinforcement ribs of the container roof. Due to the occlusion of the roof curvature and installation accessories in the top view, a multi-stage feature extraction strategy is required: first, the gap noise between the solar panel and the roof steel plate is filled by morphological closing operation, and then the intersection of the container roof welding seam is located by Harris corner detector, and then extended to the complete contour line in combination with the region growing algorithm. For example, in the feature enhancement image of the top of the refrigerated truck, the coordinate error of the intersection of the refrigeration unit vent grille does not exceed 0.3 pixels after sub-pixel correction; and in the top image of the container truck, the center coordinates of the container corner lock hole can be accurately extracted by circular Hough transform. For images with perspective distortion (such as shooting angles with a pitch angle of more than 15°), projection distortion correction must be performed first, the top image coordinate system must be converted to the orthographic projection plane, and then key point extraction must be performed to ensure that the ratio error between the measured values of the length and width of the container and the actual physical size is less than 1%.
[0037] Step S330: Perform three-dimensional spatial mapping on the key contour point set in the side image coordinate system and the key contour point set in the top image coordinate system to generate a three-dimensional contour model of the vehicle.
[0038] The core of three-dimensional spatial mapping is to establish the spatial correspondence between the key contour points on the side and top. It is necessary to calculate the external parameter matrix of the binocular vision system based on the calibration plate or a reference object of known size. Specifically, the spatial position of the tire ground contact center point in the side image coordinate system and the center point of the front edge of the container in the top image coordinate system is fitted by the least squares method, and the three-dimensional point cloud data is constructed by combining the camera focal length, pixel size and baseline distance parameters. For example, when the coordinates of the left rear tire ground contact center point in the side image are (x1, y1) and the coordinates of the corresponding left front corner point of the container in the top image are (x2, y2), the three-dimensional coordinates (X, Y, Z) are calculated by the triangulation principle, and the projection error of all matching point pairs is iteratively optimized until the root mean square error is less than 0.5 pixels. The generated vehicle three-dimensional contour model is composed of triangular mesh patches, and the mesh vertex density is adaptively adjusted according to the curvature change: high curvature areas (such as the rearview mirror surface) use dense vertices with a spacing of 5 mm, and flat areas (such as the side of the container) are sparsely sampled with a spacing of 20 mm. For special vehicles (such as tankers), it is necessary to introduce additional cylindrical parametric model constraints to fit the point cloud of the tank part into a parametric surface with a constant radius and continuously changing axis direction to ensure that the calculation accuracy of the tank volume reaches more than 98%.
[0039] Step S340: extracting the curvature features of the vehicle surface according to the curvature distribution of each contour point in the three-dimensional contour model of the vehicle.
[0040] The calculation of curvature distribution is based on the local surface differential geometry properties of each vertex in the triangular mesh model, and a combination of Gaussian curvature and mean curvature is used to characterize the concave-convex characteristics of the vehicle surface. For each vertex, the principal curvatures k1 and k2 are calculated by the rate of change of the normal vectors of its adjacent triangular facets, and then the Gaussian curvature K=k1*k2 and the mean curvature H=(k1+k2) / 2 are obtained. For example, the Gaussian curvature of the plane area of the side panel of the van container approaches zero, and the mean curvature is also close to zero; while the curved deflector on the top of the refrigerated truck presents negative Gaussian curvature (K<0) and non-zero mean curvature (H≈0.03 / mm). The extraction of curvature features requires setting multiple thresholds: areas with an absolute value of Gaussian curvature greater than 0.05 / mm² are marked as high curvature feature areas (such as the recessed area of the door handle), areas between 0.01 and 0.05 / mm² are defined as medium curvature transition areas (such as tire tread patterns), and areas less than 0.01 / mm² are considered low curvature flat areas (such as car window glass). During the implementation process, the curvature gradient histogram is used to statistically analyze the distribution of tiny wrinkles on metal stamping parts (curvature fluctuation range 0.02-0.08 / mm²) and filter out outliers caused by point cloud noise (points with gradient mutations exceeding 3σ).
[0041] Step S350: superimpose and fuse the surface curvature features of the vehicle with the edge texture features of the side feature enhanced image to generate a multi-dimensional fusion feature map of the vehicle.
[0042] The superposition fusion process adopts a hybrid strategy combining channel weighted splicing and spatial attention mechanism. First, the curvature features of the vehicle surface are converted into a two-dimensional feature map with the same resolution as the side feature enhancement image, where the channel value of each pixel corresponds to the three-dimensional curvature property of its projection position (linear combination of Gaussian curvature and mean curvature). Subsequently, the significance weight of each pixel in the edge texture feature map is calculated through the spatial attention module: for edge pixels corresponding to high curvature areas (such as door handles), a fusion weight of 0.7 is assigned to enhance the three-dimensional geometric characteristics; for low curvature areas (such as container side panels), it is reduced to 0.3, focusing on retaining two-dimensional texture details. For example, after the corrugated texture of the insulation layer in the side feature enhancement image of the refrigerated truck (edge density 60 pixels / square centimeter) and the Gaussian curvature feature of the top deflector (K=-0.04 / mm²) are fused, the generated feature map can simultaneously present the corrugation period (two-dimensional attribute) and the deflection curvature (three-dimensional attribute). The final multi-dimensional fused feature map is uniformly scaled to a standard size (such as 1024×512 pixels) through the bicubic interpolation algorithm, and normalized (each channel value is scaled to the [0,1] interval) to provide input data with both geometric structure and surface details for the subsequent vehicle model matching model.
[0043] As an implementation manner, the step S400, inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model, and obtaining the vehicle model feature matching result output by the vehicle model matching model, includes the following steps: Step S410: Divide the vehicle multi-dimensional fusion feature map into multiple local feature areas, each local feature area corresponds to a feature type among vehicle wheelbase features, vehicle body structure features or tire distribution features.
[0044] The division of the multi-dimensional fusion feature map of the vehicle needs to be based on the prior knowledge of the physical structure of the vehicle for spatial region annotation. The local feature region refers to a specific range defined by a rectangular box or a semantic segmentation mask in the feature map. The vehicle wheelbase feature region covers the horizontal projection range from the center point of the front axle to the center point of the rear axle of the vehicle, the cabin structure feature region includes the vertical section and top arc structure of the connection between the cab and the cargo box, and the tire distribution feature region focuses on the spatial distribution of each tire contact point and the tread texture. For example, for the fusion feature map of a container truck, the wheelbase feature region defines a rectangular region with a width of 50 pixels and a length of 80% of the total length of the feature map along the longitudinal beam direction of the frame, accurately covering the wheelbase connecting the tractor and the trailer; the cabin structure feature region uses a polygonal mask to mark the hexahedral contour and vent position of the refrigerated container; the tire distribution feature region detects a 50×50 pixel range around each tire contact point through a sliding window to ensure that the tire spacing measurement accuracy of the dual-tire structure reaches ±2 pixels. During the division process, the area size needs to be dynamically adjusted according to the vehicle type: the wheelbase characteristic area of the flatbed trailer needs to be expanded to include multiple sets of parallel axles, while the cabin structure characteristic area of the small passenger car needs to be shrunk to the junction of the roof sunroof and the trunk lid.
[0045] Step S420: performing feature vectorization processing on each local feature region to generate a local feature vector set.
[0046] The feature vectorization is achieved by combining depthwise separable convolution with global average pooling, which converts the two-dimensional feature map into a fixed-dimensional numerical vector. For the vehicle wheelbase feature area, a 3×3 convolution kernel is used to extract the horizontal feature response, capturing the relative distance between adjacent axles and the change in suspension height; the vehicle body structure feature area uses a 5×5 hole convolution to expand the receptive field and effectively identify the spatial relationship between the cargo box partition and the roof arc; the tire distribution feature area uses a fusion descriptor of the Histogram of Oriented Gradients (HOG) and the Local Binary Pattern (LBP) to quantify the tread pattern directionality and shoulder wear. For example, the convolution output of the wheelbase feature area of the refrigerated truck generates a 128-dimensional vector after global average pooling, and its peak component corresponds to the 3.2-meter feature of the distance between the third and fourth axes; in the structural feature vector of the container, the values of the 56th-72nd dimensions are significantly higher than other parts, representing the grid density of the corner reinforcement plate of the container; and the distribution feature vector of the twin tires is combined through the 9-directional gradient histogram of HOG and the 59 texture patterns of LBP to form a 247-dimensional hybrid descriptor.
[0047] Step S430: inputting the local feature vector set into the multi-layer convolutional network in the vehicle model matching model to obtain a feature response map output by each convolutional layer.
[0048] The multi-level convolutional network consists of three parallel branches, which process the wheelbase, car body and tire feature vectors respectively. The wheelbase branch uses the Temporal Convolutional Network (TCN) to capture the sequence dependency of multi-axle vehicles along the wheelbase direction; the car body branch uses the Residual Dense Network (RDN) to enhance the representation ability of complex structures through cross-layer feature reuse; the tire branch deploys the Graph Convolutional Network (GCN) to construct the tire position as a graph node and model the spatial topological relationship. For example, when the wheelbase feature vector of an 8-axle heavy-duty tractor is input, in the fifth-layer output feature response map of TCN, the third channel shows significant activation at the fourth axis position (response value 0.92), indicating that this is the steering axis feature; when RDN processes the structural features of refrigerated containers, the jump connection between dense blocks increases the response intensity of the vent area in the 12th-layer feature map by 40%; GCN constructs edge connections between each tire node and its adjacent tires. After three layers of graph convolution iterations, the feature similarity of the dual-tire nodes reaches 0.87, which is significantly higher than the 0.32 of the single-tire nodes.
[0049] Step S440: Determine the wheelbase matching degree corresponding to the vehicle wheelbase feature, the vehicle compartment matching degree corresponding to the vehicle compartment structure feature, and the tire matching degree corresponding to the tire distribution feature according to the activation area distribution of the feature response diagram.
[0050] The analysis of the distribution of activation areas adopts a strategy combining spatial weighted pooling and saliency detection. The wheelbase matching degree is calculated by calculating the degree of agreement between the interval of activation peaks at each axis position in the TCN output feature map and the actual calibration value. For example, if it is detected that there are three peaks with a standard deviation exceeding 2.5σ at 3.2 meters, 7.8 meters, and 11.4 meters in the feature response map, then dynamic time warping (DTW) matching is performed with the three-axis spacing template in the vehicle model library to obtain a wheelbase matching degree of 0.89; the car body matching degree is calculated by cosine similarity of the high-dimensional features output by RDN, such as the angle between the layout of the vents of the refrigerated container and the 512-dimensional vector of the template is 18°, corresponding to a similarity of 0.95; the tire matching degree requires a comprehensive tread direction consistency (HOG histogram chi-square distance) and distribution topology matching (GCN node embedding Euclidean distance). For example, the cumulative value of the tread direction difference of a certain vehicle is 125, the topological distance is 0.34, and the tire matching degree of 0.76 is obtained after normalization.
[0051] Step S450: weighted fusion of the wheelbase matching degree, the vehicle compartment matching degree and the tire matching degree to generate the vehicle model feature matching result.
[0052] Weighted fusion uses adaptive weight allocation based on feature importance, where the initial weight of wheelbase is 0.5, car body 0.3, and tire 0.2, but it is dynamically adjusted according to the vehicle type: when a complex cargo box structure (such as multi-layer partitions of a refrigerated truck) is detected, the car body weight is increased to 0.4; for multi-axle vehicles (such as 5-axle tank trucks), the wheelbase weight is increased to 0.6. The fusion formula is: comprehensive matching degree = wheelbase matching degree × W1 + car body matching degree × W2 + tire matching degree × W3. For example, the wheelbase matching degree of a container truck is 0.92 (W1=0.5), the car body matching degree is 0.88 (W2=0.3), and the tire matching degree is 0.81 (W3=0.2). The comprehensive value is 0.5×0.92+0.3×0.88+0.2×0.81=0.892; if it is an 8-axle flatbed trailer, the wheelbase weight is increased to 0.6 due to the increase in the number of axles, then W1=0.6, W2=0.25, and W3=0.15 during calculation. The calculation results are mapped to the [0,1] interval by the Sigmoid function. The final vehicle model feature matching result includes the matching degree of each sub-item and the comprehensive value, and the confidence interval is marked (such as 0.892±0.03).
[0053] Based on this, as an implementation method, the step S500, performing a similarity comparison between the vehicle model feature matching result and a pre-stored vehicle model feature library, determining the toll vehicle model category corresponding to the passing vehicle and generating a vehicle model identification mark, includes the following steps: Step S510: extracting template data of each vehicle model from the vehicle model feature library, wherein the vehicle model template data includes a template wheelbase range, a template vehicle body structure template and a template tire distribution template.
[0054] The vehicle model template data is stored in a structured form. Each record contains the template wheelbase range (such as 3.2-3.5 meters, 7.6-7.9 meters, etc.), the template car body structure template (512-dimensional feature vector) and the template tire distribution template (including the number of tires, tire position coordinates and tread coding topology map). For example, the refrigerated truck template numbered HT-45C in the vehicle model library contains three wheelbase ranges [3.18, 3.22], [7.75, 7.85], and [11.3, 11.5] meters. The car body structure template corresponds to the feature vector of a wavy container with a vent spacing of 600mm. The tire distribution template defines 6 sets of dual-tire parallel nodes and their standard value of tread direction angle of 32°. The construction of the template data requires taking the average of the feature data of 200 vehicles of the same model, and calculating the allowable fluctuation range of ±3σ to ensure that the manufacturing tolerance and measurement error are covered.
[0055] Step S520: Calculate the first similarity between the wheelbase matching degree and the template wheelbase range, calculate the second similarity between the car body matching degree and the template car body structure template, and calculate the third similarity between the tire matching degree and the template tire distribution template.
[0056] The first similarity is calculated by combining the interval overlap ratio with Gaussian weighting: if the measured wheelbase falls within the template range, the basic similarity is 1.0; if it deviates, the formula exp(-(d / σ) 2 ), where d is the deviation and σ is the template standard deviation. For example, the measured wheelbase of 3.25 meters exceeds the upper limit of the third segment template of HT-45C by 0.05 meters. If σ=0.1 meters, the similarity of the third segment is exp(-(0.05 / 0.1) 2 )=0.778; the second similarity measures the directional consistency between the measured car feature vector and the template vector through cosine similarity. For example, when the angle between the measured vector and the template is 15°, the similarity is cos(15°)=0.966; the third similarity needs to integrate the tire number matching degree (such as 12 tires in the measured vs. 12 tires in the template, which is 1.0), tire position topology difference (the average Euclidean distance of node coordinates is 0.2 meters, which is 0.82 after Sigmoid transformation) and tread direction difference (histogram intersection and union ratio is 0.75), and the weighted average is 0.86.
[0057] Step S530: Determine the comprehensive matching degree between the passing vehicle and each vehicle model template data according to the weighted sum of the first similarity, the second similarity and the third similarity.
[0058] The weighting coefficient is dynamically adjusted according to the vehicle type classification: freight vehicles use a wheelbase of 0.5, a compartment of 0.3, and a tire of 0.2; passenger vehicles focus on the comfort characteristics of the compartment, and the weights are set to 0.4 for the wheelbase, 0.4 for the compartment, and 0.2 for the tire. For example, the three similarities of a vehicle suspected to be HT-45C are 0.92 (wheelbase), 0.94 (compartment), and 0.88 (tire), and the comprehensive matching degree under the freight weight is 0.5×0.92+0.3×0.94+0.2×0.88=0.918; if the vehicle is a passenger version, it is calculated according to the passenger weight to be 0.4×0.92+0.4×0.94+0.2×0.88=0.924. The calculation results are arranged in descending order to generate a candidate list, such as HT-45C (0.918), HT-46B (0.895), and GT-32D (0.872), and the matching degree of each sub-item is marked for manual review and reference.
[0059] Step S540: The vehicle model number corresponding to the vehicle model template data with the highest comprehensive matching degree is used as the toll vehicle model category, and a vehicle model identification mark including the vehicle model number is generated.
[0060] The conversion of vehicle model numbers must comply with the JT / T 489-2019 standard, where the first two letters represent the vehicle model category (such as HT for refrigerated trucks), the middle two digits are the load rating (45 for 45 tons), and the last letter is the version number (C for the third edition). The vehicle model identification mark is encapsulated in JSON format and contains the following fields: {"Model number":"HT-45C", "Matching degree":0.918, "Timestamp":"2023-08-20T14:23:05Z", "Wheelbase details":[{"Position":"First wheelbase","Measured value":3.21,"Template range":[3.18,3.22]},...]}. The mark is encoded in Base64 and written into the transaction flow of the toll collection system, and at the same time triggers the rate calculation module to call the 0.45 yuan / ton·km rate corresponding to HT-45C. For vehicles with a comprehensive matching degree lower than 0.85 (such as 0.832), the system automatically triggers the high-definition re-shooting process and pushes the queue for review to the toll collector's console.
[0061] As an implementation mode, the method further includes a pre-training step of the vehicle type matching model, including the following steps: Step S10: Obtain a historical vehicle training set, wherein the historical vehicle training set includes multiple groups of vehicle image samples and corresponding annotated vehicle model feature data.
[0062] The construction of the historical vehicle training set needs to cover diverse data of different lighting conditions, shooting angles and vehicle models. Vehicle image samples refer to the original image sequences collected by the highway toll lane monitoring system. Each set of samples contains 10 consecutive frames of vehicle side view images and 5 frames of top view images. The frame rate is set to 30fps to ensure that motion blur is controlled within a processable range. The labeled vehicle model feature data adopts a structured label format, including vehicle wheelbase features (actual measured values of the three-axle spacing accurate to millimeters), vehicle body structure features (classification codes based on the JT / T 489 standard, such as "HT-45C" for 45-ton refrigerated trucks) and tire distribution features (pixel coordinates of each tire contact point and tread type code). For example, for a batch of training data collected in 2019, the vehicle image samples included images of high light interference in the window area caused by water film reflection in rainy environments. The wheelbase features in the labeled data were marked as "3.21±0.03 meters, 7.85±0.05 meters, 11.42±0.07 meters", the compartment structure features were marked as "double-layer partition refrigerated container", and the tire distribution features were marked as "12R22.5 vacuum tire, dual tires installed with a spacing of 200±5 mm". The training set needs to remove duplicate vehicle data through a spatiotemporal deduplication algorithm and be divided into training subsets, validation subsets, and test subsets in a ratio of 8:1:1 to ensure the generalization ability of the model.
[0063] Step S20: performing image segmentation processing on each group of vehicle image samples to extract vehicle side training images and vehicle top training images.
[0064] Image segmentation processing uses an instance segmentation algorithm based on Mask R-CNN, combined with the three-dimensional geometric constraints of the vehicle to separate the target area. The segmentation of the vehicle side training image requires generating the minimum circumscribed rectangle along the longitudinal axis of the vehicle, cutting off the guardrail, green belt and other vehicle interference areas in the background; the vehicle top training image uses a top-down projection transformation to correct the trapezoidal distortion area in the original image into a rectangular area, and crop it to the range of 10 pixels outside the outer boundary of the roof. For example, for an articulated truck with a trailer, the segmentation algorithm first identifies the junction between the tractor cab and the trailer container, generates a side mask along the extension direction of the trailer longitudinal beam, and ensures that the container side panel and the tractor rear wheel are included in the side training image at the same time; when segmenting the top image, the fisheye correction model is used to eliminate the bending of the container top cover edge caused by the wide-angle lens, and the oblique view is converted into an orthographic projection view through the perspective transformation matrix, so that the length and width ratio of the container is restored to the actual value of 1:2.5. For low-contrast areas in night images (such as the tire outline of a black car body), the auxiliary edge enhancement module needs to be enabled, combining Canny edge detection with morphological closing operations to fill broken boundaries and ensure the integrity of the closed area of the segmentation mask.
[0065] Step S30: performing noise suppression and distortion correction on the vehicle side training image to generate a side training sample set.
[0066] Noise suppression uses a cascade of the Non-Local Means (NLM) algorithm and Block-Matching and 3D filtering (BM3D) to eliminate high-frequency noise (such as raindrop noise) and low-frequency noise (such as sensor thermal noise) in layers. For example, for the side training images collected in rainy and foggy weather, the NLM algorithm (search window 21×21 pixels, similarity window 7×7 pixels, filter parameter h=15) is first applied to smooth the rain texture, and then BM3D (block size 8×8, hard threshold shrinkage) is used to remove the residual salt and pepper noise, so that the SSIM (structural similarity index) of the tire tread pattern is increased from 0.65 to 0.89. Distortion correction requires the construction of a radial distortion model based on the camera calibration parameters, and the k1, k2, and k3 coefficients in the Brown-Conrady model are used to reversely compensate for the image edge stretching. In specific implementation, for a surveillance camera with a focal length of 8mm, when the straight curvature of the frame longitudinal beam in the side image is detected to exceed 3 pixels, the correction parameters of k1=-0.25 and k2=0.06 are applied to control the straightness error within 0.5 pixels. The corrected side training sample set also needs to be normalized for brightness, adjusting the image grayscale histogram mean to 128±10 and the standard deviation to 45±5 to enhance the recognizability of dark details.
[0067] Step S40: performing shadow removal and contrast enhancement on the vehicle top training image to generate a top training sample set.
[0068] Shadow elimination can be achieved, for example, by inverting the physically based rendering equation. Specifically, the Light Estimation Network (LEN) is used to predict the direction and intensity of the main light source in the scene, and then the shadow mask generation algorithm is used to separate the projected shadow area on the roof surface (such as the strip shadow formed by the gap between the containers). For example, for the top image taken at noon, the LEN output light source azimuth angle is 120°, altitude angle is 60°, and intensity is 85,000 lux. Based on this, the theoretical projection range of the shadow is calculated, and the brightness of the shadow area is increased to the same as the non-shadow area through the Poisson image editing algorithm. Contrast enhancement uses an improved Retinex algorithm to decompose the image into reflection component and illumination component: the reflection component extracts the surface texture through multi-scale Gaussian filtering (σ=15, 80, 200 pixels), and the illumination component compresses the dynamic range through gamma correction (γ=0.6). After implementation, the contrast between the rusted area and the waterproof strip on the top of the container increased from the original value of 1.8 to 4.2, and the edge sharpness increased by 30%. The processed top training sample set needs to be geometrically aligned and verified, and a feature point matching algorithm (such as SIFT) is used to check the projection consistency between the corrected image and the 3D point cloud model to ensure that the coordinate error of the container corner points does not exceed ±2 pixels.
[0069] Step S50: Input the side training sample set and the top training sample set into the initial convolutional neural network for multi-stage training, adjust the network weight parameters until the error between the output features of the initial convolutional neural network and the labeled vehicle model feature data is lower than a preset threshold, and generate the vehicle model matching model.
[0070] For example, the initial convolutional neural network can be designed with a two-stream architecture, with ResNeXt-101 and DenseNet-201 used as the backbone network for the side and top image branches, respectively, and features fused through a cross-modal attention mechanism. The multi-stage training includes three stages: in the first stage, the top branch parameters are frozen, and only the wheelbase regression layer and the car classification layer of the side branch are trained. The learning rate is set to 1e-4, the batch size is 32, and the smooth L1 loss function is used to optimize the wheelbase prediction error; in the second stage, the top branch is unfrozen, and the tire distribution prediction module of the two-stream network is jointly trained. The learning rate is reduced to 5e-5, the batch size is 16, and the loss function is changed to a mixed loss (wheelbase MSE loss + car cross entropy loss + tire contrast loss); in the third stage, all parameters are enabled for end-to-end fine-tuning, and a curriculum learning strategy is introduced. The training data is gradually increased from easy to difficult according to the sample difficulty. The learning rate is scheduled using cosine annealing (initial value 3e-5, cycle 20 epochs). During the training process, the comprehensive error of the validation subset (wheelbase MAE + car body F1 score + tire IoU) needs to be evaluated once per epoch, and the early stopping mechanism is triggered when the error decreases by less than 0.1% for 5 consecutive epochs. For example, in a certain training, the wheelbase mean absolute error (MAE) of the initial model on the test subset was 38 mm, which dropped to 12 mm after 120 epochs of training; the macro-average F1 score of the car body structure classification increased from 0.76 to 0.93; the IoU (intersection over union) of the tire distribution was optimized from 0.65 to 0.88, and the final model was saved as a vehicle model matching model after meeting the preset thresholds (wheelbase MAE ≤ 15 mm, F1 ≥ 0.9, IoU ≥ 0.85). When the model is deployed, the floating-point model is quantized to INT8 format through the TensorRT engine, and the inference delay is reduced from 210ms to 45ms, meeting the real-time processing requirements of the toll station.
[0071] As an implementation manner, the step S20, performing image segmentation processing on each group of vehicle image samples to extract vehicle side training images and vehicle top training images, includes the following steps: Step S21: performing background separation processing on the vehicle image sample to obtain a vehicle main area mask.
[0072] Background separation processing can be achieved, for example, through a semantic segmentation network based on deep learning. The vehicle main area mask refers to a set of vehicle pixel positions represented in the form of a binary image, where foreground pixels (value 1) represent the vehicle entity area, and background pixels (value 0) represent irrelevant areas such as roads, guardrails or the sky. In specific implementation, the DeepLabv3+ network architecture is adopted, and its backbone network is Xception-65. The input image is initialized by the model weights pre-trained on the highway vehicle dataset. After the multi-scale features are extracted by the spatial pyramid pooling module, the mask image with the same resolution as the input is output. For example, for image samples containing container trucks, the network can accurately separate the side panels of the container from the background of the green belt behind. Even if the surface of the container has a light blue paint similar to the color of the sky, the vehicle body boundary can still be identified through high-level semantic features to ensure that the mask edge error does not exceed 3 pixels. For the segmentation of the highlight area of the headlights in night images, it is necessary to increase the edge-sensitive weight in the loss function to enhance the ability to distinguish the taillight outline from the reflective road surface.
[0073] Step S22: Divide the vehicle side area and the vehicle top area according to the geometric center position of the vehicle main area mask.
[0074] The calculation of the geometric center position is based on the minimum circumscribed rectangle of the vehicle main area mask, and the coordinates of the center point are the pixel coordinates of the intersection of the diagonal lines of the rectangle. The side area of the vehicle is defined as a rectangular range extending vertically downward from the center point to the ground line under the vehicle, covering the tires, chassis and door structure; the top area of the vehicle is the area extending vertically upward from the center point to the highest point of the roof, including components such as the container top cover, air conditioner outdoor unit and antenna. For example, for a refrigerated truck with a height of 4.2 meters, the geometric center of the mask is located at 380 pixels of the image ordinate (the total image height is 720 pixels), the side area is defined as the range of 380 to 720 pixels of the ordinate, and the top area is the range of 0 to 380 pixels of the ordinate. For special models such as articulated trailers, the division ratio needs to be dynamically adjusted according to the trailer connection point: the tractor part is divided according to the standard ratio, and the trailer part is divided twice according to the geometric center of its independent mask to ensure that the side and top areas of the front and rear sections of the container are fully covered.
[0075] Step S23: performing edge tracking on the side area of the vehicle to extract a continuous and closed side contour line.
[0076] For example, edge tracking can use an improved boundary tracking algorithm, starting from the lower left corner of the vehicle side area mask, traversing the pixel neighborhood in a counterclockwise direction, recording the boundary point sequence that meets the 8-connectivity condition, and finally generating a closed contour line composed of polygonal polylines. For example, when processing the side area of a refrigerated truck, the algorithm starts tracking from the left rear tire contact point, rises along the vertical line of the container side panel to the arc transition area of the roof, then extends horizontally to the right to the right front wheel fender, and finally closes downward to the starting point, forming a polygonal contour line containing 256 vertices. To ensure the smoothness of the contour line, the initial tracking result needs to be compressed by the Douglas-Peucker algorithm to reduce the number of redundant points to 30% of the original data, while keeping the contour shape error less than 1.5 pixels. For the inner edges generated by recessed structures such as door handles and reflective strips, the concave points are identified by calculating the curvature change rate, and control points are inserted into the contour line to retain the detailed features.
[0077] Step S24: determining a cropping boundary of the vehicle side training image according to the curvature change point of the side contour line.
[0078] The curvature change point refers to the local extreme point on the contour line where the curvature value exceeds the preset threshold, reflecting the geometric mutation position of the vehicle side structure. The curvature calculation adopts the discrete differential geometry method. For each point Pi in the sequence of contour line vertices, the ratio of the chord length and arc length formed by the k points before and after it (k=5) is calculated. When the ratio exceeds the threshold of 0.15, it is marked as a curvature change point. For example, in the side contour of a container truck, the curvature ratio at the connection between the front edge of the container and the cab is 0.22, which is identified as a key cropping boundary point; the periodic concave and convex structure of the corrugated plate of the refrigerated truck container produces evenly spaced curvature change points, and the interval distance corresponds to the corrugation spacing (about 200 mm). The cropping boundary is determined by the horizontal projection of the curvature change point: a straight line is extended in the vertical direction to the image boundary to form a rectangular area containing the complete side of the vehicle. If the distance between adjacent curvature change points is too small (such as less than 50 pixels), Bezier curve fitting is used to generate a smooth transition boundary to avoid jagged cropping affecting subsequent feature extraction.
[0079] Step S25: segmenting the vehicle image sample according to the cropping boundary to generate the vehicle side training image and the vehicle top training image.
[0080] The segmentation operation can be achieved by combining image cropping with mask bit operations: the vehicle side training image is the logical AND result of the pixel set of the original image within the side cropping boundary and the vehicle main area mask; the vehicle top training image is the intersection of the top area cropping result and the mask. For example, for a van image sample, the side cropping boundary is a rectangular area from 120 pixels on the left to 880 pixels on the right and 200 to 680 pixels on the vertical axis, generating a side training image with an image resolution of 760×480 pixels, in which only the door, tire and container side panel area are retained; the top training image is cropped into a 320×1280 pixel area from 0 to 200 pixels on the vertical axis, focusing on the container roof and fairing structure. For vehicles with abnormal heights (such as engineering forklifts), a dynamic scaling strategy is adopted: if the height of the cropped image exceeds the standard size (such as 512 pixels), it is proportionally reduced to the target size and the blank area on the edge is filled, keeping the aspect ratio unchanged while ensuring the integrity of key structures.
[0081] As an implementation manner, the step S30 of performing noise suppression and distortion correction on the vehicle side training image to generate a side training sample set includes the following steps: Step S31: Detect the high-frequency noise area in the vehicle side training image and generate a noise distribution map.
[0082] High-frequency noise areas refer to interference signals in the image with spatial frequencies higher than the structural features of the vehicle, such as those caused by sensor thermal noise, raindrops, or dust attached to the lens. The noise distribution map is generated by wavelet transform decomposition: the image is decomposed by 3 layers of Haar wavelets, and the sum of the absolute values of the high-frequency subbands (HL, LH, HH) is extracted as the noise intensity indicator. For example, in the vehicle side training image collected on a rainy day, the response value of raindrop noise in the HH subband (diagonal high frequency) exceeds 120 (8-bit grayscale range), forming a scattered noise area; while the sensor noise is evenly distributed in the HL subband (horizontal high frequency), and the intensity value is maintained between 30-50. The noise distribution map is represented in the form of a heat map, and the closer the color is to red, the higher the noise intensity. For periodic noise (such as LED strobe stripes), it is necessary to perform an additional Fourier transform to identify specific frequency components and mark the corresponding frequency domain coordinate area in the distribution map.
[0083] Step S32: adjusting the window size of the adaptive filter according to the density value of the noise distribution map, and performing smoothing on the high-frequency noise area.
[0084] The adaptive filter window size is negatively correlated with the noise density value. For example, a 3×3 window is used to preserve details in areas with density values higher than a threshold (such as 150 noise points per square centimeter); a 5×5 window is used to balance denoising and blurring in areas with medium density values (50-150 points / square centimeter); and a 7×7 window is used to achieve strong smoothing in low-density areas (less than 50 points). For example, for large areas of rust on the container side panels (noise density 280 points / square centimeter), a 3×3 window bilateral filter is used to suppress isolated noise points while retaining the rust texture; while a 7×7 Gaussian filter is used to eliminate sensor noise in areas with uniform door paint (density 40 points). The local noise density needs to be calculated dynamically during implementation: the sum of the pixel values of the noise distribution map in the area is counted with a sliding window (default 32×32 pixels) and normalized to a density value by area. For boundary areas (such as where the tire contacts the ground), an asymmetric window is used to avoid cross-region mixing, and the horizontal size of the window is kept at 5 pixels, while the vertical size is reduced to 3 pixels to adapt to the edge direction.
[0085] Step S33: extracting perspective distortion parameters from the vehicle side training image, wherein the perspective distortion parameters include a horizontal tilt angle and a vertical scaling ratio.
[0086] For example, perspective distortion parameters are calculated with the aid of calibration plates or estimated based on the geometric constraints of the vehicle structure. The horizontal tilt angle refers to the angle between the longitudinal axis of the vehicle and the horizontal axis of the image. It is obtained by detecting the edge straight line of the frame longitudinal beam through Hough transform and calculating its slope. The vertical scaling ratio is the ratio of the actual vehicle height to the vehicle pixel height in the image, which needs to be calibrated in combination with known calibration objects (such as reflective strips of standard size) or laser ranging data. For example, when the equation of the edge straight line of the vehicle longitudinal beam is detected as y=0.82x+120, the horizontal tilt angle is arctan(0.82)=39.3°; if the actual measured vehicle height is 3.5 meters and the height in the image is 700 pixels, the vertical scaling ratio is 3.5 / 700=0.005 meters / pixel. For scenes without calibration objects, the vanishing point estimation method is used: the horizontal tilt angle is determined by detecting the intersection of the two parallel lane lines in the image, and the scaling ratio is inferred based on the prior knowledge of the vehicle width (such as the standard truck width of 2.55 meters).
[0087] Step S34: performing affine transformation correction on the vehicle side training image according to the horizontal tilt angle, and adjusting the image ratio of the vehicle side training image according to the vertical scaling ratio to generate the side training sample set.
[0088] For example, affine transformation correction is achieved by combining a rotation matrix and a translation matrix, that is, taking the center of the image as the origin, the image is rotated around the origin by the opposite number of the horizontal tilt angle (such as -39.3°) to align the longitudinal axis of the vehicle with the horizontal axis of the image; the vertical scaling adjustment uses a bilinear interpolation algorithm to scale the image vertically to the actual scale. For example, 700 pixels in the vertical direction of the original image correspond to 3.5 meters, and the scaling ratio is 0.005 meters / pixel. If it needs to be converted to a standard ratio of 0.004 meters / pixel (corresponding to 875 pixels in height), the vertical stretching factor is 875 / 700=1.25. The corrected image needs to be edge-filled, and the blank area generated by the rotation is expanded by mirror copying to ensure that the vehicle body is intact and not cut. The final generated side training sample set contains the geometrically corrected image and the corresponding distortion parameter metadata, which can be used for spatial feature alignment and scale normalization in subsequent model training.
[0089] As an implementation mode, the step S50, inputting the side training sample set and the top training sample set into the initial convolutional neural network for multi-stage training, adjusting the network weight parameters until the error between the output feature of the initial convolutional neural network and the labeled vehicle model feature data is lower than a preset threshold, and generating the vehicle model matching model, includes the following steps: Step S51: In the first stage of training, the side training sample set is input into the first feature extraction layer of the initial convolutional neural network to obtain a side primary feature vector.
[0090] The first feature extraction layer can use the ResNet-50 architecture as the backbone network, and its input receives a three-channel side training image with a resolution of 512×512 pixels. After the initial convolution layer (7×7 kernel, stride 2) and the maximum pooling layer (3×3 kernel, stride 2), it outputs a 256×256×64 feature map, and then gradually extracts high-level semantic features through four residual blocks (each block contains 3 bottleneck structures), and finally generates a side primary feature vector with a dimension of 2048 in the global average pooling layer. For example, for a side training image of a van, in the feature map output by the first residual block, the 32nd channel shows significant activation at the hinge position of the container door (response value 0.92), and the 128th channel is sensitive to the high-frequency details of the tire contact pattern; after the fourth residual block, the 512th dimension in the feature vector corresponds to the continuity metric of the container side panel welding seam, and the 1024th dimension encodes the relative height difference between the door handle and the body plane. At this stage, the network weight parameters are initialized in a transfer learning manner, the model parameters pre-trained on the ImageNet dataset are loaded, and the weights of the first three residual blocks are frozen to accelerate convergence.
[0091] Step S52: In the second stage of training, the top training sample set is input into the second feature extraction layer of the initial convolutional neural network to obtain the top primary feature vector.
[0092] The second feature extraction layer can be designed with the DenseNet-121 architecture, with a single-channel top training image of 256×256 pixels (after grayscale conversion) as input, and multi-scale features are aggregated step by step through the initial convolution layer (7×7 kernel, step size 2) and dense block sequence (each block contains 6 dense connection layers), and finally the feature map size is compressed to 8×8×1024 through the transition layer (1×1 convolution and 2×2 average pooling), and then the top primary feature vector with a dimension of 1024 is output through global maximum pooling. For example, when processing the top image of the refrigerated truck, in the feature map output by the first dense block, the 56th channel is sensitive to the spacing of the transverse reinforcement ribs of the container roof, and the 89th channel detects the rectangular outline of the refrigeration unit cover; after the transition layer, the 256th dimension in the feature vector reflects the material difference between the solar panel and the roof steel plate, and the 768th dimension encodes the radius change rate of the curved surface of the deflector. In this stage, all parameters of the second feature extraction layer are unfrozen, and an adaptive learning rate strategy (initial value 1e-4, decaying to 0.5 times every 10 epochs) is adopted to focus on optimizing the fine-grained feature extraction capability of the roof structure.
[0093] Step S53: In the third training stage, the side primary feature vector is cross-attendedly fused with the top primary feature vector to generate a fused feature vector.
[0094] The cross-attention fusion module is built on the basis of the multi-head attention mechanism. Its core is to realize cross-modal information complementation by calculating the dynamic weight distribution between the side and top features. Specifically, the side primary feature vector is used as the query vector (Query), and the top primary feature vector is used as the key-value vector (Key-Value). They are mapped to the 64-dimensional subspace through linear transformation, and then the dot product similarity matrix between the query and the key is calculated, and the attention weight distribution is generated by Softmax normalization. For example, when the dot product score of the 512th dimension (container welding seam) in the side feature and the 256th dimension (transverse reinforcement rib) in the top feature reaches 0.87, it indicates that there is a strong spatial correlation between the two, and this weight will enhance the fusion contribution of the two features; on the contrary, if the similarity between the tire feature (the 1024th dimension of the side) and the roof fairing feature (the 768th dimension of the top) is only 0.12, then their fusion ratio is reduced to avoid noise interference. The number of attention heads is set to 8, and each head independently calculates the weight distribution of the 64-dimensional subspace, and finally splices the multi-head output into a 512-dimensional fusion feature vector. During this process, the interactive learning of side and roof features enables the model to capture three-dimensional geometric constraint relationships such as container height and roof curvature, thereby improving the feature representation capabilities.
[0095] Step S54: In the fourth training stage, the fused feature vector is input into the fully connected classification layer, and the loss function value between the predicted vehicle model feature data and the labeled vehicle model feature data is calculated.
[0096] For example, the fully connected classification layer consists of three parallel subnetworks: wheelbase regression subnetwork (2 layers of 512-dimensional fully connected + linear output), car body classification subnetwork (3 layers of 1024-dimensional fully connected + Softmax output) and tire distribution matching subnetwork (graph convolution layer + contrast loss calculation). The loss function adopts a multi-task weighted sum form: the wheelbase part uses smooth L1 loss (Smooth L1 Loss) to measure the deviation between the predicted wheelbase sequence and the labeled value, with a weight of 0.5; the car body part uses cross-entropy loss (Cross-Entropy Loss) to optimize the classification accuracy, with a weight of 0.3; the tire distribution part is based on triplet loss (Triplet Loss) to optimize the similarity of the tire position embedding vector, with a weight of 0.2. For example, when the predicted wheelbase is [3.21, 7.83, 11.41] meters and the labeled value is [3.19, 7.85, 11.43] meters, the smooth L1 loss is calculated to be 0.015; the probability of the predicted category of the carriage being "refrigerated container" is 0.92, and the label is a one-hot vector [0,1,0,...], and the cross entropy loss is 0.083; the tire distribution triplet loss is calculated by the embedding distance difference between the positive sample (same model tire position) and the negative sample (different model tire position). If the positive sample distance is 0.2, the negative sample distance is 1.3, and the interval α=0.5, then the loss value is max(0.2-1.3+0.5, 0)=0. The total loss function value is 0.5×0.015+0.3×0.083+0.2×0=0.034.
[0097] Step S55: adjusting the weight parameters of the first feature extraction layer, the second feature extraction layer and the fully connected classification layer according to the back propagation of the loss function value until the loss function value converges to the preset threshold.
[0098] Back propagation can use the Adam optimizer, set the initial learning rate to 3e-5, β1=0.9, β2=0.999, and the weight decay coefficient to 1e-4. During the training process, 32 groups of samples (16 groups of side-top image pairs) are input in each batch, and the gradients of each layer are calculated by the chain rule, and the parameters are updated according to the learning rate. For example, when training the 150th epoch, the gradient of the weight matrix W1 of the fully connected classification layer is ∂L / ∂W1=0.0023, and the update amount estimated based on Adam's momentum is , parameter update step size The training was terminated when the validation set loss decreased by less than 0.1% for 10 consecutive epochs or when the total training epochs reached 300. The wheelbase mean absolute error of the final model on the test set dropped to 8 mm, the cabin classification accuracy was 97.2%, and the tire distribution matching F1 score was 0.91, meeting the preset thresholds (wheelbase MAE ≤ 10 mm, classification accuracy ≥ 95%, F1 ≥ 0.9), generating a deployable vehicle model matching model.
[0099] As an implementation manner, in step S53, cross-attention fusion is performed on the side primary feature vector and the top primary feature vector to generate a fused feature vector, including the following steps: Step S531: Calculate the correlation matrix between the side primary feature vector and the top primary feature vector, wherein the correlation matrix reflects the correlation strength between different feature dimensions.
[0100] The relevance matrix can be calculated, for example, by a bilinear attention mechanism, given a side feature vector Vs∈R 2048 and the top eigenvector Vt∈R 1024 , first project Vs to the query space Q = Wq · Vs (Wq ∈ R 512×2048 ), Vt is projected to the key space K = Wk · Vt (Wk∈R 512×1024 ), and then calculate the correlation matrix C = Q·K T ∈R 512×512 , where each element C_ij represents the strength of association between the i-th feature dimension of the side and the j-th feature dimension of the top. For example, when the C value of the 312th dimension (encoding the door handle height) in the side feature and the 198th dimension (encoding the sunroof position) in the top feature is 0.94, it indicates that there is a strong spatial correspondence between the two; while the C value of the 1024th dimension (tire contact mark) in the side and the 56th dimension (container roof material) in the top is only 0.12, reflecting a low correlation. After the matrix is calculated, the row-wise Softmax normalization is applied so that the sum of the weights of each row is 1.
[0101] Step S532: Generate side feature weight distribution and top feature weight distribution according to the correlation matrix.
[0102] The weight distribution of the side features is obtained by summing and normalizing the correlation matrix column direction, that is, Ws=Softmax(sum(C, axis=1))∈R 512 , reflecting the importance of each side feature dimension in the global context; the top feature weight distribution is obtained by summing and normalizing the row direction Wt=Softmax(sum(C, axis=0))∈R 512, characterizing the contribution of the top feature dimension. For example, if the column sum of the 512th dimension of the side (the weld seam of the container) in the correlation matrix reaches 38.7 (the maximum value), then Ws
[512] =0.15; the row sum of the 256th dimension of the top (the transverse reinforcement rib) is 29.3, and Wt
[256] =0.12. This process strengthens the feature fusion of the key structure of the vehicle by focusing on the high response area.
[0103] Step S533: performing weighted summation on the side primary feature vectors according to the side feature weight distribution to generate a side weighted feature vector.
[0104] Side weighted eigenvector , where Ws i is the expanded weight mapped to the original 2048-dimensional space (512-dimensional Ws is expanded to 2048 dimensions through linear interpolation). For example, the door handle-related dimensions in the side features (dimensions 312-320) have a higher weight (0.08-0.12) of Ws, and their eigenvalues are increased from the original 0.75 to 0.89 after weighting; while the eigenvalues of low-weight dimensions (such as background noise in dimensions 1500-1600) are reduced from 0.32 to 0.05. The weighted Vsw∈R 2048 Preserve important features while suppressing redundant information.
[0105] Step S534: performing weighted summation on the top primary feature vectors according to the top feature weight distribution to generate a top weighted feature vector.
[0106] Top weighted eigenvector , where Wt j The dimension is expanded from 512 to 1024 through interpolation. For example, the dimension related to the curvature of the air deflector in the top feature (dimensions 768-780) is increased from 0.68 to 0.82 after weighting due to the Wt weight of 0.10-0.15; the edge feature of the container top cover (dimensions 100-120) has a lower weight (0.03-0.05), and the eigenvalue is reduced from 0.91 to 0.45. Vtw∈R 1024 Highlight the key roof structures and de-emphasize secondary areas.
[0107] Step S535: concatenate the side weighted feature vector and the top weighted feature vector to generate the fused feature vector.
[0108] The concatenation operation is performed along the feature dimension, and the fused feature vector Vf=Concat(Vsw, Vtw)∈R 3072For example, the value of the 2048th dimension (door handle feature) in Vsw is 0.89, and the value of the 1024th dimension (air deflector feature) in Vtw is 0.82. After concatenation, Vf
[2048] =0.89 and Vf
[3072] =0.82. To reduce the dimensionality, the fully connected layer (3072→512 dimensions) is used for compression, and LayerNorm normalization and ReLU activation are applied to finally generate a fused feature vector suitable for multi-task learning.
[0109] In an optional embodiment, in step S500, after determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification mark, the method may also include: detecting in real time whether the comprehensive matching degree is lower than a preset matching threshold, and when it is lower, extracting the unmatched vehicle wheelbase features, vehicle body structure features and tire distribution features from the vehicle type feature matching results; generating a temporary wheelbase code based on the unmatched vehicle wheelbase features, and constructing a new vehicle type feature data packet in combination with the vehicle body structure features and tire distribution features; matching the new vehicle type feature data packet with each vehicle type template data in the vehicle type feature library through a manual verification interface to obtain verification feature data that has been manually reviewed and approved; adding the verification feature data to the vehicle type feature library and updating the template wheelbase range, template vehicle body structure template and template tire distribution template; and dynamically correcting the vehicle type identification mark of subsequent passing vehicles according to the updated vehicle type feature library.
[0110] This embodiment is suitable for learning and updating the database of location vehicle model features. When the comprehensive matching degree is detected in real time to be lower than the preset matching threshold (such as 0.85), the system automatically triggers the unmatched feature extraction process. The unmatched vehicle wheelbase feature refers to the numerical segment in the predicted wheelbase sequence that has no intersection with the template wheelbase range of all vehicle model template data. For example, the actual measured value of a vehicle wheelbase is [3.25, 7.91, 11.62] meters, and the most similar template wheelbase range in the vehicle model library is [3.18-3.22, 7.75-7.85, 11.3-11.5] meters, then the third wheelbase of 11.62 meters is marked as an unmatched feature segment. The temporary wheelbase code is generated using a segmented hash algorithm: the unmatched wheelbase value is discretized at intervals of 10 cm (such as 11.62 meters mapped to 116.2 discrete units), and the first 8 bits are taken as the coding identifier (such as "3D5F8A2B") after performing MD5 hash operation on each discrete unit. The new vehicle model feature data package consists of a temporary wheelbase code, a vehicle body structure feature vector (512 dimensions) and a tire distribution topology map (including tire position coordinates and adjacency relationships). The data package format is a JSON-LD structured document with a timestamp and spatial location tag. The manual verification interface is deployed on the toll system management platform. Auditors can compare the new vehicle model feature data package with historical abnormal records. If it is confirmed to be a new type of vehicle (such as the first passage of a certain model of electric heavy truck), check the verification pass option and supplement the vehicle model parameters (such as load level, power type). The updated vehicle model feature library appends the new vehicle model template data to the storage node through incremental writing. The template wheelbase range is expanded to [3.18-3.22, 7.75-7.85, 11.3-11.6] meters. The template vehicle body structure template adds a battery compartment layout feature vector unique to electric trucks. The template tire distribution template adds tread coding rules for wide-body single tires. When subsequent passing vehicles are compared for similarity, the updated vehicle model feature library is dynamically loaded. For example, when an electric truck of the same model passes again, its vehicle model identification mark is updated from "UNKNOWN" to "ET-38D", and the rate calculation module is triggered to call the newly added electricity price rate of 0.52 yuan / ton·kilometer.
[0111] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification mark in step S500, the method may also include: generating an error distribution heat map based on the wheelbase matching degree, the car body matching degree and the tire matching degree, the error distribution heat map reflecting the identification deviation position of each local feature area; extracting the coordinates of the abnormal area in the error distribution heat map where the deviation value exceeds the tolerance threshold, and mapping them to the corresponding position of the vehicle multi-dimensional fusion feature map; adjusting the convolution kernel size and step size parameters of the multi-layer convolution network in the vehicle type matching model according to the coordinates of the abnormal area; using the adjusted convolution kernel parameters to incrementally train the vehicle type matching model, and updating the feature response map activation weights of the multi-layer convolution network; and applying the updated vehicle type matching model to the calculation of vehicle type feature matching results for the next batch of passing vehicles.
[0112] This embodiment is a process of dynamic optimization and incremental learning of model parameters. The error distribution heat map characterizes the recognition deviation intensity of each local feature area in the form of a three-dimensional tensor, where the heat value is calculated by the weighted residual of the wheelbase matching, the car body matching and the tire matching. For example, in the front axle area of the container (coordinates x: 120-180, y: 300-400 pixels), the wheelbase matching residual is detected to be 0.23 (threshold 0.15), the car body structure matching residual is 0.18 (threshold 0.1), and the tire matching residual is 0.32 (threshold 0.2). The heat value of this area is marked in red (RGB: 255, 0, 0). The coordinates of the abnormal area are extracted through connected domain analysis. When the heat value of a continuous 10×10 pixel area exceeds the tolerance threshold, its minimum circumscribed rectangle coordinates are recorded (such as the upper left corner point (125, 305), the lower right corner point (175, 395)). When mapping to the multi-dimensional fusion feature map of the vehicle, the spatial projection conversion algorithm is used to convert the two-dimensional coordinates into the channel index of the feature map. For example, the coordinate (150,350) corresponds to the 35th activation area of the 24th channel of the feature map. The convolution kernel size adjustment strategy is dynamically set according to the area of the abnormal area: for areas with an area less than 200 pixels² (such as abnormal tire contact points), the kernel size of the corresponding convolution layer is expanded from 3×3 to 5×5 to enhance the local feature capture capability; for areas with an area greater than 500 pixels² (such as large mismatches on the side of the container), the step size parameter is adjusted from 2 to 1 to reduce downsampling losses. Incremental training uses an online learning framework, extracting 5% of the samples from the real-time data stream as the training set, inputting 16 groups of samples per batch, setting the training cycle to 10 epochs, and the learning rate decays to 0.1 times the initial value. The updated multi-layer convolutional network has undergone significant changes in the distribution of activation weights in the feature response map: for example, the average activation value of the original model in the front axle area of the container was 0.75, which was increased to 0.88 after adjustment; the standard deviation of the feature response intensity in the tire contact point area was reduced from 0.15 to 0.09.
[0113] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification mark in step S500, the method may further include: determining whether the vehicle type identification mark includes a preset unknown vehicle type mark, and if so, initiating an image recapture instruction; controlling the shooting device to adjust the focal length and angle according to the image recapture instruction, and reacquiring a high-resolution side profile image and a top profile image of the passing vehicle; performing multi-frame deblurring processing on the reacquired image to generate a recaptured image set with enhanced clarity; inputting the recaptured image set into the vehicle type matching model for secondary feature matching, and updating the vehicle wheelbase features and the vehicle compartment structure features in the vehicle type feature matching results; and re-executing the similarity comparison according to the updated vehicle type feature matching results to cover the original vehicle type identification mark.
[0114] This embodiment is about the re-identification and result coverage of unknown vehicle models. The preset unknown vehicle model identification adopts the encoding rule of "UNKNOWN-" prefix plus 6 random characters (such as "UNKNOWN-3A5F9D"). When the vehicle model identification contains this prefix, the re-acquisition instruction system activates the multi-modal trigger mechanism: first, the best shooting angle is calculated according to the real-time position of the vehicle (positioned by the coordinates of the ground sensing coil), and the pitch angle of the top camera is adjusted from 30° to 45° to cover the full height of the container. At the same time, the focal length of the side view camera is switched from 50mm to 200mm telephoto mode to capture the details of the door rivets. The resolution of the re-acquired high-resolution image is 8192×4320 pixels, and the frame rate is increased to 60fps to support multi-frame deblurring processing. Multi-frame deblurring uses a combination of non-uniform motion estimation and iterative deconvolution algorithms: For lateral blur caused by vehicle vibration, the displacement vector of each frame is estimated by the optical flow method (such as inter-frame displacement Δx=2.3 pixels, Δy=0.8 pixels), and the point spread function matrix is constructed and restored by Wiener filtering. In the enhanced re-collected image set, the MTF (modulation transfer function) value of the tire tread pattern is increased from 0.35 to 0.62, and the edge sharpness of the door handle is increased by 40%. During secondary feature matching, the vehicle model matching model uses high-precision mode: the output dimension of the wheelbase regression subnetwork is expanded from 3 to 5, and the support axle detection for multi-axle vehicles is added; the number of categories of the vehicle body structure classifier is increased from 120 to 200 to accommodate new special vehicles. In the updated vehicle feature matching results, the vehicle wheelbase feature of an electric truck was corrected from [3.25, 7.91, 11.62] meters to [3.24, 7.89, 11.60] meters, and the cosine similarity between the vehicle structure feature vector and the template was increased from 0.72 to 0.91. The similarity comparison module implements a mandatory coverage protocol: when the comprehensive matching degree of the secondary match exceeds the original result and the difference is greater than 0.1, the original identifier "UNKNOWN-3A5F9D" is replaced with "ET-38D", and the version number (such as "Ver2.1") and the correction timestamp are recorded in the transaction log.
[0115] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification mark in step S500, it also includes: counting the identification frequency of each toll vehicle type category within a preset time period to generate a vehicle type frequency distribution table; sorting the vehicle type template data in the vehicle type feature library according to the vehicle type frequency distribution table by storage priority; migrating the template wheelbase range, template car body structure template and template tire distribution template of the high-frequency vehicle type to a cache area; establishing a fast retrieval queue according to the cache area, and giving priority to comparing the vehicle type feature matching results of subsequent passing vehicles with the vehicle type template data in the fast retrieval queue; when the comprehensive matching degree does not meet the standard in the fast retrieval queue, switching to the full data of the vehicle type feature library for secondary retrieval.
[0116] This embodiment realizes dynamic optimization and retrieval of the vehicle model feature library, counts the recognition frequency of each toll vehicle model category within the preset time period, and generates a vehicle model frequency distribution table. The preset time period is set as a dynamically adjusted time window. For example, the morning peak period (07:00-09:00) and the night period (22:00-06:00) use different statistical granularities (the former is sliced by 15 minutes, and the latter is sliced by 1 hour). The vehicle model frequency distribution table is stored in a hash map structure, with the key being the vehicle model number (such as "HT-45C") and the value being a composite object containing the number of recognitions, timestamp sequence, and average matching degree. For example, in the statistics during the morning peak period, the frequency of the logistics truck "LT-32B" is 8.7 vehicles per minute, with an average matching degree of 0.93, while the frequency of the passenger car "CT-18D" is 2.3 vehicles per minute, with an average matching degree of 0.87. When the model template data in the model feature library is sorted by storage priority according to the model frequency distribution table, a heat weighted algorithm is used: priority score = frequency × 0.6 + matching degree average × 0.4, for example, "LT-32B" score = 8.7 × 0.6 + 0.93 × 0.4 = 5.93, ranking first priority. When the template wheelbase range, template compartment structure template and template tire distribution template of high-frequency models are migrated to the cache area, the cache adopts the LRU (Least Recently Used) replacement strategy, and the capacity is set to 20% of the total number of models (such as the first 200 out of 1000). The migration process is realized through memory mapping technology, and the template data in the disk is loaded into the DDR4-3200 memory module, and the access delay is reduced from 15ms to 0.2ms. When establishing a fast retrieval queue according to the cache area, the queue index structure adopts B+ tree organization, the key is the model number, and the value points to the address pointer of the template data in the memory. For example, the wheelbase range of the template "LT-32B" is [3.18, 3.22, 7.75, 7.85, 11.3, 11.5] meters, the car structure template (512-dimensional vector) and the tire distribution template (12-node topology map) are loaded into the memory address 0x7FFA2B1C0000-0x7FFA2B1D8000 area. When the vehicle model feature matching results of subsequent passing vehicles are compared with the vehicle model template data in the fast retrieval queue, the SIMD instruction set is used to accelerate the vector similarity calculation. For example, the 512-dimensional car structure vector is processed in parallel using AVX-512 instructions, and the single comparison time is shortened from 1.2ms to 0.15ms. When the comprehensive matching degree does not meet the standard in the quick search queue (for example, the highest matching degree 0.82 is lower than the threshold value 0.85), the system switches to the full data of the vehicle feature library for secondary search. The full search enables a multi-threaded shard query mechanism, dividing the vehicle feature library into 8 shards. Each shard is scanned in parallel in the NVMe SSD storage pool through an independent thread, and the total query time is optimized from 220ms for full scan to 35ms.For example, when a new electric truck passes for the first time, there is no matching item in the quick search queue. After the system switches to full search, it matches the newly added "ET-38D" vehicle model template in shard 3 with an overall matching degree of 0.89, triggering the cache update mechanism to add it to the cache area.
[0117] Alternatively, in an optional embodiment, after determining the toll vehicle type category corresponding to the passing vehicle and generating a vehicle type identification mark in step S500, it also includes: extracting the maximum interval point coordinates of the vehicle wheelbase feature and the edge connection point sequence of the vehicle body structure feature from the vehicle type feature matching results; generating a three-dimensional vehicle body contour report based on the maximum interval point coordinates and the edge connection point sequence, the report including the wheelbase ratio parameters and the vehicle body structure topology map; binding the three-dimensional vehicle body contour report with the vehicle type identification mark, and adding a timestamp and lane position information; losslessly compressing and encoding the bound data packets to generate a standardized identification report; and sending the standardized identification report to the toll terminal and the cloud audit platform for parallel verification.
[0118] This embodiment involves three-dimensional vehicle model data encapsulation and multi-terminal verification, and extracts the maximum interval point coordinates of the vehicle wheelbase feature and the edge connection point sequence of the vehicle body structure feature from the vehicle model feature matching results. The maximum interval point coordinates refer to the two axis projection points with the largest adjacent axle spacing in the vehicle wheelbase sequence. For example, the wheelbase sequence of a five-axle truck is [3.2, 7.8, 11.4, 15.0, 18.6] meters, and the maximum interval is 3.6 meters from the fourth axis to the fifth axis, corresponding to the point coordinates (x1=1520, y1=450) and (x2=1840, y2=450) in the image coordinate system. The edge connection point sequence extracts the contour polygon vertices of the vehicle body structure feature through the Alpha Shape algorithm. For example, the edge point sequence of the side panel of the refrigerated truck container contains 56 vertices, which are stored in a clockwise direction as [(x1, y1), (x2, y2), ..., (x56, y56)]. When generating a 3D vehicle model profile report based on the coordinates of the maximum interval point and the sequence of edge connection points, a surface reconstruction technology based on point cloud is used: the maximum interval point is used as the longitudinal reference line, and the sequence of edge connection points is stretched in the vertical direction to generate a 3D mesh model, and the wheelbase ratio parameters are calculated (such as the distance from the fourth to the fifth axle accounts for 19.3% of the total vehicle length). The topology diagram of the vehicle structure is constructed using the Delaunay triangulation algorithm, and the sequence of edge connection points is converted into a non-uniform rational B-spline (NURBS) surface. For example, the corrugated structure of the container side panel is modeled as a periodic surface with an amplitude of 5mm and a wavelength of 120mm. When binding the 3D vehicle model profile report with the vehicle model identification mark and adding timestamp and lane position information, the binding data packet adopts the Protobuf binary encoding format, and the timestamp accuracy is to the millisecond level (such as "2023-08-20T14:23:05.235Z"). The lane position information is located by combining the RFID landmark coordinates of the toll station (such as the WGS-84 coordinates of lane 3 118.3245°E, 32.4567°N) with the relative position of the vehicle (12.5 meters from the entrance landmark). When the bound data packet is losslessly compressed and encoded, the DEFLATE algorithm combined with Huffman coding is used, and the compression rate is set to the highest level (Level 9). For example, the original data packet size of 2.3MB is reduced to 480KB after compression. When generating a standardized identification report, the report header contains the version number (such as "V1.2"), checksum (CRC-32) and data packet length field, and the body is embedded with the compressed 3D vehicle model data.When the standardized identification report is sent to the toll terminal and the cloud audit platform for parallel verification, the local verification of the toll terminal verifies the data integrity through SHA-256, such as calculating the hash value of the compressed package "a1b2c3d4e5f6..." and comparing it with the checksum in the report header; the cloud audit platform performs redundant verification through distributed verification nodes, such as three nodes respectively verifying the timestamp continuity, vehicle type identification compliance and coordinate validity. After all pass, a confirmation signal is returned to the toll system to trigger the rate calculation and release instructions. If a verification fails (such as the vehicle type identification "ET-38D" is not registered in the cloud), the system automatically isolates the abnormal data packet and starts the manual review process, and records the abnormal event in the audit log.
[0119] Figure 2 A hardware entity diagram of a computer system provided by an embodiment of the present invention is as follows Figure 2 As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.
[0120] The memory 1002 stores computer programs that can be run on the processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001. It can also cache data to be processed or processed by the processor 1001 and various modules in the computer system 1000 (for example, image data, audio data, voice communication data, and video communication data). This can be achieved through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0121] When the processor 1001 executes the program, the steps of any one of the above-mentioned methods for identifying toll vehicle types based on highway traffic images are implemented. The processor 1001 generally controls the overall operation of the computer system 1000.
[0122] An embodiment of the present invention provides a computer storage medium storing one or more programs, which can be executed by one or more processors to implement the steps of the toll vehicle type recognition method based on highway traffic images of any of the above embodiments.
[0123] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present invention, please refer to the description of the method embodiments of the present invention for understanding. The above processor can be at least one of a target application integrated circuit (Application Specific Integrated Circuit, ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (Programmable Logic Device, PLD), a field programmable gate array (Field Programmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device that realizes the function of the above processor can also be other, and the embodiment of the present invention is not specifically limited.
[0124] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and the like; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0125] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the size of the sequence number of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The above-mentioned sequence number of the embodiment of the present invention is only for description and does not represent the advantages and disadvantages of the embodiment. It should be noted that in this article, the term "includes", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or device including the element.
[0126] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0127] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0128] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0129] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0130] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0131] The above description is only an implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for identifying toll vehicle types based on highway traffic images, characterized in that: The method comprises: Acquire an image set of passing vehicles, wherein the image set includes multiple frames of vehicle side profile images and vehicle top profile images taken continuously; Performing multi-scale brightness compensation processing on the side profile image of the vehicle to generate a side feature enhanced image, and performing contour sharpening processing on the top profile image of the vehicle to generate a top feature enhanced image; Performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a multi-dimensional fusion feature map of the vehicle; Inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model, obtaining a vehicle model feature matching result output by the vehicle model matching model, wherein the vehicle model feature matching result includes a vehicle wheelbase feature, a vehicle compartment structure feature, and a tire distribution feature; A similarity comparison is performed based on the vehicle model feature matching result and a pre-stored vehicle model feature library to determine the toll vehicle model category corresponding to the passing vehicle and generate a vehicle model identification mark.
2. The method according to claim 1, characterized in that The performing multi-scale brightness compensation processing on the vehicle side profile image to generate a side feature enhanced image includes: Extracting a continuous brightness distribution curve in the vehicle side profile image, wherein the continuous brightness distribution curve includes a first brightness distribution in a bottom area of the vehicle and a second brightness distribution in a top area of the vehicle; generating a dynamic brightness compensation coefficient according to a gradient difference between the first brightness distribution and the second brightness distribution; Locally enhancing the low brightness area in the vehicle side profile image according to the dynamic brightness compensation coefficient to generate an initial compensated image; Performing multi-scale filtering on the initial compensated image to extract edge texture features at different resolutions; The filtering parameters are adjusted according to the density distribution of the edge texture features to generate the side feature enhanced image.
3. The method according to claim 2, characterized in that The step of performing dual-channel feature fusion on the side feature enhanced image and the top feature enhanced image to generate a multi-dimensional fusion feature map of the vehicle includes: Performing spatial coordinate transformation on the side feature enhanced image to obtain a set of key contour points in the side image coordinate system; Performing spatial coordinate transformation on the top feature enhanced image to obtain a set of key contour points in the top image coordinate system; Performing three-dimensional spatial mapping of a set of key contour points in the side image coordinate system and a set of key contour points in the top image coordinate system to generate a three-dimensional contour model of the vehicle; Extracting the curvature features of the vehicle surface according to the curvature distribution of each contour point in the three-dimensional contour model of the vehicle; The vehicle surface curvature features and the edge texture features of the side feature enhanced image are superimposed and fused to generate the vehicle multi-dimensional fusion feature map.
4. The method according to claim 1, characterized in that: The step of inputting the vehicle multi-dimensional fusion feature map into a pre-trained vehicle model matching model to obtain a vehicle model feature matching result output by the vehicle model matching model includes: Dividing the multi-dimensional fusion feature map of the vehicle into a plurality of local feature regions, each local feature region corresponding to a feature type of a vehicle wheelbase feature, a vehicle compartment structure feature or a tire distribution feature; Perform feature vectorization processing on each local feature area to generate a local feature vector set; Inputting the local feature vector set into the multi-layer convolutional network in the vehicle model matching model to obtain a feature response map output by each convolutional layer; Determining, according to the activation area distribution of the characteristic response diagram, a wheelbase matching degree corresponding to the vehicle wheelbase characteristic, a vehicle compartment matching degree corresponding to the vehicle compartment structure characteristic, and a tire matching degree corresponding to the tire distribution characteristic; The wheelbase matching degree, the vehicle compartment matching degree and the tire matching degree are weightedly integrated to generate the vehicle model feature matching result.
5. The method according to claim 4, characterized in that The method of performing a similarity comparison between the vehicle type feature matching result and a pre-stored vehicle type feature library to determine the toll vehicle type category corresponding to the passing vehicle and generate a vehicle type identification mark includes: Extracting each vehicle model template data from the vehicle model feature library, wherein the vehicle model template data includes a template wheelbase range, a template vehicle compartment structure template, and a template tire distribution template; Calculating a first similarity between the wheelbase matching degree and the template wheelbase range, calculating a second similarity between the vehicle compartment matching degree and the template vehicle compartment structure template, and calculating a third similarity between the tire matching degree and the template tire distribution template; Determining the comprehensive matching degree between the passing vehicle and the template data of each vehicle model according to the weighted sum of the first similarity, the second similarity and the third similarity; The vehicle model number corresponding to the vehicle model template data with the highest comprehensive matching degree is used as the toll vehicle model category, and a vehicle model identification mark including the vehicle model number is generated.
6. The method according to claim 1, characterized in that The method also includes a pre-training step of the vehicle type matching model: Obtain a historical vehicle training set, wherein the historical vehicle training set includes multiple groups of vehicle image samples and corresponding annotated vehicle model feature data; Perform image segmentation processing on each group of vehicle image samples to extract vehicle side training images and vehicle top training images; Performing noise suppression and distortion correction on the vehicle side training image to generate a side training sample set; Performing shadow removal and contrast enhancement on the vehicle top training image to generate a top training sample set; The side training sample set and the top training sample set are input into the initial convolutional neural network for multi-stage training, and the network weight parameters are adjusted until the error between the output features of the initial convolutional neural network and the labeled vehicle model feature data is lower than a preset threshold, thereby generating the vehicle model matching model.
7. The method according to claim 6, characterized in that The image segmentation process is performed on each group of vehicle image samples to extract the vehicle side training image and the vehicle top training image, including: Performing background separation processing on the vehicle image sample to obtain a vehicle main area mask; Dividing the vehicle side area and the vehicle top area according to the geometric center position of the vehicle main area mask; Performing edge tracking on the side area of the vehicle to extract a continuous and closed side contour line; Determining a clipping boundary of the vehicle side training image according to a curvature change point of the side contour line; Segmenting the vehicle image sample according to the cropping boundary to generate the vehicle side training image and the vehicle top training image; The step of performing noise suppression and distortion correction on the vehicle side training image to generate a side training sample set includes: Detecting a high-frequency noise area in the vehicle side training image and generating a noise distribution map; Adjusting the window size of the adaptive filter according to the density value of the noise distribution map to smooth the high-frequency noise area; Extracting perspective distortion parameters from the vehicle side training image, wherein the perspective distortion parameters include a horizontal tilt angle and a vertical scaling ratio; Affine transformation correction is performed on the vehicle side training image according to the horizontal tilt angle, and the image ratio of the vehicle side training image is adjusted according to the vertical scaling ratio to generate the side training sample set.
8. The method according to claim 6, characterized in that The step of inputting the side training sample set and the top training sample set into the initial convolutional neural network for multi-stage training, adjusting the network weight parameters until the error between the output feature of the initial convolutional neural network and the labeled vehicle model feature data is lower than a preset threshold, and generating the vehicle model matching model includes: In the first stage of training, the side face training sample set is input into the first feature extraction layer of the initial convolutional neural network to obtain a side face primary feature vector; In the second training stage, the top training sample set is input into the second feature extraction layer of the initial convolutional neural network to obtain the top primary feature vector; In the third training stage, the side primary feature vector is cross-attendedly fused with the top primary feature vector to generate a fused feature vector; In the fourth stage of training, the fused feature vector is input into the fully connected classification layer, and the loss function value between the predicted vehicle model feature data and the labeled vehicle model feature data is calculated; The weight parameters of the first feature extraction layer, the second feature extraction layer and the fully connected classification layer are adjusted according to the back propagation of the loss function value until the loss function value converges to the preset threshold.
9. The method according to claim 8, characterized in that The cross-attention fusion of the side primary feature vector and the top primary feature vector to generate a fused feature vector includes: Calculating a correlation matrix between the side primary feature vector and the top primary feature vector, wherein the correlation matrix reflects the correlation strength between different feature dimensions; generating a side feature weight distribution and a top feature weight distribution according to the correlation matrix; Performing weighted summation on the side primary feature vectors according to the side feature weight distribution to generate a side weighted feature vector; Performing weighted summation on the top primary feature vectors according to the top feature weight distribution to generate a top weighted feature vector; The side weighted feature vector is concatenated with the top weighted feature vector to generate the fused feature vector.
10. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Vehicle re-identification method and device based on roadside perception, and electronic equipment
CN114170516A
Vehicle type identification method based on deep learning fusion model
CN116863412A
Semantic segmentation-based vehicle interior quality evaluation method and system
CN117830211A
Vehicle overload identification method and system, storage medium and terminal
CN118537663A
High-precision vehicle positioning
US20240221215A1
Cited By
Tire defect detection method
CN120182275A
Road toll vehicle type identification method and device based on multi-dimensional verification
CN120411899A
Highway toll vehicle type identification method and device based on multi-dimension verification
CN120411899B
Building three-dimensional model lightweight design method and system based on artificial intelligence
CN120429937A
Lightweight design method and system of building three-dimensional model based on artificial intelligence
CN120429937B