Vehicle type classification method and system based on deep learning

By acquiring and processing images containing complete vehicle outlines and combining deep learning technology for vehicle area extraction and feature extraction, the problems of background interference and artificial feature limitations in vehicle type classification are solved, achieving higher accuracy and reliability.

CN120656004AActive Publication Date: 2025-09-16GUIZHOU HUILIANTONG ELECTRONIC COMMERCE SERVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511152140.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-09-16
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

In existing technologies, the accuracy and reliability of vehicle type classification are limited by background interference information and artificially designed features in images, making it difficult to effectively improve.

Method used

By acquiring images containing complete vehicle outlines, performing vehicle area extraction and feature extraction processing, and using deep learning technology for classification, the system eliminates background interference, automatically learns the hierarchical features of vehicle area images, and realizes end-to-end feature learning and classification decisions.

Benefits of technology

It improves the pertinence and reliability of vehicle type classification, ensures that the classification results accurately indicate the vehicle type, and improves the accuracy and reliability of vehicle type identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656004A_ABST
    Figure CN120656004A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle type classification method and system based on deep learning, and the method comprises the steps: obtaining a vehicle image set containing a plurality of vehicle types, carrying out the vehicle region extraction of the vehicle image set, and obtaining a vehicle region image set, each vehicle area image in the vehicle area image set only comprises a vehicle main body part; feature extraction processing is carried out on the vehicle area image set to obtain a vehicle feature set, and each vehicle feature in the vehicle feature set corresponds to the feature representation of one vehicle area image; and performing classification processing on the vehicle feature set through a preset vehicle classification model to obtain a vehicle classification result, the vehicle classification result being used for indicating a vehicle type corresponding to each vehicle image. The pertinence and reliability of vehicle type classification are improved, and it is effectively guaranteed that the classification result can accurately indicate the vehicle type corresponding to each vehicle image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular to a vehicle type classification method and system based on deep learning. Background Art

[0002] With the development of intelligent transportation systems, vehicle type classification technology has been widely used in various traffic scenarios. Vehicle type classification technology refers to the identification and differentiation of vehicle types in specific scenarios to achieve management and service for different types of vehicles. At present, common vehicle type classification technologies usually obtain images containing vehicles, then extract the shape, texture and other features of the vehicle through manually designed feature extraction methods, and then use traditional classification algorithms to classify the extracted features to obtain vehicle type information. However, this method often has a lot of background interference information in the image, and the manually designed features are difficult to fully capture the essential differences between vehicle types. As a result, the accuracy and reliability of vehicle type classification need to be improved. How to more effectively improve the accuracy and reliability of vehicle type classification has become an urgent problem to be solved in the field of vehicle identification. Summary of the Invention

[0003] The present invention provides a vehicle type classification method and system based on deep learning.

[0004] In a first aspect, an embodiment of the present invention provides a vehicle type classification method based on deep learning, the method comprising: Acquire a vehicle image set containing multiple vehicle types, wherein each vehicle image in the vehicle image set is an image containing a complete vehicle outline acquired by an image acquisition device; Performing vehicle region extraction processing on the vehicle image set to obtain a vehicle region image set, wherein each vehicle region image in the vehicle region image set only includes a main body of the vehicle; Performing feature extraction processing on the vehicle region image set to obtain a vehicle feature set, wherein each vehicle feature in the vehicle feature set corresponds to a feature representation of a vehicle region image; The vehicle feature set is classified and processed using a preset vehicle classification model to obtain a vehicle classification result, which is used to indicate the vehicle type corresponding to each vehicle image.

[0005] In a second aspect, an embodiment of the present invention provides a computer system, including: a memory storing a computer program; A processor is used to load the computer program to implement the vehicle type classification method based on deep learning as described above.

[0006] The vehicle type classification method based on deep learning provided by the present invention obtains a vehicle image set containing multiple vehicle types, and limits the vehicle image to an image containing a complete vehicle outline. By clarifying the image acquisition scene and quality requirements, the input data is ensured to be targeted and effective, and the interference of blurred, incomplete or non-target scene images on subsequent processing is avoided. The vehicle image set is subjected to vehicle region extraction processing to obtain a vehicle region image set containing only the main part of the vehicle. By focusing on the main area of ​​the vehicle, the processing object is accurately limited from the original image to the core target, and the influence of irrelevant information such as background and environment is eliminated, providing a pure input for subsequent feature extraction. The vehicle region image set is subjected to feature extraction processing to obtain a vehicle feature set, and the hierarchical and abstract features of the vehicle region image are automatically learned by combining deep learning technology, avoiding the limitations of manually designed features in traditional technology, so that the extracted features can better reflect the essential differences between vehicle types. The vehicle feature set is classified and processed by a preset vehicle classification model to obtain a vehicle classification result, realizing end-to-end integration of feature learning and classification decision. The various steps are closely related and work together to improve the targetedness and reliability of vehicle type classification, and effectively ensure that the classification result can accurately indicate the vehicle type corresponding to each vehicle image. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0008] Figure 1 This is a flowchart of a vehicle type classification method based on deep learning provided by an embodiment of the present invention.

[0009] Figure 2 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0011] See also Figure 1 , Figure 1A flowchart of a vehicle type classification method based on deep learning provided in an embodiment of the present invention. The vehicle type classification method based on deep learning can be executed by a computer system. The vehicle type classification method based on deep learning may include the following steps: Step S100: Acquire a vehicle image set containing multiple vehicle types, where each vehicle image in the vehicle image set is an image containing a complete vehicle outline captured by an image capture device.

[0012] In an embodiment of the present invention, a vehicle image collection is a collection of multiple vehicle images, which cover a variety of different types of vehicles, such as cars, trucks, and buses. The image acquisition device can be a camera installed in a specific location, such as a surveillance camera in a preset scene such as a highway toll station or a parking lot entrance and exit. In order to obtain a vehicle image collection, high-definition cameras can be installed at the entrance and exit of a highway toll station. These cameras collect images according to preset time intervals or trigger conditions. The camera continuously captures passing vehicles. As long as a vehicle enters the effective shooting range of the camera and the outline of the vehicle can be fully presented in the shooting picture, the captured image is included in the vehicle image collection. For example, when a vehicle passes through the entrance of a highway toll station, the camera is triggered to capture the image, and the captured image containing the complete outline of the vehicle will be added to the vehicle image collection.

[0013] Step S200: performing vehicle region extraction processing on the vehicle image set to obtain a vehicle region image set, wherein each vehicle region image in the vehicle region image set only includes a main body of the vehicle.

[0014] Vehicle region extraction involves accurately identifying and extracting the region containing only the main vehicle body from each vehicle image in the vehicle image collection. These extracted region images are then combined into a vehicle region image collection. The main vehicle body is the primary structural part of the vehicle itself, excluding other irrelevant information such as the surrounding environment.

[0015] As an embodiment, step S200, performing vehicle region extraction processing on the vehicle image set to obtain a vehicle region image set, may specifically include the following steps S210 to S250: Step S210: performing image grayscale processing on each vehicle image in the vehicle image set, converting the color vehicle image into a grayscale vehicle image.

[0016] Image grayscale conversion converts a color image into a grayscale image. A color vehicle image has three color channels: red, green, and blue, and contains rich color information. A grayscale vehicle image, on the other hand, contains only one channel, and the value of each pixel represents the grayscale value at that point, reflecting the image's brightness information.

[0017] In practice, a weighted average method can be used to grayscale a color vehicle image. Specifically, for each pixel in the color vehicle image, the pixel values ​​of its red, green, and blue channels are weighted and summed according to the set weights to obtain the pixel's grayscale value. For example, the weights assigned are: 0.299 for the red channel, 0.587 for the green channel, and 0.114 for the blue channel. In this way, each pixel in the color vehicle image is converted to its corresponding grayscale value, thereby converting the entire color vehicle image into a grayscale image.

[0018] Step S220: performing noise removal processing on the grayscale vehicle image to reduce interfering pixels in the grayscale vehicle image to obtain a denoised grayscale vehicle image.

[0019] Denoising is used to eliminate interfering pixels in grayscale vehicle images due to various factors. These interfering pixels may be caused by factors such as noise from the image acquisition device and lighting variations, which can affect the subsequent accurate extraction of the vehicle area. A denoised grayscale vehicle image is one in which the number of interfering pixels has been significantly reduced after noise removal.

[0020] As an embodiment, step S220 performs noise removal on the grayscale vehicle image to reduce interfering pixels in the grayscale vehicle image to obtain a denoised grayscale vehicle image. Specifically, the following steps S221 to S225 may be included: Step S221: performing pixel value statistical processing on the grayscale vehicle image, and calculating the pixel value mean and pixel value variance of all pixels in the grayscale vehicle image.

[0021] The pixel value mean is the average of the pixel values ​​of all pixels in the grayscale vehicle image, reflecting the overall brightness level of the image; the pixel value variance measures the degree of dispersion of the pixel value relative to the mean, reflecting the fluctuation of the pixel value in the image.

[0022] To calculate the pixel mean, the pixel values ​​of all pixels in the grayscale vehicle image are added together and then divided by the total number of pixels. To calculate the pixel variance, the square of the difference between each pixel value and the pixel mean is calculated. These squared values ​​are then added together and divided by the total number of pixels. This gives the pixel variance. For example, by performing pixel value statistical processing on a grayscale vehicle image collected at a highway toll booth, the pixel mean and pixel variance can be determined. This statistical information is useful for determining the noise threshold.

[0023] Step S222: determining a noise judgment threshold based on the pixel value mean and the pixel value variance, where the noise judgment threshold is the pixel value mean plus or minus a preset multiple of the pixel value variance.

[0024] The noise threshold is used to determine whether a pixel in a grayscale vehicle image is a noise pixel. The preset multiple is a constant set in advance based on actual conditions and can be determined through experiments or experience.

[0025] When determining the noise threshold, the upper and lower limits of the noise threshold are determined by adding and subtracting a preset multiple of the pixel value variance from the pixel value mean calculated in step S221. For example, if the preset multiple is 2, the pixel value mean is 100, and the pixel value variance is 10, then the lower limit of the noise threshold is 100 - 2 × 10 = 80, and the upper limit is 100 + 2 × 10 = 120. In vehicle images at highway toll booths, the noise threshold determined in this manner can effectively identify pixels that deviate from the normal pixel value range as noise pixels.

[0026] Step S223: traverse each pixel in the grayscale vehicle image and mark the pixel whose pixel value exceeds the noise judgment threshold range as a noise pixel.

[0027] The traversal is to sequentially visit each pixel in the grayscale vehicle image. For each pixel, its pixel value is compared with the noise judgment threshold range determined in step S222. If the pixel value is less than the lower limit or greater than the upper limit, the pixel is marked as a noise pixel.

[0028] In practice, a nested loop can be used to iterate over each pixel in a grayscale vehicle image. For each pixel, its value is obtained and then compared with the noise threshold. For example, in a grayscale vehicle image taken at a highway toll booth, if a pixel with a value of 130 is found, while the noise threshold is between 80 and 120, the pixel will be marked as a noise pixel.

[0029] Step S224: performing pixel value correction processing on the noise pixel point, replacing the pixel value of the noise pixel point with the average pixel value of the non-noise pixel points in the preset neighborhood around it.

[0030] Pixel value correction is used to eliminate the effects of noise pixels on the image and adjust the pixel values ​​of the noise pixels to reasonable values. The preset neighborhood is a specific size area centered on the noise pixel, such as 3×3, 5×5, etc.

[0031] During pixel value correction, for each pixel marked as a noise pixel, the non-noise pixels within a preset neighborhood are found, the average pixel value of these non-noise pixels is calculated, and the pixel value of the noise pixel is then replaced with this average. For example, within a preset 3×3 neighborhood, for a noise pixel, the non-noise pixels in the eight surrounding pixels are found, their average pixel values ​​are calculated, and the pixel value of the noise pixel is updated to this average. In vehicle images taken at highway toll booths, pixel value correction effectively eliminates the interference of noise pixels and improves image quality.

[0032] Step S225: Smoothing the corrected grayscale vehicle image, performing weighted average calculation on the pixel values ​​in the corrected grayscale vehicle image through a sliding average window, and obtaining a denoised grayscale vehicle image.

[0033] Smoothing is done to further eliminate noise and discontinuities in the image, making it smoother. The sliding average window is a fixed-size window that slides on the corrected grayscale vehicle image at a certain step size.

[0034] During smoothing, a sliding average window is placed at a certain position within the corrected grayscale vehicle image. The weighted average of the pixel values ​​for all pixels within the window is calculated, and this average is used as the new pixel value for the center pixel of the window. The window is then moved to the next position using a preset step size, and the above calculation process is repeated until the entire corrected grayscale vehicle image is traversed. For example, using a 3×3 sliding average window, the weight of each pixel within the window can be set equal, that is, each pixel has a weight of 1 / 9.

[0035] Step S230: performing edge detection processing on the denoised grayscale vehicle image to identify edge lines of the vehicle outline and obtain a vehicle edge image.

[0036] Edge detection involves finding the edge lines of a vehicle's outline in a de-noised grayscale vehicle image. These lines represent the boundary between the vehicle and the surrounding background. By identifying these lines, the vehicle's position and shape can be accurately determined. A vehicle edge image consists solely of the vehicle's outline edge lines.

[0037] As an embodiment, step S230 performs edge detection processing on the denoised grayscale vehicle image to identify the edge lines of the vehicle outline to obtain a vehicle edge image. Specifically, the following steps S231 to S237 may be included: Step S231: performing gradient calculation processing on the denoised grayscale vehicle image, calculating the pixel value change rates in the horizontal and vertical directions respectively, and obtaining a horizontal gradient image and a vertical gradient image.

[0038] Gradient calculation is performed by calculating the rate of change of pixel values ​​in the horizontal and vertical directions for each pixel in the denoised grayscale vehicle image. This determines areas of image where pixel values ​​vary dramatically. These areas typically correspond to image edges. The horizontal gradient image reflects the rate of change of pixel values ​​in the horizontal direction of the denoised grayscale vehicle image, while the vertical gradient image reflects the rate of change of pixel values ​​in the vertical direction.

[0039] When calculating the rate of change of pixel values ​​in the horizontal and vertical directions, a differential operator, such as the Sobel operator, can be used. For each pixel in the denoised grayscale vehicle image, the Sobel operator is used to calculate the rate of change of its pixel values ​​in the horizontal and vertical directions, respectively. For example, for a pixel, the horizontal and vertical gradient values ​​of the pixel are calculated by calculating the difference between its pixel values ​​and those of adjacent pixels. The horizontal and vertical gradient values ​​of all pixels are combined to form a horizontal gradient image and a vertical gradient image, respectively. In the highway toll booth scenario, gradient calculation processing can highlight the changes in pixel values ​​at the edges of the vehicle image.

[0040] Step S232: performing gradient amplitude calculation processing on the horizontal gradient image and the vertical gradient image, performing square and square root operations on the horizontal gradient values ​​and the vertical gradient values ​​at corresponding positions to obtain gradient amplitudes, thereby obtaining a gradient amplitude image.

[0041] The gradient magnitude calculation process comprehensively considers the changes in pixel values ​​in the horizontal and vertical directions to obtain the gradient magnitude of each pixel. The gradient magnitude reflects the edge strength of the pixel in the image. The larger the gradient magnitude, the more likely the pixel is an edge point.

[0042] When calculating the gradient magnitude, for each pixel at the corresponding position in the horizontal gradient image and the vertical gradient image, the horizontal gradient value and the vertical gradient value are squared, respectively. These two squared values ​​are then added together, and the square root of the result is taken to obtain the gradient magnitude of the pixel. The gradient magnitudes of all pixels are combined to form a gradient magnitude image.

[0043] Step S233: performing gradient direction calculation processing on the gradient magnitude image, calculating the gradient direction angle of each pixel point by using the inverse tangent function, and obtaining a gradient direction image.

[0044] The gradient direction calculation process is used to determine the gradient direction of each pixel in the gradient magnitude image. The gradient direction angle represents the edge direction of the pixel. The inverse tangent function is a mathematical function used to calculate angles.

[0045] When performing gradient direction calculation processing, for each pixel in the gradient magnitude image, the inverse tangent function is used to calculate the gradient direction angle of the pixel based on its horizontal gradient value and vertical gradient value. For example, if the horizontal gradient value of a pixel is x and the vertical gradient value is y, then the gradient direction angle of the pixel is arctan(y / x). The gradient direction angles of all pixels are combined to form a gradient direction image. In the highway toll station scenario, the gradient direction calculation processing can obtain the edge direction information of each pixel in the vehicle image, which is very important for the subsequent non-maximum suppression processing.

[0046] Step S234: performing non-maximum suppression processing on the gradient magnitude image, retaining the pixel points with local maximum gradient magnitude in the gradient direction, and obtaining a suppressed gradient magnitude image.

[0047] Non-maximum suppression is used to refine edges by retaining only pixels with local maxima in the gradient direction and setting the gradient magnitudes of all other non-local maxima pixels to 0. The suppressed gradient magnitude image is the image obtained after non-maximum suppression, containing only the local maxima pixels on the edges.

[0048] When performing non-maximum suppression, for each pixel in the gradient magnitude image, the gradient magnitude of that pixel is compared with that of its adjacent pixels along its gradient direction. If the gradient magnitude of that pixel is not a local maximum, its gradient magnitude is set to 0; if it is a local maximum, its gradient magnitude is retained. For example, along the gradient direction of a pixel, the gradient magnitudes of that pixel are compared with those of its two adjacent pixels. If the gradient magnitude of that pixel is the largest, its gradient magnitude is retained; otherwise, its gradient magnitude is set to 0.

[0049] Step S235: performing double threshold processing on the suppressed gradient amplitude image, and classifying the pixels into strong edge points, weak edge points, and non-edge points according to the first threshold and the second threshold.

[0050] Double threshold processing is used to further filter out true edge points while retaining some possible edge points. The first threshold and the second threshold are two pre-set gradient amplitude thresholds used to distinguish different types of edge points. The first threshold is a high threshold and the second threshold is a low threshold. Strong edge points are pixels with a gradient amplitude greater than the first threshold. These pixels are likely to be true edge points; weak edge points are pixels with a gradient amplitude between the second threshold and the first threshold. These pixels may be edge points or noise points; non-edge points are pixels with a gradient amplitude less than the second threshold. These pixels are unlikely to be edge points.

[0051] When performing dual threshold processing, for each pixel in the suppressed gradient magnitude image, its gradient magnitude is compared with a first threshold and a second threshold. If the gradient magnitude is greater than the first threshold (high threshold), the pixel is marked as a strong edge point; if the gradient magnitude is between the second threshold and the first threshold, it is marked as a weak edge point; if the gradient magnitude is less than the second threshold (low threshold), it is marked as a non-edge point.

[0052] Step S236: performing edge connection processing on the weak edge points, determining the weak edge points connected to the strong edge points as edge points, and obtaining an edge point set.

[0053] Edge connection processing is used to connect possible edge points (weak edge points) with confirmed edge points (strong edge points) to further determine the true edge. The edge point set is the set of all pixels that are confirmed to be edge points after edge connection processing.

[0054] As an implementation manner, step S236 performs edge connection processing on the weak edge points, determines the weak edge points connected to the strong edge points as edge points, and obtains an edge point set. Specifically, the following steps S2361 to S2366 may be included: Step S2361: define a neighborhood search range of the weak edge point, where the neighborhood search range is a rectangular area of ​​a preset size centered on the weak edge point.

[0055] The neighborhood search range is used to determine the search area when searching for strong edge points connected to weak edge points. The preset rectangular area is a rectangular area pre-set according to actual conditions, such as 3×3, 5×5, etc.

[0056] When defining a neighborhood search range, a preset rectangular area is defined with the weak edge point as the center. For example, a 3×3 rectangular area centered on the weak edge point and encompassing the weak edge point and its eight surrounding pixels is selected. In highway toll booth scenarios, defining a neighborhood search range clarifies the search range for strong edge points connected to the weak edge point.

[0057] Step S2362: traverse each weak edge point in the weak edge point set, and search for a strong edge point within the neighborhood search range of the currently traversed weak edge point.

[0058] Traversal is to visit each weak edge point in the weak edge point set in turn. For each weak edge point, find out whether there is a strong edge point in its neighborhood search range.

[0059] In actual operation, a loop is used to traverse each weak edge point in the weak edge point set. For the currently traversed weak edge point, the pixels in its neighborhood search range are checked to see if they are strong edge points.

[0060] Step S2363: If there is at least one strong edge point in the neighborhood search range, the pixel distance between the current weak edge point and each searched strong edge point is calculated, where the pixel distance is the Euclidean distance between the two points in the image coordinate system.

[0061] Pixel distance is used to measure the spatial distance between two pixels, and Euclidean distance is a method for calculating the distance between two points.

[0062] When calculating pixel distance, for the currently traversed weak edge point and each strong edge point searched within its neighborhood search range, the distance between them is calculated using the Euclidean distance formula based on their coordinates in the image coordinate system.

[0063] Step S2364: Compare the pixel distance with a preset distance threshold. If there is a strong edge point whose pixel distance is less than the preset distance threshold, determine that the current weak edge point is a valid weak edge point.

[0064] The preset distance threshold is a pre-set distance standard used to judge whether the distance between the weak edge point and the strong edge point is close enough to determine whether the weak edge point is a valid weak edge point.

[0065] During the comparison, the pixel distance calculated in step S2363 is compared with the preset distance threshold. If there is a strong edge point whose pixel distance is less than the preset distance threshold, it is considered that the distance between the current weak edge point and the strong edge point is close enough, and the weak edge point is determined to be a valid weak edge point.

[0066] Step S2365: If there is no strong edge point in the neighborhood search range, or the pixel distances of all searched strong edge points are greater than or equal to the preset distance threshold, the current weak edge point is determined to be an invalid weak edge point.

[0067] When there is no strong edge point in the neighborhood search range, or although there is a strong edge point, the pixel distance between them and the current weak edge point is greater than or equal to the preset distance threshold, it is considered that the association between the weak edge point and the strong edge point is not close enough and it is judged as an invalid weak edge point.

[0068] In actual judgment, for a currently traversed weak edge point, if there are no strong edge points within its neighborhood search range, or if the pixel distances between all strong edge points and the weak edge point are greater than or equal to a preset distance threshold, the weak edge point is marked as invalid. For example, if the preset distance threshold is 5 and all calculated pixel distances are greater than or equal to 5, the weak edge point is considered invalid. In the highway toll booth scenario, this judgment can eliminate weak edge points that are not closely associated with strong edge points.

[0069] Step S2366: Add all valid weak edge points to the strong edge point set to obtain an updated edge point set, and determine the updated edge point set as the edge point set.

[0070] After the judgment in steps S2364 and S2365, all pixel points judged to be valid weak edge points are added to the strong edge point set to form an updated edge point set, which is the final edge point set.

[0071] During the operation, all pixels identified as valid weak edge points are traversed and added to the strong edge point set in sequence. For example, in a vehicle image at a highway toll booth, after edge connection processing, all valid weak edge points are added to the strong edge point set, resulting in a more complete edge point set that includes all confirmed edge points in the vehicle image.

[0072] Step S237: Mark corresponding pixel points on the blank image based on the edge point set to obtain a vehicle edge image.

[0073] A blank image is an image whose initial pixel values ​​are all 0. The corresponding pixels in the edge point set are marked on the blank image, that is, the pixel values ​​of these pixels are set to specific values ​​(such as 255), thereby forming a vehicle edge image.

[0074] In practice, for each pixel in the edge point set, its corresponding position is found in the blank image, and the pixel value at that position is set to 255. For example, if a pixel in the edge point set has coordinates (10, 20), the pixel value at the position with coordinates (10, 20) in the blank image is set to 255. In the case of a highway toll booth, this method can clearly display the edge lines of a vehicle on a blank image, resulting in a vehicle edge image.

[0075] Step S240: determining a minimum circumscribed rectangular boundary of the vehicle area based on the edge lines in the vehicle edge image. The minimum circumscribed rectangular boundary is the minimum rectangular area that completely contains all the vehicle edge lines.

[0076] The minimum bounding rectangle boundary is a minimum rectangular area that can completely contain all edge lines in the vehicle edge image, and can accurately define the position and range of the vehicle in the image.

[0077] To determine the minimum enclosing rectangle (MCR), the team first locates the leftmost, rightmost, topmost, and bottommost pixels of all edge lines in the vehicle edge image. Based on these boundary pixels, a rectangular region is then determined that completely encompasses all edge lines and is the smallest of all possible rectangular regions that could encompass these lines. For example, in a vehicle edge image at a highway toll booth, the MRC boundary is determined by finding the boundary pixels of the edge lines. This boundary accurately defines the vehicle's position.

[0078] Step S250: cropping a corresponding image region from the vehicle image according to the minimum circumscribed rectangle boundary to obtain a vehicle region image; and combining all the cropped vehicle region images into a vehicle region image set.

[0079] Cropping is to cut out the corresponding image area from the original vehicle image according to the minimum circumscribed rectangle boundary. The cut out image area is the vehicle area image, which only contains the main part of the vehicle.

[0080] During the cropping operation, the image corresponding to the rectangular region of the vehicle image is extracted based on the coordinate information of the minimum bounding rectangle determined in step S240. For example, if the coordinates of the upper left corner of the minimum bounding rectangle are (x1, y1) and the coordinates of the lower right corner are (x2, y2), then the image region with coordinates ranging from (x1, y1) to (x2, y2) is extracted from the vehicle image to obtain the vehicle region image. All vehicle region images cropped in this manner are combined to form a vehicle region image set.

[0081] Step S300: performing feature extraction processing on the vehicle area image set to obtain a vehicle feature set, where each vehicle feature in the vehicle feature set corresponds to a feature representation of a vehicle area image.

[0082] Feature extraction involves extracting information representing the characteristics of each vehicle region image from the vehicle region image set, and then assembling this information into a vehicle feature set. Vehicle features are an abstract representation of a vehicle region image that can reflect information such as the vehicle's type, shape, and structure.

[0083] As an embodiment, step S300 performs feature extraction processing on the vehicle area image set to obtain a vehicle feature set, which may specifically include the following steps S310 to S350: Step S310: performing size unification processing on each vehicle area image in the vehicle area image set, and adjusting vehicle area images of different sizes into standard vehicle area images of a preset size.

[0084] The purpose of resizing is to eliminate the impact of size differences between vehicle area images, facilitating subsequent feature extraction and processing. The preset size is a fixed size that is predetermined, such as 224×224, 300×300, etc.

[0085] During resizing, each vehicle area image in the vehicle area image set is resized to a standard vehicle area image of the preset size using an image scaling algorithm. For example, if the preset size is 224×224, a 300×400 vehicle area image can be resized to a standard 224×224 vehicle area image using a bilinear interpolation algorithm. In the highway toll booth scenario, resizing ensures that all vehicle area images have the same size, facilitating subsequent feature extraction.

[0086] Step S320: performing channel separation processing on the standard vehicle area image, separating the multi-channel pixel information of the standard vehicle area image into multiple single-channel pixel matrices, and obtaining a single-channel pixel matrix set.

[0087] Channel separation is the process of extracting pixel information from multiple channels contained in the standard vehicle area image to form multiple single-channel pixel matrices. The standard vehicle area image is usually a color image containing pixel information from three channels: red, green, and blue.

[0088] During channel separation, for each pixel in the standard vehicle area image, the pixel values ​​of the red, green, and blue channels are extracted separately to form three single-channel pixel matrices. For example, for a pixel with a value of (200, 150, 100), the pixel value of the red channel (200), the pixel value of the green channel (150), and the pixel value of the blue channel (100) are extracted separately and stored in the corresponding single-channel pixel matrices. Combining all the single-channel pixel matrices together forms a set of single-channel pixel matrices. In the highway toll booth scenario, channel separation can separate the multi-channel information of the standard vehicle area image, facilitating subsequent convolution filtering.

[0089] Step S330: performing convolution filtering processing on each single-channel pixel matrix in the single-channel pixel matrix set, performing pixel value weighted calculation on the single-channel pixel matrix through a sliding window, and obtaining multiple initial feature matrices.

[0090] Convolution filtering uses a sliding window to perform weighted calculations on pixel values ​​on a single-channel pixel matrix to extract the image's feature information. The initial feature matrix is ​​the matrix obtained after convolution filtering and contains the feature information of the single-channel pixel matrix.

[0091] As an embodiment, step S330 performs convolution filtering on each single-channel pixel matrix in the single-channel pixel matrix set, performs pixel value weighted calculation on the single-channel pixel matrix through a sliding window, and obtains multiple initial feature matrices. Specifically, the following steps S331 to S335 may be included: Step S331: for each single-channel pixel matrix in the single-channel pixel matrix set, initialize multiple convolution kernels of different sizes, each convolution kernel including a preset number of weight coefficients.

[0092] A convolution kernel is a small matrix containing a set of pre-set weight coefficients. Convolution kernels of different sizes can extract features at different scales. The pre-set number of weight coefficients is predetermined based on the size of the convolution kernel.

[0093] When initializing the convolution kernel, multiple convolution kernels of different sizes are created for each single-channel pixel matrix in the set, such as 3×3, 5×5, and so on. The weight coefficients in each convolution kernel can be set through random initialization or other methods. For example, a 3×3 convolution kernel contains 9 weight coefficients, which can be randomly initialized to a set of decimals. In the highway toll booth scenario, by initializing the convolution kernels of different sizes, feature information of different scales can be extracted from the single-channel pixel matrix.

[0094] Step S332: Slide each convolution kernel as a sliding window on the corresponding single-channel pixel matrix according to a preset step size. At each sliding position, multiply and sum the weight coefficient of the convolution kernel with the pixel value of the corresponding area of ​​the single-channel pixel matrix to obtain the eigenvalue at the sliding position.

[0095] The preset step size is the distance the convolution kernel moves each time it slides across the single-channel pixel matrix. At each sliding position, the convolution kernel weight coefficient is multiplied and summed with the pixel values ​​of the corresponding area in the single-channel pixel matrix to obtain an eigenvalue.

[0096] When sliding and calculating, the convolution kernel is placed at a starting position of the single-channel pixel matrix, the weight coefficient of the convolution kernel is multiplied one by one by the pixel values ​​of the corresponding area of ​​the single-channel pixel matrix, and then these products are added to obtain the eigenvalue at the sliding position. For example, for a 3×3 convolution kernel and a 3×3 single-channel pixel matrix corresponding area, the 9 weight coefficients of the convolution kernel are multiplied by the 9 pixel values ​​of the corresponding area of ​​the single-channel pixel matrix respectively, and then these 9 products are added to obtain the eigenvalue at the sliding position. Then, the convolution kernel is moved to the next position according to the preset step size, and the above calculation process is repeated until the entire single-channel pixel matrix is ​​traversed. In the scenario of a highway toll station, multiple eigenvalues ​​can be extracted on the single-channel pixel matrix in this way.

[0097] Step S333: Arrange the eigenvalues ​​at all sliding positions in a sliding order to form an initial eigenmatrix related to the size of the single-channel pixel matrix.

[0098] The eigenvalues ​​at all sliding positions calculated in step S332 are arranged according to the sliding order of the convolution kernel on the single-channel pixel matrix to form a matrix. This matrix is ​​the initial feature matrix, and its size is related to the size of the single-channel pixel matrix and the sliding step size of the convolution kernel.

[0099] When arranging the eigenvalues, the eigenvalues ​​at each sliding position are sequentially filled into the corresponding positions of the initial feature matrix according to the sliding order of the convolution kernel. For example, if the convolution kernel starts sliding from the upper left corner of a single-channel pixel matrix, the eigenvalues ​​at each sliding position are sequentially filled into the initial feature matrix in a left-to-right and top-to-bottom order. In the highway toll booth scenario, this method can be used to obtain an initial feature matrix that reflects the feature information of a single-channel pixel matrix.

[0100] Step S334: Perform the above convolution filtering process on each single-channel pixel matrix using convolution kernels of all different sizes to obtain multiple initial feature matrices.

[0101] For each single-channel pixel matrix in the single-channel pixel matrix set, all convolution kernels of different sizes initialized in step S331 are used to perform convolution filtering processing in steps S332 and S333 respectively to obtain multiple initial feature matrices.

[0102] During operation, for a single-channel pixel matrix, convolution filtering is performed using convolution kernels of different sizes in sequence to obtain the initial feature matrix corresponding to each convolution kernel. For example, for a single-channel pixel matrix, convolution filtering is performed using convolution kernels of two different sizes, 3×3 and 5×5, to obtain two initial feature matrices respectively. In the highway toll station scenario, by using convolution kernels of different sizes, feature information of different scales can be extracted from the single-channel pixel matrix, resulting in multiple initial feature matrices.

[0103] Step S335: multiple initial feature matrices corresponding to the same single-channel pixel matrix are grouped into an initial feature matrix group. The multiple initial feature matrix groups corresponding to the single-channel pixel matrix set together constitute multiple initial feature matrices.

[0104] Multiple initial feature matrices obtained by convolution filtering the same single-channel pixel matrix using convolution kernels of different sizes are combined to form an initial feature matrix group. Each single-channel pixel matrix in the single-channel pixel matrix set corresponds to an initial feature matrix group, and these initial feature matrix groups together constitute multiple initial feature matrices.

[0105] When combining initial feature matrices, for each single-channel pixel matrix, the corresponding multiple initial feature matrices are arranged in a certain order to form an initial feature matrix group. For example, for a single-channel pixel matrix, two initial feature matrices are obtained by performing convolution filtering using two different convolution kernel sizes of 3×3 and 5×5. These two initial feature matrices are then combined into an initial feature matrix group. By combining the initial feature matrix groups corresponding to all single-channel pixel matrices in the single-channel pixel matrix set, multiple initial feature matrices are formed.

[0106] Step S340: performing feature aggregation processing on the multiple initial feature matrices, superimposing and combining the initial feature matrices obtained from different sliding windows in a preset order to obtain an aggregated feature matrix.

[0107] The purpose of feature aggregation is to integrate the feature information in the initial feature matrices obtained from different sliding windows to form a more representative aggregate feature matrix. The preset order is the matrix superposition order predetermined according to the actual situation.

[0108] As an implementation method, step S340 performs feature aggregation processing on multiple initial feature matrices, superimposing and combining the initial feature matrices obtained from different sliding windows in a preset order to obtain an aggregated feature matrix. Specifically, the following steps S341 to S345 may be included: Step S341: resizing each of the multiple initial characteristic matrices, and unifying the initial characteristic matrices of different sizes into a standard initial characteristic matrix of the same size through interpolation operation.

[0109] Resizing is done to eliminate size differences between different initial feature matrices, facilitating subsequent overlay and combination. Interpolation is a method used to estimate the value of unknown data points between known data points. It can be used to resize initial feature matrices of different sizes to the same size.

[0110] During resizing, each initial feature matrix is ​​resized to a preset standard size using an interpolation algorithm. For example, a 20×20 initial feature matrix can be resized to a 30×30 standard initial feature matrix using bilinear interpolation. In the highway toll booth scenario, resizing ensures that all initial feature matrices have the same size, facilitating subsequent feature aggregation operations.

[0111] Step S342: performing channel dimension expansion processing on the standard initial feature matrix, adding a channel dimension identifier to each standard initial feature matrix, and obtaining a standard initial feature matrix with a channel identifier.

[0112] The channel dimension expansion process is to introduce channel dimension identification into the standard initial feature matrix so that different initial feature matrices can be accurately distinguished in subsequent splicing operations.

[0113] During channel dimension expansion, a channel dimension identifier is added to each standard initial feature matrix. For example, a 30×30 standard initial feature matrix is ​​expanded to a 30×30×1 standard initial feature matrix with a channel identifier, where 1 represents the channel dimension identifier. In the highway toll booth scenario, channel dimension expansion can assign a unique channel dimension identifier to each standard initial feature matrix, facilitating subsequent channel dimension concatenation operations.

[0114] Step S343: All standard initial feature matrices with channel identifiers are spliced ​​in the order of channel dimension identifiers to obtain a multi-channel feature matrix.

[0115] Channel dimension splicing is to splice multiple standard initial feature matrices with channel identifiers in the channel dimension to form a multi-channel feature matrix.

[0116] When performing channel-dimensional concatenation, all standard initial feature matrices with channel identifiers are concatenated sequentially in the order of their channel dimension identifiers. For example, if there are three standard initial feature matrices with channel identifiers 1, 2, and 3, concatenating them along the channel dimension yields a 30×30×3 multi-channel feature matrix. In the highway toll booth scenario, channel-dimensional concatenation allows the feature information of multiple initial feature matrices to be integrated into a single multi-channel feature matrix.

[0117] Step S344: performing feature importance weighting processing on the multi-channel feature matrix, and assigning a preset importance weight to the feature matrix of each channel.

[0118] The feature importance weighting process is to assign different importance weights to the feature matrix of each channel according to its importance, so that important features can be highlighted and unimportant features can be suppressed in subsequent calculations.

[0119] As an implementation manner, step S344 performs feature importance weighting processing on the multi-channel feature matrix, assigning a preset importance weight to the feature matrix of each channel, which may specifically include the following steps S3441 to S3448: Step S3441: Obtain the size parameters of the convolution kernel corresponding to the feature matrix of each channel, where the size parameters include the height and width of the convolution kernel.

[0120] The size parameters of the convolution kernel are the height and width of the convolution kernel. These parameters reflect the size and shape of the convolution kernel. The importance of feature information extracted by convolution kernels of different sizes may be different.

[0121] When obtaining the size parameters, for the feature matrix of each channel in the multi-channel feature matrix, find the corresponding convolution kernel and record the height and width values ​​of the convolution kernel.

[0122] Step S3442: Calculate the receptive field area of ​​the convolution kernel based on the height value and width value of the convolution kernel. The receptive field area is the product of the height value and the width value of the convolution kernel.

[0123] The receptive field area is the size of the area that the convolution kernel can cover on a single-channel pixel matrix, reflecting the range of feature information that the convolution kernel can extract.

[0124] When calculating the receptive field area, multiply the height and width of the convolution kernel to get the receptive field area of ​​the convolution kernel. By calculating the receptive field area of ​​the convolution kernel, we can understand the range of feature information that each convolution kernel can extract.

[0125] Step S3443: Count the number of weight coefficients contained in each convolution kernel to obtain the number of convolution kernel parameters.

[0126] The number of convolution kernel parameters is the total number of weight coefficients contained in the convolution kernel, reflecting the complexity of the convolution kernel.

[0127] When counting the number of convolution kernel parameters, for each convolution kernel, the number of weight coefficients it contains is calculated. For example, for a 3×3 convolution kernel, which contains 9 weight coefficients, the number of convolution kernel parameters is 9.

[0128] Step S3444: performing an association calculation between the receptive field area and the number of convolution kernel parameters to obtain an initial weight coefficient, which is the ratio of the receptive field area to the number of convolution kernel parameters.

[0129] The initial weight coefficient is a coefficient calculated based on the receptive field area of ​​the convolution kernel and the number of convolution kernel parameters, which reflects the importance of the convolution kernel in feature extraction.

[0130] When performing the association calculation, the receptive field area calculated in step S3442 is divided by the number of convolution kernel parameters obtained by counting in step S3443 to obtain the initial weight coefficient.

[0131] Step S3445: Calculate the pixel value variance of the feature matrix of each channel in the multi-channel feature matrix to obtain the feature matrix variance value.

[0132] Pixel value variance is an indicator to measure the degree of discreteness of pixel values ​​in the feature matrix, reflecting the fluctuation of pixel values ​​in the feature matrix.

[0133] When calculating pixel value variance, the variance of all pixel values ​​in each channel's feature matrix is ​​calculated. For example, for a channel's feature matrix, the mean of all pixel values ​​is first calculated. The square of the difference between each pixel value and the mean is then calculated. These squared values ​​are then added together and divided by the total number of pixels to obtain the pixel value variance for that channel's feature matrix. In the case of a highway toll booth, calculating the variance of the feature matrix can provide an understanding of the fluctuations in pixel values ​​in each channel's feature matrix, providing a basis for the subsequent calculation of dynamic weight coefficients.

[0134] Step S3446: Multiply the initial weight coefficient by the characteristic matrix variance value of the corresponding channel to obtain a dynamic weight coefficient.

[0135] The dynamic weight coefficient is a coefficient calculated based on the initial weight coefficient and the variance value of the feature matrix, which comprehensively considers the importance of the convolution kernel and the fluctuation of the pixel values ​​in the feature matrix.

[0136] When performing the product operation, the initial weight coefficient calculated in step S3444 is multiplied by the characteristic matrix variance value of the corresponding channel calculated in step S3445 to obtain a dynamic weight coefficient.

[0137] Step S3447: normalize the dynamic weight coefficients so that the sum of the dynamic weight coefficients of all channels is 1, thereby obtaining a normalized dynamic weight coefficient.

[0138] Normalization is performed to make the dynamic weight coefficients of all channels comparable and adjust their sum to 1.

[0139] During normalization, the sum of the dynamic weight coefficients of all channels is calculated, and then the dynamic weight coefficient of each channel is divided by this sum to obtain the normalized dynamic weight coefficient.

[0140] Step S3448: Using the normalized dynamic weight coefficient as the importance weight of the feature matrix of each channel, assigning a corresponding importance weight to each channel feature matrix in the multi-channel feature matrix.

[0141] After obtaining the normalized dynamic weight coefficient, it is used as the importance weight of the feature matrix of each channel, and each channel feature matrix in the multi-channel feature matrix is ​​weighted.

[0142] When performing weighted processing, for each channel feature matrix in the multi-channel feature matrix, all its pixel values ​​are multiplied by the corresponding normalized dynamic weight coefficient. For example, for a channel feature matrix, its normalized dynamic weight coefficient is 0.3, and each pixel value in the channel feature matrix is ​​multiplied by 0.3.

[0143] Step S345: performing pixel value addition operation on the weighted multi-channel feature matrix in the channel dimension to obtain a two-dimensional aggregated feature matrix.

[0144] The pixel value addition operation in the channel dimension is to add the pixel values ​​of different channels at each position in the weighted multi-channel feature matrix to obtain a two-dimensional matrix, which is the aggregate feature matrix.

[0145] When performing the addition operation, for each position in the weighted multi-channel feature matrix, the pixel values ​​of all channels at that position are added together to obtain the pixel value of that position in the aggregated feature matrix. For example, for a 30×30×3 weighted multi-channel feature matrix, at each 30×30 position, the pixel values ​​of the three channels are added together to obtain a 30×30 two-dimensional aggregated feature matrix.

[0146] Step S350: performing dimensionality reduction processing on the aggregated feature matrix, converting the two-dimensional aggregated feature matrix into a one-dimensional feature vector, obtaining vehicle features, and forming all vehicle features into a vehicle feature set.

[0147] Dimensionality reduction is used to reduce the dimension of features, converting the two-dimensional aggregate feature matrix into a one-dimensional feature vector to facilitate subsequent classification. Vehicle features are the one-dimensional feature vectors obtained after dimensionality reduction, which can represent the characteristic information of the vehicle area image.

[0148] When performing dimensionality reduction, you can use vector flattening to arrange the elements of a two-dimensional aggregate feature matrix into a one-dimensional vector in a certain order. For example, for a 30×30 two-dimensional aggregate feature matrix, you can arrange it row by row into a one-dimensional feature vector of length 900.

[0149] Step S400: classifying the vehicle feature set using a preset vehicle classification model to obtain a vehicle classification result. The vehicle classification result is used to indicate the vehicle type corresponding to each vehicle image.

[0150] The preset vehicle classification model is a pre-trained model that can classify vehicles based on the input vehicle features. The vehicle classification result is the result of the model classifying the vehicle feature set, clearly defining the vehicle type corresponding to each vehicle image, such as car, truck, bus, etc.

[0151] As an embodiment, step S400 classifies the vehicle feature set using a preset vehicle classification model to obtain a vehicle classification result, which may specifically include the following steps S410 to S440: Step S410: Input each vehicle feature in the vehicle feature set into a feature encoding module of a preset vehicle classification model, convert the vehicle feature into an encoding feature vector of a fixed length, and obtain an encoding feature vector set.

[0152] The feature encoding module is a module in the pre-defined vehicle classification model. Its function is to convert the input vehicle features into fixed-length encoded feature vectors to facilitate subsequent processing and classification. The encoded feature vector set is a set of fixed-length encoded feature vectors obtained after all vehicle features are converted by the feature encoding module.

[0153] As an embodiment, step S410 inputs each vehicle feature in the vehicle feature set into a feature encoding module of a preset vehicle classification model, converts the vehicle feature into a fixed-length encoding feature vector, and obtains an encoding feature vector set. Specifically, the following steps S411 to S415 may be included: Step S411: Input the vehicle features into the first coding layer of the feature coding module, perform linear transformation on the vehicle features, and obtain a first coding vector by multiplying the weight matrix and the vehicle features.

[0154] The first encoding layer, the first layer in the feature encoding module, performs a linear transformation on the input vehicle features. This linear transformation multiplies the vehicle features with a weight matrix to convert them into another vector, which is the first encoding vector.

[0155] During linear transformation, the vehicle features are represented as a vector, and the weight matrix of the first coding layer is multiplied by this vector to produce the first coding vector. For example, if the vehicle features are a vector of length 900 and the weight matrix of the first coding layer is a 900×500 matrix, the matrix multiplication yields a first coding vector of length 500. In the highway toll booth scenario, the linear transformation of the first coding layer can provide a preliminary feature conversion of the vehicle features.

[0156] Step S412: performing nonlinear activation processing on the first coding vector, performing nonlinear transformation on each element in the first coding vector through a preset activation function to obtain an activated coding vector.

[0157] Nonlinear activation processing is used to introduce nonlinear factors and increase the expressive power of the model. The default activation function is a nonlinear function, such as the ReLU function and the Sigmoid function.

[0158] When performing nonlinear activation processing, for each element in the first code vector, it is input into the preset activation function for calculation to obtain a new value, and these new values ​​are composed of the activation code vector. For example, if the ReLU function is used as the preset activation function, for an element x in the first code vector, if x>0, then the element at the corresponding position in the activation code vector is x; if x≤0, then the element at the corresponding position in the activation code vector is 0.

[0159] Step S413: Input the activated coding vector to the second coding layer of the feature coding module, perform linear transformation on the activated coding vector, and obtain a second coding vector by multiplying the activated coding vector by another weight matrix.

[0160] The second encoding layer is the second layer in the feature encoding module, and further linearly transforms the activation encoding vector. The other weight matrix is ​​the weight matrix of the second encoding layer, which is different from the weight matrix of the first encoding layer.

[0161] During the linear transformation, the activation code vector is represented as a vector, and the weight matrix of the second coding layer is multiplied by the vector to obtain the second code vector. For example, if the activation code vector is a vector of length 500 and the weight matrix of the second coding layer is a 500×200 matrix, the matrix multiplication operation results in a second code vector of length 200.

[0162] Step S414: performing regularization processing on the second coding vector, adjusting the element values ​​of the second coding vector to within a preset numerical range, and obtaining a regularized coding vector as the coding feature vector.

[0163] Regularization is used to prevent overfitting of the model by adjusting the element values ​​of the second encoding vector to a reasonable range. The preset value range is a predetermined value interval, such as [0, 1].

[0164] As an implementation manner, step S414 performs regularization processing on the second code vector, adjusting the element values ​​of the second code vector to within a preset numerical range to obtain a regularized code vector. Specifically, the following steps S4141 to S4148 may be included: Step S4141: Calculate the mean of all element values ​​in the second coding vector to obtain the coding mean.

[0165] The coding mean is the average value of all element values ​​in the second coding vector, reflecting the overall level of the second coding vector.

[0166] To calculate the encoding mean, add the values ​​of all elements in the second encoding vector and divide by the total number of elements to get the encoding mean. For example, if the second encoding vector is a vector of length 200, add the values ​​of its 200 elements and divide by 200 to get the encoding mean.

[0167] Step S4142: Calculate the sum of squares of deviations between all element values ​​in the second coding vector and the coding mean, and then divide it by the number of elements to obtain the coding variance.

[0168] The coding variance is an indicator to measure the degree of dispersion of element values ​​in the second coding vector, reflecting the fluctuation of element values ​​relative to the coding mean.

[0169] To calculate the encoding variance, first calculate the square of the difference between each element value and the encoding mean in the second encoding vector. These squared values ​​are summed and then divided by the total number of elements to obtain the encoding variance. For example, for a second encoding vector of length 200, calculate the square of the difference between each element value and the encoding mean, sum these 200 squared values, and divide by 200 to obtain the encoding variance.

[0170] Step S4143: Determine a dynamic adjustment range based on the coding mean and coding variance, wherein the lower limit of the dynamic adjustment range is the coding mean minus the coding variance of a preset multiple, and the upper limit is the coding mean plus the coding variance of a preset multiple.

[0171] The dynamic adjustment range is a numerical interval determined according to the coding mean and coding variance, and the preset multiple is a constant pre-set according to actual conditions.

[0172] When determining the dynamic adjustment range, the encoding variance of a preset multiple is subtracted and added to the encoding mean value as the center to obtain the lower limit value and the upper limit value of the dynamic adjustment range.

[0173] Step S4144: traverse each element value in the second encoding vector, and if the element value is less than the lower limit value of the dynamic adjustment range, adjust the element value to the lower limit value.

[0174] The traversal is to sequentially access each element value in the second encoding vector. For each element value, it is compared with the lower limit of the dynamic adjustment range. If it is less than the lower limit, it is adjusted to the lower limit.

[0175] Step S4145: If the element value is greater than the upper limit of the dynamic adjustment range, the element value is adjusted to the upper limit.

[0176] Step S4146: If the element value is between the lower limit value and the upper limit value, the element value is kept unchanged to obtain a preliminary adjusted coding vector.

[0177] When the element value in the second code vector is between the lower limit and the upper limit of the dynamic adjustment range, no adjustment is performed and the element value remains unchanged. After processing all element values, a preliminary adjustment code vector is obtained.

[0178] Step S4147: Perform linear scaling processing on the preliminary adjusted coding vector, and map the element values ​​of the preliminary adjusted coding vector to a preset standard numerical interval, where the preset standard numerical interval is a pre-set fixed numerical range.

[0179] The linear scaling process is to further adjust the element values ​​of the initially adjusted coding vector to within a preset standard value range so that the element values ​​have a uniform scale.

[0180] During linear scaling, a linear transformation formula is used to map the element values ​​of the preliminary adjusted code vector to a preset standard numerical interval. For example, if the preset standard numerical interval is [0, 1], an element value x in the preliminary adjusted code vector is mapped to the interval [0, 1] using the linear transformation formula.

[0181] Step S4148: Determine the coding vector after linear scaling as the regularized coding vector.

[0182] After completing the linear scaling process, the obtained encoding vector is determined as the regularized encoding vector, which is the final encoding feature vector.

[0183] Step S415: Execute the above processing for each vehicle feature in the vehicle feature set to obtain a set of encoded feature vectors.

[0184] Each vehicle feature in the vehicle feature set is processed according to the processing procedures of steps S411 to S414 to obtain a corresponding coded feature vector, and all coded feature vectors are combined into a coded feature vector set.

[0185] In practice, a loop is used to sequentially process each vehicle feature in a vehicle feature set. For example, for a vehicle feature set containing 100 vehicle features, each feature is encoded sequentially to obtain 100 encoded feature vectors, which form a set of encoded feature vectors. In the case of a highway toll booth, encoding each vehicle feature produces a set of encoded feature vectors that accurately represent the image feature information of the vehicle area.

[0186] Step S420: Input each coded feature vector in the coded feature vector set into a feature fusion module of a preset vehicle classification model, associate and integrate features of different dimensions in the coded feature vectors, and obtain a fused feature vector set.

[0187] The feature fusion module is a module in the pre-defined vehicle classification model. Its function is to correlate and integrate features of different dimensions in the encoded feature vectors, explore potential relationships between features, and improve the expressiveness of features. The fused feature vector set is a set of fused feature vectors obtained by processing all encoded feature vectors through the feature fusion module.

[0188] When performing feature fusion processing, the feature fusion module can use structures such as fully connected layers and convolutional layers to process the encoded feature vector. For example, when using a fully connected layer, each neuron in the fully connected layer is connected to each element in the encoded feature vector. The different dimensional features of the encoded feature vector are weighted and summed through a series of weight coefficients, thereby achieving the correlation and integration of features of different dimensions.

[0189] In the highway toll booth scenario, features of different dimensions may represent different vehicle attributes, such as length, height, and color distribution. By correlating and integrating these features through the feature fusion module, the fused feature vector can more comprehensively and accurately reflect the overall characteristics of the vehicle, providing stronger support for subsequent classification decisions. All encoded feature vectors are processed in this manner, ultimately resulting in a set of fused feature vectors.

[0190] Step S430: Input each fused feature vector in the fused feature vector set into the classification decision module of the preset vehicle classification model, determine the vehicle type probability distribution corresponding to the fused feature vector through classification operation, and obtain the vehicle type probability distribution set.

[0191] The classification decision module is the module used to make classification decisions within the pre-set vehicle classification model. It receives fused feature vectors as input and, through a series of classification operations, determines the vehicle type probability distribution corresponding to each fused feature vector. The vehicle type probability distribution represents the likelihood that each fused feature vector corresponds to a different vehicle type. For example, for a particular fused feature vector, the probability of it being a car is 0.7, the probability of it being a truck is 0.2, the probability of it being a bus is 0.1, and so on. The vehicle type probability distribution set is the set of vehicle type probability distributions corresponding to all fused feature vectors.

[0192] As an embodiment, step S430 inputs each fused feature vector in the fused feature vector set into a classification decision module of a preset vehicle classification model, determines the vehicle type probability distribution corresponding to the fused feature vector through a classification operation, and obtains a vehicle type probability distribution set. Specifically, the following steps S431 to S435 may be included: Step S431: Input the fused feature vector to the first decision layer of the classification decision module, perform linear transformation on the fused feature vector, and obtain the first decision vector by multiplying the decision weight matrix and the fused feature vector.

[0193] The first decision layer, the first layer of the classification decision module, performs a linear transformation on the input fused feature vector. The decision weight matrix is ​​a pre-defined matrix in the first decision layer. Its function is to weight the different dimensional features of the fused feature vector to extract more discriminative feature information.

[0194] When performing linear transformation processing, the fused feature vector is represented as a vector, and the decision weight matrix is ​​subjected to matrix multiplication operation with the vector.

[0195] Step S432: performing nonlinear activation processing on the first decision vector, performing nonlinear transformation on each element in the first decision vector through a preset decision activation function to obtain an activated decision vector.

[0196] Nonlinear activation processing is used to introduce nonlinear factors, increase the model's expressiveness, and enable the model to learn more complex feature relationships. Preset decision activation functions can include Sigmoid, ReLU, Softmax, etc.

[0197] During nonlinear activation processing, each element in the first decision vector is input into a preset decision activation function for calculation, resulting in a new value. These new values ​​are then combined into the activation decision vector. For example, if the ReLU function is used as the preset decision activation function, for an element x in the first decision vector, if x > 0, the corresponding element in the activation decision vector is x; if x ≤ 0, the corresponding element in the activation decision vector is 0. In the highway toll booth scenario, nonlinear activation processing enables the model to better capture the complex relationships between the characteristics of different vehicle types, improving classification accuracy.

[0198] Step S433: Input the activation decision vector into the second decision layer of the classification decision module, perform linear transformation on the activation decision vector, and obtain a second decision vector by multiplying another decision weight matrix with the activation decision vector. The dimension of the second decision vector is consistent with the preset number of vehicle types.

[0199] The second decision layer is the second layer of the classification decision module, which performs further linear transformation processing on the activation decision vector. The other decision weight matrix is ​​the weight matrix of the second decision layer, which is different from the decision weight matrix of the first decision layer.

[0200] During the linear transformation process, the activation decision vector is represented as a vector, and the other decision weight matrix is ​​multiplied by this vector. Because the dimension of the second decision vector matches the number of predefined vehicle types, assuming there are five predefined vehicle types (e.g., sedan, truck, bus, motorcycle, and van), the design of the other decision weight matrix results in a second decision vector with a length of 5.

[0201] Step S434: normalizing the second decision vector, converting the element values ​​of the second decision vector into probability values ​​whose sum is 1 through a normalization function, and obtaining a vehicle type probability distribution.

[0202] Normalization is performed to convert the element values ​​of the second decision vector into probability values ​​so that the sum of these probability values ​​is 1, in order to intuitively represent the likelihood of each vehicle type. A commonly used normalization function is the Softmax function.

[0203] During normalization, the Softmax function is used to calculate each element in the second decision vector. The calculation process of the Softmax function is as follows: for the i-th element in the second decision vector, it is first exponentially operated, and then the exponential value is divided by the sum of the exponential values ​​of all elements in the second decision vector. The result is the probability value of the vehicle type. For example, the second decision vector is [2,3,1,4,0]. For the first element 2, e² (e is a natural constant) is first calculated, and then e² is divided by e²+e³+e¹+e 4 +e 0 The sum of is used to get the probability value of the vehicle type corresponding to the element. After calculating all elements, the set of probability values ​​obtained is the vehicle type probability distribution.

[0204] Step S435: Execute the above processing on each fused feature vector in the fused feature vector set to obtain a vehicle type probability distribution set.

[0205] Each fused feature vector in the fused feature vector set is processed according to the processing procedures of steps S431-S434 to obtain the corresponding vehicle type probability distribution, and all vehicle type probability distributions are combined into a vehicle type probability distribution set.

[0206] In practice, a loop is used to process each fused feature vector in the fused feature vector set in sequence. For example, for a fused feature vector set containing 200 fused feature vectors, a classification operation is performed on each fused feature vector in sequence to obtain 200 vehicle type probability distributions, which constitute the vehicle type probability distribution set.

[0207] Step S440: Based on each vehicle type probability distribution in the vehicle type probability distribution set, the vehicle type with the largest probability value is selected as the predicted vehicle type of the corresponding vehicle image to obtain a vehicle classification result.

[0208] After obtaining the vehicle type probability distribution set, for each vehicle type probability distribution, find the vehicle type with the largest probability value, and use the vehicle type as the predicted vehicle type of the corresponding vehicle image.

[0209] In practice, each vehicle type probability distribution in the set is iterated over, the probability values ​​for each vehicle type are compared, and the vehicle type with the highest probability value is selected. For example, given a vehicle type probability distribution [0.1, 0.2, 0.05, 0.6, 0.05], the highest probability value is 0.6. Assuming the corresponding vehicle type is a sedan, the sedan is used as the predicted vehicle type for the corresponding vehicle image. The predicted vehicle types for all vehicle images are summed up to obtain the vehicle classification result.

[0210] In summary, the deep learning-based vehicle classification method described above, starting with acquiring a vehicle image set and proceeding through a series of steps, including vehicle region extraction, feature extraction, and classification processing, ultimately accurately classifies vehicle types. This method has important application value in scenarios such as highway toll booths. This method fully leverages the advantages of deep learning, improving the accuracy and reliability of vehicle classification through multi-level processing and feature extraction of vehicle images.

[0211] It is understandable that the various modules involved in the above-mentioned introduction of the embodiments of the present invention, such as the encoding module and the decision module, can be adaptively selected and implemented according to the processes involved in the introduction. For example, for the encoding layer, a linear transformation process is performed on the vehicle features, and the first encoding vector is obtained by multiplying the weight matrix and the vehicle features. In addition, when implementing the scheme of the present invention, those skilled in the art can supplement the details according to the common knowledge in the field. For example, according to the common knowledge in the field, normalization can be used to eliminate dimensional conflicts before feature fusion, interpolation can be used to eliminate dimensional differences, and thresholds can be reasonably set based on historical data, experience or business scenario requirements. The model can be trained based on a general model training method, and the number of layers in the model structure can be set based on actual needs, the activation function can be selected, etc. The present invention will no longer provide redundant introductions to overly detailed implementation processes.

[0212] See also Figure 2 , Figure 2 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 may be connected via a bus or other means. The processor 101 (also known as the Central Processing Unit (CPU)) is the computing and control core of the computer system, capable of parsing various instructions within the computer system and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), which can be used to send and receive data under the control of the processor 101. The communication interface 102 may also be used for data transmission and interaction within the computer system. The memory 103 is a storage device in the computer system for storing programs and data. It is understood that the memory 103 herein may include both the built-in memory of the computer system and, of course, the extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system, but this is not limited to this in the present invention.

[0213] In one embodiment, the processor 101 executes the vehicle type classification method based on deep learning provided in the above embodiment of the present invention by running the computer program in the memory 103.

Claims

1. A vehicle type classification method based on deep learning, characterized in that: The method comprises: Acquire a vehicle image set containing multiple vehicle types, wherein each vehicle image in the vehicle image set is an image containing a complete vehicle outline acquired by an image acquisition device; Performing image grayscale processing on each vehicle image in the vehicle image set to convert the color vehicle image into a grayscale vehicle image; performing noise removal processing on the grayscale vehicle image to reduce interfering pixels in the grayscale vehicle image to obtain a denoised grayscale vehicle image; Performing edge detection processing on the denoised grayscale vehicle image to identify edge lines of the vehicle outline and obtain a vehicle edge image; Determining a minimum circumscribed rectangular boundary of a vehicle area based on edge lines in the vehicle edge image, wherein the minimum circumscribed rectangular boundary is a minimum rectangular area that can completely contain all vehicle edge lines; Cutting out a corresponding image area from the vehicle image according to the minimum circumscribed rectangular boundary to obtain a vehicle area image; All the cropped vehicle area images are combined into the vehicle area image set, wherein each vehicle area image in the vehicle area image set only includes a main body of the vehicle; Performing feature extraction processing on the vehicle region image set to obtain a vehicle feature set, wherein each vehicle feature in the vehicle feature set corresponds to a feature representation of a vehicle region image; The vehicle feature set is classified and processed using a preset vehicle classification model to obtain a vehicle classification result, which is used to indicate the vehicle type corresponding to each vehicle image.

2. The method according to claim 1, characterized in that The performing feature extraction processing on the vehicle area image set to obtain a vehicle feature set includes: performing size unification processing on each vehicle area image in the vehicle area image set, and adjusting vehicle area images of different sizes into standard vehicle area images of a preset size; performing channel separation processing on the standard vehicle area image to separate multi-channel pixel information of the standard vehicle area image into a plurality of single-channel pixel matrices to obtain a set of single-channel pixel matrices; Performing convolution filtering on each single-channel pixel matrix in the single-channel pixel matrix set, and performing pixel value weighted calculation on the single-channel pixel matrix through a sliding window to obtain a plurality of initial feature matrices; Performing feature aggregation processing on the multiple initial feature matrices, superimposing and combining the initial feature matrices obtained from different sliding windows in a preset order to obtain an aggregated feature matrix; Performing dimensionality reduction processing on the aggregated feature matrix to convert the two-dimensional aggregated feature matrix into a one-dimensional feature vector to obtain vehicle features; All vehicle features are grouped into the vehicle feature set.

3. The method according to claim 1, characterized in that The classifying process of the vehicle feature set by using a preset vehicle classification model to obtain a vehicle classification result includes: Inputting each vehicle feature in the vehicle feature set into a feature encoding module of a preset vehicle classification model, converting the vehicle feature into an encoding feature vector of a fixed length, and obtaining an encoding feature vector set; Inputting each coded feature vector in the coded feature vector set into a feature fusion module of a preset vehicle classification model, correlating and integrating features of different dimensions in the coded feature vectors to obtain a fused feature vector set; Inputting each fused feature vector in the fused feature vector set into a classification decision module of a preset vehicle classification model, determining the vehicle type probability distribution corresponding to the fused feature vector through a classification operation, and obtaining a vehicle type probability distribution set; Based on each vehicle type probability distribution in the vehicle type probability distribution set, the vehicle type with the largest probability value is selected as the predicted vehicle type of the corresponding vehicle image to obtain a vehicle classification result.

4. The method according to claim 1, wherein The performing noise removal processing on the grayscale vehicle image to reduce interference pixels in the grayscale vehicle image to obtain a denoised grayscale vehicle image includes: Performing pixel value statistical processing on the grayscale vehicle image to calculate the pixel value mean and pixel value variance of all pixels in the grayscale vehicle image; Determine a noise judgment threshold based on the pixel value mean and the pixel value variance, wherein the noise judgment threshold is the pixel value mean plus or minus a preset multiple of the pixel value variance; Traversing each pixel in the grayscale vehicle image, marking the pixel whose pixel value exceeds the noise judgment threshold range as a noise pixel; Performing pixel value correction processing on the noise pixel point, replacing the pixel value of the noise pixel point with the average pixel value of the non-noise pixel points in a preset neighborhood around it, to obtain a corrected grayscale vehicle image; The corrected grayscale vehicle image is smoothed, and pixel values ​​in the corrected grayscale vehicle image are weighted averaged using a sliding average window to obtain a denoised grayscale vehicle image.

5. The method according to claim 1, characterized in that The performing edge detection processing on the denoised grayscale vehicle image to identify edge lines of the vehicle outline to obtain a vehicle edge image includes: Performing gradient calculation processing on the denoised grayscale vehicle image to calculate the pixel value change rates in the horizontal and vertical directions, respectively, to obtain a horizontal gradient image and a vertical gradient image; performing gradient amplitude calculation processing on the horizontal gradient image and the vertical gradient image, performing square and square root operations on the horizontal gradient values ​​and the vertical gradient values ​​at corresponding positions to obtain gradient amplitudes, thereby obtaining a gradient amplitude image; Performing gradient direction calculation processing on the gradient magnitude image, calculating the gradient direction angle of each pixel point by an inverse tangent function, and obtaining a gradient direction image; Performing non-maximum suppression on the gradient magnitude image, retaining pixels with local maximum gradient magnitudes in the gradient direction, and obtaining a suppressed gradient magnitude image; performing double-threshold processing on the suppressed gradient amplitude image, classifying pixels into strong edge points, weak edge points, and non-edge points based on a first threshold and a second threshold, wherein the strong edge points are pixels having a gradient amplitude greater than the first threshold, the weak edge points are pixels having a gradient amplitude between the second threshold and the first threshold, and the non-edge points are pixels having a gradient amplitude less than the second threshold; Performing edge connection processing on the weak edge points, determining the weak edge points connected to the strong edge points as edge points, and obtaining an edge point set; Based on the edge point set, corresponding pixel points are marked on the blank image to obtain a vehicle edge image.

6. The method according to claim 2, characterized in that The convolution filtering process is performed on each single-channel pixel matrix in the single-channel pixel matrix set, and pixel value weighted calculation is performed on the single-channel pixel matrix through a sliding window to obtain multiple initial feature matrices, including: For each single-channel pixel matrix in the single-channel pixel matrix set, initializing a plurality of convolution kernels of different sizes, each convolution kernel including a preset number of weight coefficients; Each convolution kernel is used as a sliding window to slide on the corresponding single-channel pixel matrix according to a preset step size. At each sliding position, the weight coefficient of the convolution kernel is multiplied and summed with the pixel value of the corresponding area of ​​the single-channel pixel matrix to obtain the eigenvalue at that sliding position; Arrange the eigenvalues ​​at all sliding positions in the sliding order to form an initial feature matrix related to the size of the single-channel pixel matrix; Perform the above convolution filtering process on each single-channel pixel matrix using convolution kernels of all different sizes to obtain multiple initial feature matrices; Multiple initial feature matrices corresponding to the same single-channel pixel matrix are grouped into an initial feature matrix group, and the multiple initial feature matrix groups corresponding to the single-channel pixel matrix set together constitute the multiple initial feature matrices.

7. The method according to claim 2, characterized in that The performing feature aggregation processing on the multiple initial feature matrices, superimposing and combining the initial feature matrices obtained from different sliding windows in a preset order to obtain an aggregated feature matrix, includes: Performing a size adjustment process on each of the multiple initial feature matrices, and unifying the initial feature matrices of different sizes into a standard initial feature matrix of the same size through an interpolation operation; Performing channel dimension expansion processing on the standard initial feature matrix, adding a channel dimension identifier to each standard initial feature matrix, and obtaining a standard initial feature matrix with a channel identifier; All standard initial feature matrices with channel identifiers are concatenated in the order of channel dimension identifiers to obtain a multi-channel feature matrix; Performing feature importance weighting processing on the multi-channel feature matrix, assigning a preset importance weight to the feature matrix of each channel, wherein the importance weight is determined according to the convolution kernel size corresponding to the feature matrix; The pixel values ​​of the weighted multi-channel feature matrix are added in the channel dimension to obtain a two-dimensional aggregate feature matrix.

8. The method according to claim 7, characterized in that The step of performing feature importance weighting processing on the multi-channel feature matrix and assigning a preset importance weight to the feature matrix of each channel includes: Obtain the size parameters of the convolution kernel corresponding to the feature matrix of each channel, wherein the size parameters include the height and width of the convolution kernel; Calculating the receptive field area of ​​the convolution kernel based on the height value and the width value of the convolution kernel, wherein the receptive field area is the product of the height value and the width value of the convolution kernel; Count the number of weight coefficients contained in each convolution kernel to obtain the number of convolution kernel parameters; The receptive field area is correlated with the number of convolution kernel parameters to obtain an initial weight coefficient, where the initial weight coefficient is the ratio of the receptive field area to the number of convolution kernel parameters; Calculating the pixel value variance of the feature matrix of each channel in the multi-channel feature matrix to obtain a feature matrix variance value; Multiplying the initial weight coefficient by the characteristic matrix variance value of the corresponding channel to obtain a dynamic weight coefficient; Normalizing the dynamic weight coefficients so that the sum of the dynamic weight coefficients of all channels is 1, thereby obtaining a normalized dynamic weight coefficient; The normalized dynamic weight coefficient is used as the importance weight of the feature matrix of each channel, and the corresponding importance weight is assigned to each channel feature matrix in the multi-channel feature matrix.

9. The method according to claim 3, characterized in that The step of inputting each vehicle feature in the vehicle feature set into a feature encoding module of a preset vehicle classification model, converting the vehicle feature into an encoding feature vector of a fixed length, and obtaining an encoding feature vector set includes: Inputting the vehicle features into the first encoding layer of the feature encoding module, performing linear transformation processing on the vehicle features, and obtaining a first encoding vector by multiplying the weight matrix and the vehicle features; Performing nonlinear activation processing on the first coding vector, performing nonlinear transformation on each element in the first coding vector using a preset activation function to obtain an activated coding vector; Inputting the activation coding vector into the second coding layer of the feature coding module, performing a linear transformation on the activation coding vector, and obtaining a second coding vector by multiplying the activation coding vector by another weight matrix; Regularizing the second coding vector to adjust element values ​​of the second coding vector to within a preset value range, thereby obtaining a regularized coding vector as a coding feature vector; Performing the above processing on each vehicle feature in the vehicle feature set to obtain a set of encoded feature vectors; The step of inputting each fused feature vector in the fused feature vector set into a classification decision module of a preset vehicle classification model, determining the vehicle type probability distribution corresponding to the fused feature vector through a classification operation, and obtaining a vehicle type probability distribution set includes: Inputting the fused feature vector into the first decision layer of the classification decision module, performing linear transformation on the fused feature vector, and obtaining a first decision vector by multiplying the decision weight matrix and the fused feature vector; Performing nonlinear activation processing on the first decision vector, performing nonlinear transformation on each element in the first decision vector using a preset decision activation function to obtain an activated decision vector; Inputting the activation decision vector into the second decision layer of the classification decision module, performing a linear transformation on the activation decision vector, and obtaining a second decision vector by multiplying the activation decision vector by another decision weight matrix, wherein the dimension of the second decision vector is consistent with the number of preset vehicle types; Normalizing the second decision vector, converting element values ​​of the second decision vector into probability values ​​whose sum is 1 using a normalization function, to obtain a vehicle type probability distribution; The above processing is performed on each fused feature vector in the fused feature vector set to obtain a vehicle type probability distribution set.

10. A computer system, characterized in that: include: a memory storing a computer program; A processor, configured to load the computer program to implement the vehicle type classification method based on deep learning as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Automatic vehicle detection method based on deep learning

    CN105184271A

  • Vehicle type classification method and classification device

    CN105224951A

  • Method and apparatus for identifying vehicle

    CN108171203A

  • Method for identifying axle type of green channel vehicle based on convolutional neural network

    CN110532946A

  • Vehicle classification method and device, computer equipment and storage medium

    CN116188832A