Lightweight vehicle identification method based on image
Through lightweight deep learning model and multi-label loss function optimization, the problem of vehicle identification model deployment problems and uneven categories of categories are solved, and efficient and accurate vehicle identification on low-cost equipment is achieved.
Patent Information
- Application Number
- CN202510393885.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing image-based vehicle recognition models are difficult to deploy to edge devices or low-cost servers due to the complex model structure and numerous parameters, and the recognition accuracy of a few categories is difficult to balance the learning efficiency of different vehicle categories, especially the low recognition accuracy of a few categories.
A lightweight deep learning model is adopted, including selection modules, location modules, channel modules and coordination modules, combined with multi-label loss function, and optimized through data augmentation and feature extraction optimization, and optimized for the uneven distribution of vehicle categories.
Real-time operation of vehicle recognition on low-cost equipment is achieved, which improves the accuracy of vehicle recognition, especially the recognition accuracy of a few categories, reduces hardware requirements, and enhances the generalization ability and recognition efficiency of the model.
Smart Images

Figure CN120259845A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a lightweight vehicle recognition method based on images. Background Art
[0002] Safety accidents caused by traffic congestion due to excessive traffic flow are also gradually increasing. A reliable and simple vehicle recognition method can help traffic management departments accurately and quickly identify vehicles, reduce labor costs, and reduce the response time of the operation and management departments.
[0003] Existing image-based vehicle recognition models have a complex model structure and numerous model parameters, resulting in slow model inference speed, poor real-time performance, and difficulty in being deployed to edge devices or low-cost servers; they have a high dependence on hardware performance and are difficult to adapt to terminal devices with different computing powers (such as camera edge computing nodes), restricting the expansion of actual application scenarios; the sample numbers of different vehicle categories (such as ordinary cars, trucks, and special vehicles) in traffic images vary significantly, and conventional loss functions are difficult to balance the learning efficiency of minority categories. The above problems urgently require a lightweight vehicle recognition method based on images. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a lightweight vehicle recognition method based on images, which is particularly suitable for a vehicle recognition method model that needs to be lightweight deployed to edge devices or in-vehicle servers.
[0005] The technical solution adopted by the present invention is as follows: In the first aspect, a lightweight vehicle recognition method based on images is provided, including:
[0006] Receiving a vehicle image;
[0007] Normalizing the size and channels of the vehicle image;
[0008] Using a trained deep learning model to label the bounding boxes and categories of all vehicles in the normalized vehicle image;
[0009] Converting the labeled vehicle image to the original image size and then outputting it.
[0010] Further, the training method of the trained deep learning model includes the following steps:
[0011] Collecting vehicle images and enhancing data by rotation, scaling, and cropping;
[0012] Labeling the bounding boxes and categories of all vehicles in the vehicle image;
[0013] Normalizing the size and channels of the vehicle image;
[0014] Dividing the vehicle images into a training set, a validation set, and a test set;
[0015] Input the training set and the validation set into the deep learning model for iterative loop;
[0016] Evaluate the performance of the deep learning model by the ACC metric for the validation set.
[0017] Furthermore, in the iterative loop process of the deep learning model, a multi-label loss function is adopted for optimization. The positive sample loss formula is L pos =-y i ·(1 - p i ) γ+ ·log(p i ), and the negative sample loss formula is where y i represents the positive data label in data i, p i represents the probability that i belongs to positive data in the recognition result, γ+ represents the adjustment parameter for positive data, (1 - y i ) represents the negative data label, p m represents the modified probability, p m = max(p i - m, 0), m represents the probability boundary, m >= 0, γ- represents the adjustment parameter for negative data.
[0018] Furthermore, the deep learning model is based on a convolutional neural network, including a selection module, a position module, a channel module, and a coordination module.
[0019] Furthermore, the working process of the selection module includes the following steps:
[0020] Perform convolution operations with convolution kernels of 3×3 and 5×5 on the vehicle image respectively to obtain corresponding feature maps and
[0021] For and perform an element-by-element summation operation to obtain the merged feature map U;
[0022] Perform global average pooling operation on the feature map U through the equation to obtain the feature map S, where H represents the length of the image, W represents the width of the image, and F gp represents the global average pooling operation;
[0023] Perform a fully connected layer operation on the feature map S through the equation Z = F fc (S)= δ(Ψ(WS)), where F fc represents the fully connected layer operation, δ represents the ReLU function, and Ψ represents the batch normalization operation;
[0024] Through the equation Perform soft attention calculations on the feature map Z for the feature maps U~ and respectively, and perform dot product operations on the results obtained with the feature maps and to obtain and
[0025] Then and are subjected to dot product operations through the equation to obtain the feature map V ∈ R C×H×W .
[0026] Furthermore, the working process of the position module includes the following steps:
[0027] Perform a Con convolution operation on the feature map V to obtain H, I ∈ R C×N , where N = H × W represents the number of pixels in the spatial range;
[0028] Through the equation transpose H and perform a dot product operation with I, and then calculate through the softmax function to obtain G ∈ R N×N , where g j,i represents the influence of the i-th position on the j-th position, and exp represents the exponential function;
[0029] Through the equation perform a dot product operation on the feature map V and the feature map G, multiply the result by a scale parameter α, and then perform an element-by-element summation operation with the feature map V to obtain the feature map E ∈ R C×H×W .
[0030] Furthermore, the working process of the channel module includes the following steps:
[0031] Vectorize the feature map E to obtain the feature map K ∈ R C×N , where N = H × W;
[0032] Through the equation multiply the feature map K by the transposed matrix of the feature map K and then obtain the feature map L ∈ R through the softmax function C×C , where l j,i represents the influence of the i-th position on the j-th position, and exp represents the exponential function;
[0033] Through the equation perform a dot product operation on the feature map E and the feature map L, multiply the result by a scale parameter β, and then perform an element-by-element summation operation with the feature map E to obtain the feature map M ∈ R C×H×W .
[0034] Further, the working process of the coordination module includes the following steps:
[0035] For the feature map M, through the equation the feature map P1 is obtained, where W represents encoding each channel vertically using a pooling kernel of size (1, W), and h is the image height;
[0036] For the feature map M, through the equation the feature map P2 is obtained, where H represents encoding each channel horizontally using a pooling kernel of size (1, H), and w is the image width;
[0037] Perform 1×1 convolution operations on the feature map P1 and the feature map P2 respectively to obtain the feature map O1 and the feature map O2;
[0038] Perform an element-wise multiplication operation on the feature map O1, the feature map O2, and the feature map M to obtain the feature map R;
[0039] Perform a Con convolution operation on the feature map R to obtain the vehicle recognition image.
[0040] The advantages and positive effects of the present invention are: Due to the above technical solution, the proposed lightweight vehicle recognition method significantly reduces the number of architecture parameters, can reduce the hardware requirements for running the model, and makes it easier to migrate to various edge devices with lower costs and poorer performance for real-time operation; By using the multi-label loss function, the problem of uneven vehicle categories in the image can be solved, the accuracy of vehicle recognition can be improved, especially the false negative rate of minority-class vehicles is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic flowchart of the lightweight vehicle recognition method based on images according to an embodiment of the present invention
[0042] Figure 2 is a schematic flowchart of the training method of the trained deep learning model according to an embodiment of the present invention
[0043] Figure 3 is a schematic architecture diagram of the module recognition method in the model according to an embodiment of the present invention
[0044] Figure 4 is a schematic diagram of the vehicle image with output annotation completed according to an embodiment of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which exemplary embodiments of the present disclosure are shown. The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0046] The present invention proposes a method for annotating the bounding boxes and categories of vehicles in vehicle images in an artificial intelligence manner. In the traditional manual method, due to the large amount of surveillance videos or pictures faced by the staff, they will quickly feel fatigued. Therefore, this recognition and annotation method is inefficient and inaccurate.
[0047] As Figure 1 shown, the present invention provides a lightweight vehicle recognition method based on images, including:
[0048] S100. Receive vehicle images;
[0049] S200. Standardize the size and channels of the vehicle images;
[0050] S300. Use the trained deep learning model to annotate the bounding boxes and categories of all vehicles in the standardized vehicle images;
[0051] S400. Convert the annotated vehicle images to the original image size and then output them.
[0052] By adopting the above method, by establishing a deep learning model and replacing the manual annotation method with an artificial intelligence method, the efficiency and accuracy of the recognition and annotation of vehicle images are improved; various vehicles in traffic video images can be accurately recognized, which can not only provide road condition vehicle information for autonomous driving, but also provide decision-making assistance for the operation and maintenance management department, and improve traffic management efficiency.
[0053] In order to solve the problems of the accuracy, robustness and generalization ability of the vehicle recognition model, an implementation manner is provided in this embodiment.
[0054] As Figure 2 shown, in one embodiment, the training method of the trained deep learning model includes the following steps:
[0055] S301. Collect vehicle images and enhance the data by rotation, scaling and cropping;
[0056] S302. Annotate the bounding boxes and categories of all vehicles in the vehicle images;
[0057] S303. Standardize the size and channels of the vehicle images;
[0058] S304. Divide the vehicle images into a training set, a validation set, and a test set;
[0059] S305. Input the training set and the validation set into the deep learning model for iterative loop;
[0060] S306. Evaluate the performance of the deep learning model by using the ACC metric for the validation set.
[0061] By using the above method, processing vehicle images through data augmentation techniques such as rotation, scaling, and cropping can effectively increase the quantity and diversity of training samples, which helps improve the generalization ability of the model and reduce the risk of overfitting; accurately annotating the bounding boxes and categories of all vehicles in the vehicle images can help the model more precisely learn the features of the target objects; standardizing the size and channels of the vehicle images to ensure the consistency of the data input into the model helps improve the stability and efficiency of model training; by inputting the training set and validation set data to repeatedly adjust the structure and parameters of the method, and using the multi-label loss function calculation method to identify the gap between the recognized value and the true value during each adjustment process to more directionally adjust the structure and parameters of the method.
[0062] To solve the problem that the sample quantity differences of different vehicle categories in vehicle images are significant and it is difficult for conventional loss functions to balance the learning efficiency of minority categories, an implementation method is provided in this embodiment.
[0063] In one embodiment, during the iterative loop process of the deep learning model, a multi-label loss function is used for optimization. The positive sample loss formula is L pos = -y i ·(1 - p i ) γ+ ·log(p i ), and the negative sample loss formula is where y i represents the positive data label in data i, p i represents the probability that i belongs to positive data in the recognition result, γ+ represents the adjustment parameter for positive data, (1 - y i ) represents the negative data label, p m represents the modified probability, p m = max(p i - m, 0), m represents the probability boundary, m >= 0, γ- represents the adjustment parameter for negative data. The values of γ+ and γ- are respectively set to preset values, such as [0, 2, 4, 6, 8, 10]. During the training and validation process of the model, the combination of γ+ and γ- with the optimal recognition effect is selected. m is a parameter that adjusts the influence of the model on various types of negative samples, and the range of m is [0 - 1].
[0064] By adopting the above method, a positive and negative sample weighted loss function for uneven vehicle category distribution is proposed to optimize the problem of uneven vehicle category distribution in the actual dataset. In particular, the miss detection rate of minority class vehicles (such as fire trucks) is significantly reduced.
[0065] To solve the problem of the large number of architecture parameters in traditional deep learning models, an implementation method is provided in this embodiment.
[0066] As Figure 3 shown, in one embodiment, the deep learning model is based on a convolutional neural network and includes a selection module, a position module, a channel module, and a coordination module.
[0067] By adopting the above method, a processing framework cascaded by multiple modules in sequence is used to replace the traditional complex network structure, greatly reducing the number of architecture parameters, lowering the hardware requirements for servers or edge devices, and enabling real-time operation on low-cost devices.
[0068] To solve the problem of the difficulty in extracting key target information from the original vehicle images through traditional models, an implementation method is provided in this embodiment.
[0069] As Figure 3 shown, in one embodiment, the working process of the selection module includes the following steps:
[0070] Perform convolution operations with convolution kernels of 3×3 and 5×5 on the vehicle image respectively to obtain corresponding feature maps and
[0071] For and perform an element-by-element summation operation to obtain the merged feature map U;
[0072] Perform global average pooling operation on the feature map U through the equation to obtain the feature map S, where H represents the length of the image, W represents the width of the image, and F gp represents the global average pooling operation;
[0073] Perform a fully connected layer operation on the feature map S through the equation Z = F fc (S) = δ(Ψ(WS)), where F fc represents the fully connected layer operation, δ represents the ReLU function, and Ψ represents the batch normalization operation;
[0074] Through the equation perform soft attention calculations on the feature map Z for the feature maps and respectively, and multiply the obtained results with the feature maps and Perform a dot product operation to obtain and
[0075] Let and Perform a dot product operation through the equation to obtain the feature map V ∈ R C×H×W . The initial values of a and b are respectively set to preset values, limited to [0 - 1], increasing from small to large, and the combination of a and b with the best recognition effect is selected as the actual soft attention coefficient.
[0076] Using the above method, through the processing of the vehicle image by the selection module, the extraction of useful information is enhanced during feature extraction, achieving the effect of suppressing image background noise and focusing on the features of the vehicle main body area.
[0077] To solve the problem that the representation ability of local features of context potential information in traditional models is weak, an implementation method is provided in this embodiment.
[0078] As Figure 3 shown, in one embodiment, the working process of the position module includes the following steps:
[0079] Perform a Con convolution operation on the feature map V to obtain H, I ∈ R C×N , where N = H × W represents the number of pixels in the spatial range;
[0080] Through the equation transpose H and perform a dot product operation with I, and then calculate through the softmax function to obtain G ∈ R N×N , where g j,i represents the influence of the i-th position on the j-th position, and exp represents the exponential function;
[0081] Through the equation perform a dot product operation on the feature map V and the feature map G, multiply the result by a scale parameter α, and then perform an element-by-element summation operation with the feature map V to obtain the feature map E ∈ R C×H×W .
[0082] Using the above method, through the processing of the feature map by the position module, the representation ability of local features of context potential information is improved.
[0083] To solve the problem that the recognition ability of the importance of each channel of the image in traditional models is poor, an implementation method is provided in this embodiment.
[0084] As Figure 3 shown, in one embodiment, the working process of the channel module includes the following steps:
[0085] Vectorize the feature map E to obtain the feature map K ∈ R C×N , where N = H × W;
[0086] Through the equation Multiply the feature map K by the transpose matrix of the feature map K and then pass it through the softmax function to obtain the feature map L ∈ R C×C , where l j,i represents the influence of the i-th position on the j-th position, and exp represents the exponential function;
[0087] Through the equation Perform a dot product operation on the feature map E and the feature map L, multiply the result by a scale parameter β, and then perform an element-wise summation operation with the feature map E to obtain the feature map M ∈ R C×H×W .
[0088] Using the above method, through the processing of the feature map by the channel module, the recognition ability of the importance of each channel of the image is improved, and greater weights are given to the important channels, focusing on the important feature information.
[0089] To solve the problem of poor recognition ability of vehicles in images in traditional models, an implementation method is provided in this embodiment.
[0090] As Figure 3 shown, in one embodiment, the working process of the coordination module includes the following steps:
[0091] For the feature map M, through the equation Obtain the feature map P1, where W represents encoding each channel vertically using a pooling kernel of size (1, W), and h is the image height;
[0092] For the feature map M, through the equation Obtain the feature map P2, where H represents encoding each channel horizontally using a pooling kernel of size (1, H), and w is the image width;
[0093] Perform 1×1 convolution operations on the feature map P1 and the feature map P2 respectively to obtain the feature maps O1 and O2;
[0094] Perform a dot product operation on the feature maps O1 and O2 and the feature map M to obtain the feature map R;
[0095] Perform a Con convolution operation on the feature map R to obtain the vehicle recognition image.
[0096] Using the above method, through the processing of the feature map by the coordination module, global pooling in the vertical and horizontal directions is achieved to encode position information, improving the recognition ability of the method for vehicles.
[0097] The following will describe the content involved in the above embodiments in conjunction with a preferred embodiment.
[0098] As Figures 1-4 shown, the vehicle in the traffic flow is photographed by a surveillance camera or a drone to obtain its image information. Data augmentation operations such as rotation, scaling, and cropping are performed on the obtained images. The resolution of all images is uniformly modified to 640*640, and the number of channels is modified to three RGB color channels. The vehicles in the images are manually labeled, and the ranges and types of all vehicles in the images are labeled respectively. The types include bus, car, van, etc. All images are randomly divided into 4130 training sets, 500 validation sets, and 500 test sets. A deep learning model based on a convolutional neural network is constructed, and the training set and the validation set are input into the model for training. In the iterative loop process, a multi-label loss function is used for optimization. By inputting the training set and the validation set data, the structure and parameters of the method are adjusted multiple times. In each adjustment process, the multi-label loss function is used to calculate the gap between the recognition value and the true value, so as to more directionally adjust the structure and parameters of the method. The termination condition of the iterative loop is two preset conditions. Once either condition is met, the iteration is terminated. Condition one is that the number of iterative loops reaches 500 times, and condition two is that the change rate of the loss function in 5 consecutive iterative loops does not exceed 0.01%. The test set is input into the model to evaluate the performance of the model. The accuracy index ACC is used to evaluate the recognition accuracy of the vehicles in the images. The formula is: ACC = (TP + TN) / (TP + TN + FN + FP). The calculated recognition accuracy ACC = 97.4%.
[0099] The above has described the embodiments of the present invention in detail, but the content described is only the preferred embodiments of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. An image-based lightweight vehicle recognition method, characterized in that, Including: Receiving vehicle images; Normalizing the size and channels of vehicle images; Using a trained deep learning model to label the bounding boxes and categories of all vehicles in the normalized vehicle images; Converting the labeled vehicle images to the original image size and then outputting them.
2. The image-based lightweight vehicle recognition method according to claim 1, wherein, The training method of the trained deep learning model includes the following steps: Collecting vehicle images and augmenting data through rotation, scaling, and cropping; Labeling the bounding boxes and categories of all vehicles in the vehicle images; Normalizing the size and channels of vehicle images; Dividing the vehicle images into a training set, a validation set, and a test set; Inputting the training set and the validation set into the deep learning model for iterative loops; Evaluating the performance of the deep learning model using the ACC metric for the validation set.
3. The image-based lightweight vehicle recognition method according to claim 2, characterized in that: During the iterative loop process of the deep learning model, a multi-label loss function is used for optimization. The positive sample loss formula is L pos =-y i ·(1 - p i ) γ+ ·log(p i ). The negative sample loss formula is where y i represents the positive data label in data i, p i represents the probability that i belongs to positive data in the recognition result, γ+ represents the adjustment parameter for positive data, (1 - y i ) represents the negative data label, p m represents the modified probability, p m = max(p i - m, 0), m represents the probability boundary, m >= 0, and γ- represents the adjustment parameter for negative data.
4. The image-based lightweight vehicle recognition method according to claim 1, characterized in that: The deep learning model is based on a convolutional neural network and includes a selection module, a position module, a channel module, and a coordination module.
5. The image-based lightweight vehicle recognition method according to claim 4, characterized in that The working process of the selection module includes the following steps: Perform convolution operations on the vehicle images with convolution kernels of 3×3 and 5×5 respectively to obtain the corresponding feature maps and For and perform an element-by-element summation operation to obtain the merged feature map U; Perform global average pooling operation on the feature map U through the equation to obtain the feature map S, where H represents the length of the image, W represents the width of the image, and F gp represents the global average pooling operation; Perform a fully connected layer operation on the feature map S through the equation Z = F fc (S) = δ(Ψ(WS)), where F fc represents the fully connected layer operation, δ represents the ReLU function, and Ψ represents the batch normalization operation; Through the equation Perform soft attention calculations on the feature map Z for the feature maps and respectively. For the results obtained, perform dot product operations with the feature maps and respectively to obtain and Multiply and through the equation to perform a dot product operation, obtaining the feature map V ∈ R C×H×W , where a and b are soft attention coefficients.
6. The image-based lightweight vehicle recognition method according to claim 5, characterized in that, The working process of the position module includes the following steps: Perform a Con convolution operation on the feature map V to obtain H, I ∈ R C×N , where N = H × W represents the number of pixels in the spatial range; Through the equation Perform a dot product operation on the transpose of H and I, and then calculate through the softmax function to obtain G ∈ R N×N , where g j,i represents the influence of the i-th position on the j-th position, and exp represents the exponential function; Through the equation Perform an element-wise multiplication operation on the feature map V and the feature map G, multiply the result by a scale parameter α, and then perform an element-wise summation operation with the feature map V to obtain the feature map E ∈ R C×H×W .
7. The image-based lightweight vehicle recognition method according to claim 6, wherein The working process of the channel module includes the following steps: Vectorize the feature map E to obtain the feature map K ∈ R C×N , where N = H × W; Through the equation Multiply the feature map K by the transposed matrix of the feature map K and then obtain the feature map L ∈ R through the softmax function C×C , where l j,i represents the influence of the i-th position on the j-th position, and exp represents the exponential function; Through the equation perform a dot product operation on the feature map E and the feature map L, multiply the result by a scale parameter β, and then perform an element-wise summation operation with the feature map E to obtain the feature map M ∈ R C×H×W .
8. The image-based lightweight vehicle recognition method according to claim 7, wherein, The working process of the coordination module includes the following steps: The feature map M is processed through the equation to obtain the feature map P1, where W represents encoding each channel vertically using a pooling kernel of size (1, W), and h is the image height; The feature map M is processed through the equation to obtain the feature map P2, where H represents encoding each channel horizontally using a pooling kernel of size (1, H), and w is the image width; Performing 1×1 convolutional operations on the feature map P1 and the feature map P2 respectively to obtain the feature map O1 and the feature map O2; Performing a dot product operation on the feature map O1, the feature map O2, and the feature map M to obtain the feature map R; Performing a Con convolutional operation on the feature map R to obtain a vehicle recognition image.
Citation Information
Patent Citations
Vehicle and pedestrian identification method based on improved YOLOv7
CN118230286A
YOLOv8 vehicle identification method, system and device and storage medium
CN118470494A
Lightweight double-dynamic vehicle detection and classification method and system in complex road environment
CN119579953A