Vehicle-mounted image recognition method
By adopting a combination method of frame-by-frame separation and CNN-Transformer model in the on-board image recognition system, the recognition challenges in complex traffic scenarios in the prior art are solved, and high-accuracy image recognition and classification are achieved, which improves the robustness and security of the system.
Patent Information
- Application Number
- CN202510035366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
AI Technical Summary
When existing vehicle image recognition technology deals with complex traffic scenarios, there are confrontational attacks, data privacy and ethical issues, and there is still room for improvement in identification and classification accuracy.
A frame-by-frame separation method based on vehicle camera images is adopted to generate a three-dimensional image data set, and a CNN-Transformer model is constructed to generate image recognition results through an optimization algorithm. This model combines the local feature extraction ability of CNN and the global feature exploration ability of Transformer, and generates feature representations suitable for classification through self-attention and multi-head self-attention mechanisms.
The high accuracy of on-board image recognition is achieved, and the classification accuracy reaches 98.35%, which significantly improves the recognition and classification performance of on-board images and enhances the robustness and security of the system.
Smart Images

Figure CN120014594A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image recognition, and in particular relates to a vehicle-mounted image recognition method. Background Art
[0002] As a core component of the development of autonomous vehicles, on-board image recognition technology relies on advanced artificial intelligence technology, especially the application of deep learning algorithms in image recognition. Detailed image data of the surrounding environment is obtained through cameras installed on the vehicle. This data is combined with three-dimensional spatial information from radar and lidar to provide the car with accurate environmental perception capabilities. This comprehensive perception system enables autonomous vehicles to recognize and understand important information such as road conditions, pedestrians, other vehicles, and traffic signs. In terms of deep learning models, convolutional neural networks (CNNs) are widely used for feature extraction and classification of images, while recurrent neural networks (RNNs) show their advantages in processing sequence data such as target tracking in video streams. In addition, autonomous vehicles need to recognize and understand not only static images, but also analyze video stream data to predict the movement trajectory of objects. For example, lane line detection and deviation warning functions rely on forward-looking cameras to monitor whether the vehicle deviates from the lane, while the recognition of pedestrians and other vehicles helps prevent possible collisions. The application of these technologies not only improves driving safety, but also enhances the ability of autonomous driving systems to cope with complex traffic scenarios. With the continuous advancement of technology, on-board image recognition also faces challenges such as adversarial attacks, data privacy, and ethical issues. Researchers are using emerging technologies such as federated learning, simulated learning, and explainable AI to solve these problems and improve the robustness and safety of the system. In the future, self-driving cars will provide a safe and efficient driving experience in a more complex and changing environment, while also finding a reasonable balance between protecting personal privacy and data utilization.
[0003] By integrating the above technologies, the on-board image recognition system can quickly and accurately detect and identify targets in real-time dynamic environments, including lane lines, pedestrians, vehicles, and traffic signs. The realization of these functions depends on the real-time processing and analysis of a large amount of dynamic data, involving advanced tasks such as behavior prediction and target tracking. For example, by analyzing the posture and movement trajectory of pedestrians, the system can predict their walking routes and adjust the driving strategy in time to avoid potential collisions. Similarly, by tracking the movement status of multiple targets, self-driving cars can make more reasonable driving decisions, thereby improving driving safety and efficiency. At the same time, the combination of real-time positioning and map update technology makes it possible for self-driving cars to locate and plan paths in unknown environments. Using lidar data and image processing results, the self-driving system can adapt to changing road conditions and environments while maintaining high accuracy. The development of this technology not only promotes the practical application of self-driving cars, but also promotes the evolution of transportation systems towards a more intelligent and automated direction, indicating that future transportation will be safer, more efficient and more environmentally friendly. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a vehicle-mounted image recognition method, comprising:
[0005] Based on the vehicle camera images, frame-by-frame separation is performed to generate a three-dimensional image dataset;
[0006] Based on the 3D image dataset, a CNN-Transformer model is constructed;
[0007] Based on the CNN-Transformer model, an optimization algorithm is used to generate image recognition results.
[0008] Furthermore, the CNN-Transformer model is constructed based on the three-dimensional image dataset, including:
[0009] Based on the 3D image dataset, CNN is used for feature encoding, class token and position encoding and dropout layer are introduced to generate processed feature data;
[0010] Based on the processed feature data, the Transformer’s self-attention and multi-head self-attention mechanisms and multi-layer neural perceptrons are used to generate feature representations suitable for classification.
[0011] Furthermore, based on the three-dimensional image data set, CNN is used for feature encoding, class token and position encoding and dropout layer are introduced to generate processed feature data, including:
[0012] Based on the three-dimensional image dataset, CNN is used for feature encoding to generate two-dimensional feature data;
[0013] Based on two-dimensional feature data, class token is introduced to generate feature aggregation structure;
[0014] Based on the feature aggregation structure, the position encoding is calculated and a dropout layer is added to generate feature data with position awareness and anti-overfitting mechanism.
[0015] Furthermore, based on the processed feature data, the self-attention and multi-head self-attention mechanisms of Transformer and multi-layer neural perceptron are used to generate feature representations suitable for classification, including:
[0016] Based on the self-attention mechanism, the position attention weight is calculated to capture long-distance dependencies;
[0017] Based on the multi-head self-attention mechanism, the self-attention score is calculated to generate multi-subspace features and position information learning mechanism;
[0018] Based on multi-subspace features, a multi-layer neural perceptron conversion mapping is used to generate feature representations suitable for classification.
[0019] Furthermore, the frame-by-frame separation based on the vehicle-mounted camera image to generate a three-dimensional image data set includes:
[0020] Using the on-board camera, the images are separated frame by frame in chronological order through programming in the MATLAB environment to obtain a series of RGB image data sets with three-dimensional structures.
[0021] Furthermore, the feature coding includes:
[0022] The input frame-by-frame images are processed using a convolution operation with an output channel number of 3, a convolution kernel size of 4×4, and a stride of 4, and the final output is two-dimensional feature data with a size dimension of 841×496.
[0023] Furthermore, the class token is a learnable tensor with a dimension of 1×U, where U is the number of output channels. The class token is randomly set and then continuously updated during the training process.
[0024] Furthermore, the Transformer is composed of a multi-layer multi-head self-attention module and a multi-layer neural perceptron module, and a Layernorm layer is used to connect each module, and the Layernorm layer normalizes the input of each layer.
[0025] Furthermore, the self-attention score is calculated based on the multi-head self-attention mechanism to generate multi-subspace features and position information learning mechanism including:
[0026] The multi-head self-attention mechanism divides the input into 8 parts, then calculates the self-attention score on each part in parallel, and finally concatenates the attention outputs of these 8 parts and obtains the final result by multiplying them with another trainable parameter matrix.
[0027] Furthermore, the optimization algorithm is adopted, including:
[0028] The Adam optimization algorithm was selected to adjust the parameters of the CNN-Transformer model, the learning rate was set to make the CNN-Transformer model converge stably during the training process, and the difference between the predicted results of the CNN-Transformer model and the true label was measured by minimizing the cross entropy loss.
[0029] Compared with the prior art, the present invention has the following advantages:
[0030] 1. The classification accuracy of the CNN-Transformer model reaches 98.35%, which is a significant improvement compared to existing advanced methods and can more accurately classify and identify vehicle images.
[0031] 2. Combining the advantages of CNN in extracting local features and the sensitivity of Transformer in exploring global features by relying on the attention mechanism, the local-global features of vehicle images are effectively mined and extracted.
[0032] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0034] Figure 1 A flow chart of a vehicle-mounted image recognition method provided in an embodiment of the present invention is shown;
[0035] Figure 2 A flowchart of S2 in a vehicle-mounted image recognition method provided in an embodiment of the present invention is shown;
[0036] Figure 3 A flowchart of S21 in a vehicle-mounted image recognition method provided in an embodiment of the present invention is shown;
[0037] Figure 4 A flowchart of S22 in a vehicle-mounted image recognition method provided in an embodiment of the present invention is shown;
[0038] Figure 5 A schematic diagram showing the architecture of a vehicle-mounted image recognition method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] See also Figure 1 , an embodiment of the present invention provides a vehicle-mounted image recognition method, comprising:
[0041] S1: Based on the vehicle camera images, frame-by-frame separation is performed to generate a three-dimensional image dataset.
[0042] Using the on-board camera, the images are separated frame by frame in chronological order. This process can be implemented by programming in the matlab environment, and a series of RGB image data sets with a three-dimensional structure (116×116×3) are obtained. Each such image becomes the basic unit for subsequent processing and analysis.
[0043] S2: Based on the 3D image dataset, build a CNN-Transformer model. Figure 2 .
[0044] S21: Based on the 3D image dataset, CNN is used for feature encoding, class token and position encoding and dropout layer are introduced to generate processed feature data. Figure 3 .
[0045] S211: Based on the three-dimensional image dataset, CNN is used for feature encoding to generate two-dimensional feature data.
[0046] Convolutional neural networks (CNN) use multiple convolutional layers to process input frame-by-frame images (such as See also Figure 5 ) for processing. The convolution layer uses the set convolution kernel to slide on the image and encode the features of the local area of the image.
[0047] CNN can effectively capture features in local areas of an image. The number of output channels in the convolution operation parameters is set to 3, the convolution kernel size is 4×4, and the step size is 4. After a series of convolution operations, the final output is two-dimensional data with a size of 841×496. This two-dimensional data is an abstract representation of the features of the original image, retaining the key feature information in the image.
[0048] S212: Based on the two-dimensional feature data, class token is introduced to generate a feature aggregation structure.
[0049] After CNN completes the feature encoding of the image features, the class token is introduced. The "class token" is borrowed from the idea of the Transformer architecture. In the natural language processing tasks of the Transformer architecture, the "class token", such as CLS (Classification Token) in BERT, is used to summarize the semantic information of the entire sequence to facilitate downstream tasks such as classification.
[0050] The introduced class token is a learnable tensor with a dimension of 1×U, where U is the number of output channels. Its initial value is randomly set and then continuously updated during the training process of the entire network. Its important role is to aggregate feature information from different positions and levels, so that the model can directly use the class token containing global feature information to make judgments when making predictions and classifications, thereby simplifying the classification process and improving efficiency.
[0051] S213: Based on the feature aggregation structure, calculate the position encoding and add a dropout layer to generate feature data with position awareness and anti-overfitting mechanism.
[0052] In order to make the CNN-Transformer model understand the relationship between different positions in the image feature vector, position encoding is introduced. The calculation formula of position encoding is:
[0053]
[0054] Among them, pos represents the position of the feature, d model represents the dimension of CNN output, and c represents the dimension of features.
[0055] Through calculation, the features at each position are assigned coded values. These coded values contain position information, which helps the model take position factors into account when processing features. At the same time, in order to prevent the CNN-Transformer model from overfitting and improve its generalization ability, a dropout layer with a dropout rate of 0.5 is added to the CNN-Transformer model. The dropout layer randomly discards the connections of some neurons during training, forcing the model to learn more robust feature representations.
[0056] S22: Based on the processed feature data, the Transformer self-attention and multi-head self-attention mechanisms and multi-layer neural perceptrons are used to generate feature representations suitable for classification. Figure 4 .
[0057] The attention neural network (Transformer) consists of multiple layers of multi-head self-attention modules and multi-layer neural perceptrons (MLP), and the Layernorm layer is used to connect each module. The Layernorm layer normalizes the input of each layer, which helps to accelerate the convergence speed of the model and improve the stability of the model.
[0058] S221: Based on the self-attention mechanism, calculate the position attention weight to capture long-distance dependencies.
[0059] Self-attention (SA) is an attention mechanism. In Transformer, the self-attention mechanism is used to calculate the attention weights of each position in the input vector in order to capture long-range dependencies in the input vector.
[0060] Specifically, the input vector Z is firstly transformed into a trainable parameter matrix W Q ,W K ,W V Multiply them together to get the query vector Q, key vector K and value vector V respectively.
[0061] Q=ZW Q ,K=ZW K ,V=ZW V
[0062] Then the attention weight of each position is determined by the following calculation:
[0063]
[0064] The softmax function normalizes the attention scores so that the sum of the attention weights of all positions is 1. k represents the dimension of the key vector, The role of is to properly normalize the attention score, so that the gradient is more stable during back propagation, which is helpful for model training.
[0065] S222: Based on the multi-head self-attention mechanism, calculate the self-attention score and generate multi-subspace feature and position information learning mechanism.
[0066] The multi-head self-attention (MSA) mechanism is a further extension of the self-attention mechanism. MSA divides the input into H = 8 parts, and then calculates the self-attention score on each part in parallel (i.e., the attention calculation under each head). Finally, the attention outputs of these 8 parts are spliced together and combined with another trainable parameter matrix W O Multiply to get the final result. The calculation formula of MSA is as follows:
[0067] head i =Attention(QW i Q ,KW i K ,VW i V )
[0068] MSA(Q,K,V)=Concat(head1,head2,…,head i )W O
[0069] The advantage of this multi-head structure is that it allows the CNN-Transformer model to learn features and position information simultaneously in different representation subspaces, and different heads can focus on different aspects of the image. In this way, the CNN-Transformer model can understand the information in the image more comprehensively and carefully, thereby improving the accuracy of recognition.
[0070] S223: Based on multi-subspace features, a multi-layer neural perceptron conversion mapping is used to generate feature representations suitable for classification.
[0071] The multi-layer neural perceptron (MLP) consists of two fully connected layers, with a Gaussian error linear layer (GELU) between the two layers. GELU is a nonlinear activation function that provides better expressiveness for the CNN-Transformer model than traditional activation functions. The role of MLP is to further transform and map the features processed by the multi-head self-attention module to obtain a feature representation that is more suitable for classification tasks. The overall calculation formula of Transformer can be summarized as:
[0072]
[0073] z l ′=MSA(LN(Z l-1 ))+Z l-1
[0074] Z l =MLP(LN(z l ′))+z l '
[0075]
[0076] in, Z0 represents all data sequences with position information, MSA represents the multi-head attention function, and LN represents the regularization operation.
[0077] S3: Based on the CNN-Transformer model, an optimization algorithm is used to generate image recognition results.
[0078] In Python 3.6 environment, the entire CNN-Transformer model was built using the deep learning framework PyTorch. During the training process, Adam was selected as the optimization algorithm to adjust the parameters of the model so that the model can better fit the training data. The learning rate was set to 0.001, which is an important parameter for controlling the update step of the model parameters. A smaller learning rate helps the model converge more stably during the training process. The model was trained for 500 iterations, and the parameters of the model were updated in each iteration to gradually improve the performance of the model. At the same time, 20 samples were processed each time. This batch processing method can improve training efficiency and reduce memory usage to a certain extent. The training goal of the model is to minimize the cross entropy loss, which measures the difference between the model prediction result and the true label. The model reduces this loss value by continuously adjusting parameters.
[0079] In order to accurately evaluate the performance of the model in the vehicle image recognition task, a five-fold cross-validation method is used. The specific operation is to randomly divide all the vehicle image data into five subsets. In each round of validation, one of the subsets is selected as the test set to test the generalization ability of the model, that is, the model's ability to predict unseen data; and the remaining four subsets are combined as the training set for training the model. This process is repeated five times, each time a different subset is selected as the test set, and finally the results of the five tests are summarized.
[0080] The embodiment of the present invention uses accuracy (ACC) as the evaluation index, and the calculation formula is as follows:
[0081]
[0082] Among them, TP (True positive) represents the number of correctly predicted positive samples (that is, the number of samples that are actually positive and predicted as positive by the model), TN (True negative) represents the number of correctly predicted negative samples (that is, the number of samples that are actually negative and predicted as negative by the model), FP (False positive) represents the number of false positive samples (that is, the number of samples that are actually negative but predicted as positive by the model), and FN (False negative) represents the number of false negative samples (that is, the number of samples that are actually positive but predicted as negative by the model).
[0083] Through this indicator, the prediction accuracy of the model on samples of different categories can be comprehensively evaluated, thereby judging the overall performance of the model. After evaluation, the classification accuracy of the model reached 98.35%, indicating that the model has high accuracy and reliability in the task of in-vehicle image recognition, see Table 1.
[0084] Table 1 Accuracy
[0085] Method ACC(%) MLP 95.31(1.39) LSTM 92.48(1.58) SVM 94.44(2.33) CNN-Transform 98.38(1.61)
[0086] An embodiment of the present invention provides a vehicle-mounted image recognition device, comprising:
[0087] Image separation module: Based on the vehicle camera image, frame-by-frame separation is performed to generate a three-dimensional image data set.
[0088] Model building module: Build a CNN-Transformer model based on a 3D image dataset.
[0089] The model building modules include CNN module and Transformer module:
[0090] CNN module: Based on the 3D image dataset, CNN is used for feature encoding, class token and position encoding and dropout layer are introduced to generate processed feature data.
[0091] The CNN module includes a feature encoding unit, a class token unit, and a position encoding unit:
[0092] Feature encoding unit: Based on the three-dimensional image data set, CNN is used for feature encoding to generate two-dimensional feature data;
[0093] Class token unit: Based on two-dimensional feature data, class token is introduced to generate feature aggregation structure;
[0094] Position encoding unit: Based on the feature aggregation structure, the position encoding is calculated and a dropout layer is added to generate feature data with position awareness and anti-overfitting mechanism.
[0095] Transformer module: Based on the processed feature data, the Transformer’s self-attention and multi-head self-attention mechanisms and multi-layer neural perceptrons are used to generate feature representations suitable for classification.
[0096] The Transformer module includes self-attention units, multi-head self-attention units, and multi-layer neural perceptron units:
[0097] Self-attention unit: Based on the self-attention mechanism, it calculates the position attention weight and captures long-distance dependencies;
[0098] Multi-head self-attention unit: Based on the multi-head self-attention mechanism, it calculates the self-attention score and generates multi-subspace features and position information learning mechanism;
[0099] Multi-layer neural perceptron unit: Based on multi-subspace features, a multi-layer neural perceptron conversion mapping is used to generate feature representations suitable for classification.
[0100] Model optimization module: Based on the CNN-Transformer model, it uses optimization algorithms to generate image recognition results.
[0101] An application embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the vehicle-mounted image recognition method.
[0102] An application embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the vehicle-mounted image recognition method.
[0103] An embodiment of the present invention further provides a computer program product corresponding to the vehicle-mounted image recognition method provided by the aforementioned embodiment. The computer program product includes a computer program, which is executed by a processor to implement the vehicle-mounted image recognition method provided by the aforementioned embodiment.
[0104] The above description and the accompanying drawings fully illustrate the embodiments of the present invention so that those skilled in the art can practice them. Other embodiments may include structural and other changes. The embodiments represent only possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The embodiments of the present invention are not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A vehicle-mounted image recognition method, characterized in that: include: Based on the vehicle camera images, frame-by-frame separation is performed to generate a three-dimensional image dataset; Based on the 3D image dataset, a CNN-Transformer model is constructed; Based on the CNN-Transformer model, an optimization algorithm is used to generate image recognition results.
2. The vehicle-mounted image recognition method according to claim 1, characterized in that: The CNN-Transformer model is constructed based on the three-dimensional image dataset, including: Based on the 3D image dataset, CNN is used for feature encoding, class token and position encoding and dropout layer are introduced to generate processed feature data; Based on the processed feature data, the Transformer’s self-attention and multi-head self-attention mechanisms and multi-layer neural perceptrons are used to generate feature representations suitable for classification.
3. The vehicle-mounted image recognition method according to claim 2, characterized in that: Based on the three-dimensional image dataset, CNN is used for feature encoding, class token and position encoding and dropout layer are introduced to generate processed feature data, including: Based on the three-dimensional image dataset, CNN is used for feature encoding to generate two-dimensional feature data; Based on two-dimensional feature data, class token is introduced to generate feature aggregation structure; Based on the feature aggregation structure, the position encoding is calculated and a dropout layer is added to generate feature data with position awareness and anti-overfitting mechanism.
4. The vehicle-mounted image recognition method according to claim 2, characterized in that: Based on the processed feature data, the self-attention and multi-head self-attention mechanisms of Transformer and multi-layer neural perceptron are used to generate feature representations suitable for classification, including: Based on the self-attention mechanism, the position attention weight is calculated to capture long-distance dependencies; Based on the multi-head self-attention mechanism, the self-attention score is calculated to generate multi-subspace features and position information learning mechanism; Based on multi-subspace features, a multi-layer neural perceptron conversion mapping is used to generate feature representations suitable for classification.
5. The vehicle-mounted image recognition method according to claim 1, characterized in that: The method of performing frame-by-frame separation based on the vehicle-mounted camera image to generate a three-dimensional image data set includes: Using the on-board camera, the images are separated frame by frame in chronological order through programming in the MATLAB environment to obtain a series of RGB image data sets with three-dimensional structures.
6. The vehicle-mounted image recognition method according to claim 3, characterized in that: The feature coding includes: The input frame-by-frame images are processed using a convolution operation with an output channel number of 3, a convolution kernel size of 4×4, and a stride of 4, and the final output is two-dimensional feature data with a size dimension of 841×496.
7. The vehicle-mounted image recognition method according to claim 3, characterized in that: The class token is a learnable tensor with a dimension of 1×U, where U is the number of output channels. The class token is randomly set and then continuously updated during the training process.
8. The vehicle-mounted image recognition method according to claim 2, characterized in that: The Transformer consists of a multi-layer multi-head self-attention module and a multi-layer neural perceptron module, and a Layernorm layer is used to connect each module, and the Layernorm layer normalizes the input of each layer.
9. The vehicle-mounted image recognition method according to claim 4, characterized in that: The multi-head self-attention mechanism is based on which the self-attention score is calculated and the multi-subspace feature and position information learning mechanism is generated. The multi-head self-attention mechanism divides the input into 8 parts, then calculates the self-attention score on each part in parallel, and finally concatenates the attention outputs of these 8 parts and obtains the final result by multiplying them with another trainable parameter matrix.
10. The vehicle-mounted image recognition method according to claim 1, characterized in that: The optimization algorithm adopted includes: The Adam optimization algorithm was selected to adjust the parameters of the CNN-Transformer model, the learning rate was set to make the CNN-Transformer model converge stably during the training process, and the difference between the predicted results of the CNN-Transformer model and the true label was measured by minimizing the cross entropy loss.