A road traffic sign recognition method based on image semantic understanding
By employing an image-based semantic understanding approach, utilizing the Blip network and the multimodal Transformer model, the problems of low recognition rate and poor environmental adaptability of traditional traffic sign recognition methods are solved, achieving efficient and accurate traffic sign recognition and description.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGAN UNIV
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing road traffic sign recognition methods based on traditional image recognition technology suffer from low recognition rates and poor environmental adaptability.
We employ an image semantic understanding-based approach, which involves image acquisition, preprocessing, and dataset construction. We then use the Blip network to build a road traffic sign detection model, which is trained by combining a visual encoder, a text encoder, and a decoder to achieve the recognition of road traffic signs.
It can efficiently and accurately identify road traffic signs in complex environments, generate sentence descriptions, avoid dependence on already marked signs, and improve the practical value of identification.
Smart Images

Figure BDA0005132309040000052 
Figure BDA0005132309040000062 
Figure BDA0005132309040000071
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and intelligent transportation technology, and in particular to a method for road traffic sign recognition based on image semantic understanding. Background Technology
[0002] With the development of intelligent transportation systems, autonomous driving technology has become a crucial trend in future transportation. Road traffic sign recognition is a key component of autonomous driving, providing vehicles with vital data such as traffic rules and instructions. Currently, traffic sign recognition methods based on traditional image recognition technology suffer from low recognition rates and poor environmental adaptability. Therefore, designing a road traffic sign recognition method based on image semantic understanding is essential. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a road traffic sign recognition method based on image semantic understanding.
[0004] To achieve the above objectives, the present invention provides the following solution:
[0005] This invention provides a road traffic sign recognition method based on image semantic understanding, comprising:
[0006] Image acquisition is performed using image acquisition equipment;
[0007] The collected images are preprocessed to construct a road traffic sign dataset;
[0008] A road traffic sign detection model was constructed based on the Blip network, and the model was trained on a road traffic sign dataset to obtain the trained road traffic sign detection model.
[0009] The effectiveness of the trained road traffic sign detection model was tested based on a road traffic sign dataset.
[0010] Preferably, image acquisition is performed using an image acquisition device, specifically as follows:
[0011] Based on camera equipment installed on the vehicle, image data of the road scene in front of the vehicle is acquired in real time. The camera equipment includes high-definition cameras and night vision cameras, which are used to switch according to different environmental requirements.
[0012] Preferably, the acquired images are preprocessed to construct a road traffic sign dataset, specifically as follows:
[0013] Acquire the collected image data, resize the images, and standardize the image dimensions.
[0014] Denoise removal is performed on the image data after it has been resized.
[0015] Perform color correction on the denoised image data;
[0016] Perform data augmentation processing on the color-corrected image data;
[0017] Manually annotate the image data after data augmentation.
[0018] The labeled image data is aggregated to generate a road traffic sign dataset.
[0019] Preferably, the image data after size unification is denoised by Gaussian filtering or median filtering to remove noise from the image.
[0020] Preferably, color correction is performed on the denoised image data, specifically as follows:
[0021] The image data after denoising is acquired and inspected. For images with uneven lighting, white balance adjustment and color correction are performed.
[0022] Preferably, the color-corrected image data undergoes data enhancement processing, specifically as follows:
[0023] Image enhancement is performed on color-corrected image data based on an image enhancement model. The image enhancement model includes a separate feature processing module, a joint feature processing module, and a feature fusion module. The separate feature processing module is used to learn five private feature maps, the joint feature processing module is used to learn a public feature map, and the feature fusion module is used to uniformly represent the private and public feature maps, ultimately achieving effective image enhancement.
[0024] Preferably, a road traffic sign detection model is constructed based on the Blip network, specifically as follows:
[0025] A road traffic sign detection model is constructed, comprising a visual encoder, a text encoder, a visual-text encoder, and a visual-text decoder. The visual encoder uses a VisionTransformer architecture and a Cross-Attention mechanism to extract feature information of road signs from images, which is then used as one of the joint representations. The text encoder uses a BERT architecture to extract text features. The visual-text encoder employs a Cross-Attention mechanism, with the attention part using a bidirectional Self-Attention mechanism, introducing an additional [Encode] token for binary classification using image and text features. The visual-text decoder also employs a Cross-Attention mechanism, with the attention part using Casual-Attention, introducing an additional [DEcode] token and a termination token to extract text features of road traffic signs. The road traffic sign detection model combines the image features extracted by the visual encoder and the text features extracted by the visual-text decoder through a cross-attention mechanism to form a new feature representation containing multimodal information of road signs.
[0026] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0027] This invention provides a road traffic sign recognition method based on image semantic understanding. The method includes image acquisition using an image acquisition device, preprocessing the acquired images to construct a road traffic sign dataset, building a road traffic sign detection model based on the Blip network, training the model on the road traffic sign dataset to obtain the trained road traffic sign detection model, and testing the performance of the trained model on the road traffic sign dataset. This invention can recognize road traffic signs and generate descriptive statements based on color, shape, and composition, identifying lane types from a macroscopic perspective and avoiding reliance on pre-labeled signs. It can achieve efficient and accurate traffic sign recognition in complex environments and has high practical value. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 A flowchart of the method provided in an embodiment of the present invention;
[0030] Figure 2 This is a schematic diagram of the Blip network structure;
[0031] Figure 3 Generate a schematic diagram of the task network structure for Blip;
[0032] Figure 4 This is a schematic diagram of the Vision Transformer structure;
[0033] Figure 5 This is a schematic diagram of the multi-head attention mechanism structure;
[0034] Figure 6 This is a schematic diagram of the cross-attention mechanism.
[0035] Figure 7 This is a schematic diagram of a dataset example;
[0036] Figure 8 A diagram illustrating the data format for text annotation;
[0037] Figure 9A Example A schematic diagram of the image description generated for the model;
[0038] Figure 9B Example B is a schematic diagram of the image description generated for the model;
[0039] Figure 9C Example C is a schematic diagram of the image description generated for the model;
[0040] Figure 10A Schematic diagram of example I of the incomplete description generated for the model;
[0041] Figure 10B Example II: A schematic diagram of an incomplete description generated for the model. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] The purpose of this invention is to provide a road traffic sign recognition method based on image semantic understanding, which can realize road traffic sign recognition based on image semantic understanding and is easy to use.
[0044] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0045] Figure 1 The method flowchart provided in the embodiments of the present invention is as follows: Figure 1 As shown, this invention provides a road traffic sign recognition method based on image semantic understanding, including:
[0046] Step 100: Acquire images using the image acquisition device;
[0047] Step 200: Preprocess the acquired images to construct a road traffic sign dataset;
[0048] Step 300: Construct a road traffic sign detection model based on the Blip network, and train the model based on the road traffic sign dataset to obtain the trained road traffic sign detection model;
[0049] Step 400: Test the performance of the trained road traffic sign detection model based on the road traffic sign dataset.
[0050] In step 100, image acquisition is performed using an image acquisition device, specifically as follows:
[0051] Based on camera equipment installed on the vehicle, image data of the road scene in front of the vehicle is acquired in real time. The camera equipment includes high-definition cameras and night vision cameras, which are used to switch according to different environmental requirements.
[0052] In step 200, the acquired images are preprocessed to construct a road traffic sign dataset, specifically as follows:
[0053] Acquire the collected image data, adjust its image size, and standardize the image size, for example, to 224×224 or larger, in order to unify the model input format;
[0054] Denoising is performed on the image data after it has been resized, for example by using Gaussian filtering or median filtering, to remove noise from the image and ensure image clarity.
[0055] Color correction is performed on the denoised image data. For images with uneven lighting, white balance adjustment and color correction are performed to ensure that the color information accurately reflects the real scene.
[0056] Perform data augmentation processing on the color-corrected image data;
[0057] Manually annotate the image data after data augmentation.
[0058] The labeled image data is aggregated to generate a road traffic sign dataset.
[0059] Data augmentation processing is performed on the color-corrected image data, specifically as follows:
[0060] Image enhancement is performed on color-corrected image data based on an image enhancement model. This model includes a separate feature processing module, a joint feature processing module, and a feature fusion module. The separate feature processing module learns five private feature maps, the joint feature processing module learns a common feature map, and the feature fusion module uniformly represents the private and common feature maps, ultimately achieving effective image enhancement. Each module is described in detail below:
[0061] Separate feature processing module:
[0062] First, feature vectors of a given image and its I variants are extracted, and then they are processed separately. A separate feature processing module receives multiple input visual signals, i.e., given image y... c [m, n] and its variant z i c [m, n], i = 1, ..., I, 0 ≤ m ≤ M-1, 0 ≤ n ≤ N-1, where c ∈ {1, 2, 3} represents the color channels of the image, M and N represent the height and width of the image respectively, and I = 4 is taken as the image height. i c [m, n], i = 1, 2, 3, 4, image enhancement based on a physical model is implemented, and the process of generating visual signals is defined as follows:
[0063]
[0064] In the formula, f1(), i = 1, 2, 3, 4 represents the i-th physical model-based image enhancement scheme. This application uses four classic methods, namely RoWS, MIP, UDCP and ULAP.
[0065] After acquiring all visual signals, a separate feature processing module extracts features from these signals. and the original image y c Features are extracted from [m, n]. This single feature processing module employs 5 parallel branches in its network architecture, each branch is designated to process one of the prior visual signals. Each branch uses a U-Net architecture with mirror skip connections. Each U-Net subnetwork is formed by a cascade of 5 strided convolutional operations and 5 transposed convolutional operations, each followed by a batch normalization operation and a LeakyReLU activation function.
[0066] Each straddle convolution and transpose convolution uses a straddle value of 2 and a kernel size of 4×4. The number of filters in the first and second straddle convolution operations is set to 32 and 128, respectively, while the remaining three straddle convolution operations use 256 filters. The first and second transpose convolution operations use 256 filters, and the remaining three transpose convolution operations use 128, 32, and 32 filters, respectively. This is achieved by representing the operation of each U-Net subnetwork in the five branches as g. i (), i = 1, 2, 3, 4, 5, we can obtain:
[0067]
[0068] u5[m,n]=g i (y c [m, n])
[0069] Among them, the feature tensor u i [m, n], i = 1, 2, 3, 4, 5 shows the outputs of the 5 U-Net branches. Finally, in a separate feature processing module, a convolution operation is performed once for each feature tensor, outputting the feature tensor:
[0070] v i [m, n] = W i (u i [m, n]), i=1,2,3,4,5
[0071] In the formula, W i (), i = 1, 2, 3, 4, 5 represents a convolution operation, which uses three filters with a kernel size of 4×4, followed by a sigmoid activation function.
[0072] Joint Feature Processing Module: Each of the five branches of the individual feature processing module acquires a specific set of information for recovering high-quality images, due to the varying input visual signals. i=1,2,3,4 and y c[m, n] each has unique characteristics. The feature tensors obtained by individual feature processing modules contain various different sets of information. Fusing them in a hierarchical manner can further enhance network performance. In the joint feature processing module, the feature tensors obtained after the first transposed convolution operation in the five parallel branches are initially connected together along the channel dimension, and the resulting tensor is input into another transposed convolution operation with a kernel size of 4×4 and a stride of 2. The above process is repeated four times on the feature tensors at the output of the second, third, fourth and fifth transposed convolution operations in the five branches. It should be noted that the number of filters used in the transposed convolution operation in the joint feature processing module are 512, 256, 64, 32 and 3, respectively. The last convolution operation is followed by a sigmoid activation function. The output feature tensor obtained by the joint feature processing module is called v6[m, n].
[0073] Feature fusion module: In the feature fusion module, the feature tensors V[m,n], i=1,2,3,4,5 obtained from the individual feature processing modules and the feature tensors v6[m,n] obtained from the joint feature processing module are fused together to produce the final estimated high-quality image;
[0074] Specifically, these six feature tensors are concatenated along the channel dimension and convolved using three 3×3 kernel filters to produce an enhanced, high-quality image x. c [m, n];
[0075] The proposed model can be trained using supervised learning, employing the C1 norm loss function:
[0076]
[0077] In the formula, T represents the total number of images in the batch, and ⊙ represents the entire parameter set in the network;
[0078] The training process for this model can be carried out using conventional training methods, which can be implemented using PyTorch programming.
[0079] After the above image enhancement methods, road traffic sign images can be enhanced to obtain an additional batch of road traffic sign images. Other operations, such as rotation and cropping, can also be performed on the road traffic sign images to finally obtain the dataset. The significance of this operation is that when it is impossible to collect a large number of road traffic sign images, a batch of artificially created road traffic sign images can be obtained through image enhancement operations and operations such as rotation and cropping, which can improve the training accuracy of the detection model.
[0080] In step 300, a road traffic sign detection model is constructed based on the Blip network, specifically as follows:
[0081] A road traffic sign detection model is constructed, comprising a visual encoder, a text encoder, a visual-text encoder, and a visual-text decoder. The visual encoder uses a VisionTransformer architecture and a Cross-Attention mechanism to extract feature information of road signs from images, which is then used as one of the joint representations. The text encoder uses a BERT architecture to extract text features. The visual-text encoder employs a Cross-Attention mechanism, with the attention part using a bidirectional Self-Attention mechanism, introducing an additional [Encode] token for binary classification using image and text features. The visual-text decoder also employs a Cross-Attention mechanism, with the attention part using Casual-Attention, introducing an additional [DEcode] token and a termination token to extract text features of road traffic signs. The road traffic sign detection model combines the image features extracted by the visual encoder and the text features extracted by the visual-text decoder through a cross-attention mechanism to form a new feature representation containing multimodal information of road signs.
[0082] First, the Blip model is analyzed. This invention uses the downstream task Image Caption of Blip to generate text descriptions for road traffic sign images. The visual encoder of this network uses the ViT architecture, which mainly utilizes the image features extracted by ViT and the text input of the visual text decoder to jointly perform the text generation task.
[0083] Blip (Bootstrapping Language-Image Pre-training) is a multimodal Transformer model for unified visual understanding and generation. Blip is a novel VLP (Vision-Language Pre-training) framework that enables a wider range of downstream tasks compared to existing methods. Figure 2 The image shows the Blip model structure MED (Multimodal mixture of Encoder-Decoder). The analysis of the Blip model structure is as follows.
[0084] 1. Visual encoder, in Figure 2The leftmost image shows Blip's visual encoder, which uses the VisionTransformer architecture. This architecture segments the input image into patches and encodes them into a series of image embeddings, using [CLS] tokens to represent global image features. Research on ViLT and SimVLM has demonstrated that ViT is computationally more efficient, therefore Blip's visual encoder adopted the ViT architecture.
[0085] 2. Text encoder, in Figure 2 In the diagram, the second column represents the text encoder, which uses the BERT architecture. The key feature of this structure is the addition of a [CLS] token to the beginning of the text input to summarize the entire sentence. This design aims to extract text features for comparative learning.
[0086] 3. Visual Text Encoder: The third column represents the visual text encoder, employing a Cross-Attention mechanism. This structure aims to perform binary classification tasks using image features from ViT and text input. Therefore, an encoder architecture is adopted, and the attention part uses a bidirectional Self-Attention mechanism. Furthermore, an additional [Encode] token is introduced to represent the joint representation of the image and text.
[0087] 4. Visual text decoder Figure 2 The fourth column of the structure is the visual text decoder, which employs a Cross-Attention mechanism. The main function of this part is to perform a text generation task based on image features from ViT and the text input. Therefore, this structure belongs to the decoder type, and its attention mechanism employs Casual-Attention to predict the next token. In addition, to achieve the text generation task, extra [Decode] tokens and end tokens are introduced, serving as the start and end points of the generated result, respectively.
[0088] During pre-training, BLIP jointly optimizes three objective functions: the Image-Text Contrastive Loss (ITC), the Image-Text Matching Loss (ITM), and the Language Modeling Loss (LM). ITC and ITM are objective functions for the understanding task, while LM is an objective function for the generation task. The Image-Text Contrastive Loss (ITC) promotes similarity between positive and negative image-text pair representations between a visual encoder and a text encoder, thereby aligning the visual and text feature spaces. The Image-Text Matching Loss (ITM) focuses on learning multimodal image-text representations to capture fine-grained alignment between visual and linguistic elements. ITM operates on a visual encoder and a visual-text encoder to perform a binary classification task, using a linear layer to predict whether image-text pairs match based on multimodal features. BLIP includes a decoder for the generation task. The Language Modeling Loss (LM), operating on a visual encoder and a visual-text decoder, aims to generate a textual description of a given image. The purpose of this invention is to use a language model objective function to generate a text description of a road traffic sign by providing the model with an image of the sign.
[0089] The image captioning of road traffic sign images in this invention is one of the downstream tasks of Blip, and its specific network structure is as follows: Figure 3 As shown, this model utilizes a Vision Transformer as an image feature encoder to extract feature information of road signs in an image, serving as one of the joint representations of the image and text. Then, a visual text decoder is used to extract the text features of the road signs, serving as another joint representation of the image and text. Finally, the image features extracted by Vision Transformer and the text features extracted by the visual text decoder are combined using a cross-attention mechanism to form a new feature representation containing rich multimodal information about the road signs.
[0090] The feature extraction process for road sign images is as follows: The Vision Transformer model structure mainly includes the following components: image serialization, linear transformation, position encoding, Transformer encoder, and fully connected network. The ViT model structure is as follows: Figure 4 As shown;
[0091] Taking a 224×224 input image as an example, the training process can be summarized as follows: First, the image is segmented into blocks in the Embedding layer. Next, each image block is linearly projected into a d-dimensional feature vector using a trained weight matrix. Then, these feature vectors are fed into the Transformer's encoding module, which contains a learned classification vector and a location vector. The core of this encoding module is a multi-head attention mechanism, which effectively maps the information of the original image to multiple spatial scales, thereby achieving comprehensive capture and effective encoding of image information. Finally, a one-dimensional vector representing the classification is extracted from the encoding module and fed into the MLP Head for further processing to obtain the final classification result.
[0092] Vision Transformer employs a scaling dot product attention mechanism, which enables it to achieve remarkable results in the field of computer vision. This mechanism differs from the feature extraction methods of traditional convolutional neural networks.
[0093] The calculation formula is shown below;
[0094]
[0095] In the formula, Attention is the final attention score, and Q, K, and V are Queries(Q), Key(K), and Values(V), respectively. This represents the dimension of matrix Q; the formulas for calculating Q, K, and V are shown below. Where W... q W k and W v This represents the weight matrix.
[0096] Q = W q X
[0097] V = W v X
[0098] K = W k X
[0099] The computation process of the scaled dot product attention mechanism can be divided into the following steps:
[0100] 1. The input image X first undergoes a linear transformation to obtain the Query matrix (Q), Key matrix (K), and Value matrix (V), all of which have the same dimensions.
[0101] 2. Calculate the dot product of the Q matrix and the K matrix as shown in the above formula to obtain the score matrix.
[0102] 3. Scale the score matrix by dividing the dot product of Q and K by... The purpose is to prevent the dot product from becoming too large when the dimension is large, thus pushing the softmax function into the region of minimum gradient and affecting the training effect of the model.
[0103] 4. Finally, multiply the results by V to obtain the output tensor Z.
[0104] To improve the efficiency of attention mechanisms in capturing feature information, multi-head attention is widely used. For example... Figure 5 As shown, compared to single-head attention mechanisms, multi-head attention mechanisms allow for multiple independent learning processes on the same set of information, enabling multi-faceted information mining. These diverse outputs are then concatenated and fused through a fully connected neural network to produce the final output. This design significantly enhances the attention mechanism's ability to learn features from data, enabling the model to collect and integrate information more comprehensively, thereby improving model performance and accuracy.
[0105] The image and text feature fusion process is as follows: The visual encoder of this invention uses a cross-attention (CA) mechanism, which performs image-to-text generation based on image features extracted by ViT and text input. The cross-attention mechanism has the ability to capture the correlation information between different input sequences, which is crucial for the model to understand their semantic relationships. In natural language processing tasks, the cross-attention mechanism is widely used in machine translation, reading comprehension, and other scenarios, where the source language sentence and the target language sentence are used as two separate input sequences. Cross-attention also plays an important role in computer vision tasks, especially in processing the correlation between images and text, such as image caption generation and image question answering. Through the cross-attention mechanism, the model can more accurately grasp the intrinsic relationship between images and text, thereby improving the performance of related tasks. The basic principle of the cross-attention mechanism used in this invention is as follows: Figure 6 As shown.
[0106] The core idea of this method is to first define one input as a query (Q) and the other inputs as a key (K) and a value (V). Then, an attention mechanism is used to calculate the correlation between these two inputs. The calculated attention weights are multiplied by the value (V) and summed to obtain joint interaction characteristics. Finally, these interaction features are combined with the original input to construct a new feature representation containing rich multimodal information. This method successfully overcomes the limitations of single-modality in image description tasks, greatly improving the model's versatility and robustness. Cross-attention, as a special form of multi-head attention, is based on dividing the input tensor into two distinct parts, H... x and H yIn this process, one part is designated as the query set, while the other part serves as the key-value set. This configuration allows the cross-attention mechanism to efficiently capture the relationships between the different parts, thereby improving the model's performance when processing complex data.
[0107] The present invention sets the evaluation criteria as follows: using five evaluation indicators commonly used in the field of image description, mainly including BLEU, METEOR, CIDEr, ROUGE-L and SPICE. These evaluation methods will be introduced below.
[0108] BLEU (Bilingual Evaluation Understudy) is a widely used machine translation quality assessment standard. Its core lies in evaluating the degree of overlap between the translated text and the standard text's N-grams by calculating the BLEU-N score, thereby measuring the accuracy and fluency of the translation. Different BLEU-N values reflect different aspects of the translation; BLEU-1 focuses on lexical integrity, while BLEU-2 to BLEU-4 progressively emphasize sentence coherence and fluency. Before calculating BLEU-N, the precision Pn of the N-grams must be determined, as shown in the following expression.
[0109]
[0110] In the formula, c i This represents the i-th generated statement, s ij Let m represent the j-th truth statement of the i-th generated statement; let w be the k-th N-gram phrase. k h k (c i ) represents w k The number of times h appears in the generated statement k (s ij ) represents w k The number of times a statement appears in a truth statement. Since short sentences may inflate N-gram scores, BLEU introduces a length penalty factor (BP) to correct this bias and ensure evaluation accuracy. Its calculation formula is shown below.
[0111]
[0112] In the formula, l c That is, the length of the statement generated by the model; l s That is, the length of the reference (real) sentence. When l s Not less than l c When this happens, a penalty factor is introduced, and the BLEU-N score is calculated according to the following formula.
[0113]
[0114] METEOR (Metric for Evaluation of Translation with Explicit Ordering) calculates the harmonic mean of precision and recall, with recall considered more important than precision. In addition to evaluating exact matches, METEOR considers stemming and synonym matching when evaluating machine translation quality, effectively compensating for the shortcomings of BLEU. It achieves a mapping evaluation between the generated text and the reference text through exact word, stemming, and synonym matching, which is highly correlated with human judgment and improves the accuracy of the evaluation. Its calculation process is as follows:
[0115] METEOR=(1-Pen)F mean
[0116] in,
[0117]
[0118]
[0119]
[0120]
[0121] The default parameters are α, β, and θ, and a key preset correction value m is introduced. This correction value is based on the WordNet thesaurus and is obtained by minimizing consecutive ordered chunks in the sentence.
[0122] CIDEr can be seen as an evaluation metric combining BLEU and vector space, specifically designed for image captioning. In image captioning, one key evaluation criterion is whether the generated description contains crucial information. Therefore, different phrases should be assigned different weights to reduce the weight of non-keywords. Its calculation formula is shown below.
[0123]
[0124] Among them, c i s ij The meanings of 'm' and 's' are the same as those described in the above formula. i This represents the i-th truth statement, g n (c i ) indicates the generation statement c i The TF-IDF values of all n-tuples in the set are given by the formula below.
[0125]
[0126] Where |D| represents the total number of statements in the corpus, |{j:c i ∈d j}| indicates the number of sentences in the corpus that contain the phrase ci.
[0127] The ROUGE-L evaluation method is based on the longest common subsequence principle and focuses on measuring the consistency of two sentences in terms of part of speech and word order, with particular emphasis on the recall rate of N-grams from the reference text. In image description tasks, an image is usually accompanied by a text description; however, due to the possibility of multiple true descriptions, the following formula is used to comprehensively calculate the evaluation result of an image description.
[0128]
[0129] Where β is a constant, R lcs and P lcs Recall and precision are represented respectively:
[0130]
[0131]
[0132] Where LCS(c, s) represents the longest common subsequence of the generating statement c and the truth statement s, l s and l c same as BLEU s and l c .
[0133] SPICE is an efficient method for evaluating image description tasks. It relies on descriptions of target, attribute, and relation triples and uses a probabilistic context-free grammatical dependency parser to construct a syntactic dependency tree, which is then mapped to a graph structure to derive the SPICE score. SPICE evaluation considers not only the content of the description but also its syntactic structure and semantic information. The calculation formula is shown below.
[0134]
[0135]
[0136]
[0137] Here, c and s represent the generating statement and the truth statement, respectively. The G() function is responsible for mapping the triples in the sentence to a graph, while the T() operation is responsible for converting the graph into a tuple.
[0138] The experimental results and analysis of this invention are as follows:
[0139] To facilitate use, this invention uses the Ceymo dataset of manually annotated text from the prior art to conduct experiments and verify the model of this invention.
[0140] The Ceymo dataset contains 11 categories of road traffic signs, and descriptive text annotations have been provided for each of these 11 categories, based on their shape, color, and composition. The annotated data format is one image per TXT file, with the TXT file containing the annotated text description. Examples of some images and their annotated text in the dataset are shown below. Figure 7 As shown. Based on the model's input requirements, the txt format file needs further data processing to convert it into the JSON format required by the network. Specific details include the image location, the caption required for the model input, and the image ID, such as... Figure 8 As shown
[0141] To verify whether the model is overfitting during training, the 2099 images in the Ceymo dataset training set were divided into a training set and a validation set in an 8:2 ratio. The resulting training set contained 1679 images, and the validation set contained 420 images. The test set consisted of 788 images.
[0142] The experimental environment for this invention is as follows: The experiments were conducted using Blip as the experimental benchmark. The comparative experiments primarily compared it with the Blip2 and InstructBlip algorithms. The experimental environment for InstructBlip was the same as for Blip, and the development environment for Blip2 was the same as for Blip, but the training hardware environment configuration for Blip2 differed slightly from that of Blip. The hardware environment configuration is shown in Table 1. The development environment consisted of Python 3.8 and CUDA 11.3, and the algorithm support libraries were mainly PyTorch 1.11.0 and OpenCVPython. The training hardware environment configuration for the Blip algorithm is shown in Table 2.
[0143] Table 1 Hardware Environment Configuration Table for Blip2
[0144]
[0145] Table 2 Blip's Hardware Environment Configuration Table
[0146]
[0147] The algorithm training parameters are shown in Table 3, and the model training parameters are shown in the following table.
[0148] Table 3 Training Parameters
[0149]
[0150]
[0151] The sentence generation strategy experiment comparison is as follows: In the Image Caption testing phase, two search methods are typically used to obtain the output sentence. The first method is the greedy sampling method, which selects the word with the highest probability at each time step as the output. Another commonly used search method is Beam Search. Here, "beamsize" refers to the parameter used to control the size of the search space when generating image descriptions. Specifically, beam size determines the number of candidate words selected at each time step. By retaining the top beam-size most likely candidate words, the system can more accurately estimate the optimal solution during the generation process. This invention explores the performance of models Blip, Blip2, and InstructBlip with beam size parameters set to 3, 4, and 5. Table 4 shows that Blip achieves the best performance on B@1, B@2, B@3, B@4, ROUGE-L, and CIDEr metrics when the beam size is 3. Table 5 shows that Blip2 achieves the best performance on B@1, B@2, B@3, B@4, and ROUGE-L metrics when the beam size is 5. Table 6 shows that InstructBlip achieves the best performance across all metrics when the beam size is 4. Although Blip2 and InstructBlip achieve the best performance with beam sizes of 4 or 5, they are slightly less effective than Blip. A larger beam size may increase computational costs. Therefore, this invention ultimately sets the beam size to 3.
[0152] Table 4 Comparison of experimental results for different beam sizes in Blip
[0153]
[0154] Table 5 Comparison of experimental results for Blip2 with different beam sizes
[0155]
[0156] Table 6. Comparison of experimental results for different beam sizes in InstructBlip
[0157]
[0158] The performance comparison analysis of multiple algorithms is as follows: To demonstrate the performance of the network model, the Blip model was compared with the Blip2 and InstructBlip models on the dataset. As shown in Table 7, the Blip model outperformed the Blip2 and InstructBlip models in all evaluation metrics. These results indicate that the Blip model can generate high-quality sentences, thus verifying the superiority of the Blip model.
[0159] Table 7 Performance Comparison Results of Different Models
[0160] Beamsize B@1 B@2 B@3 B@4 METEOR ROUGE-L CIDEr SPICE Blip 48.7 35.1 26.5 20.5 27.7 40.8 99.7 37.8 Blip2 47.4 34.0 25.3 19.4 27.5 39.5 87.6 36.1 InstructBlip 45.4 32.0 23.7 18.3 24.3 37.2 69.3 33.1
[0161] An example of image description generated by the model is as follows: In order to qualitatively analyze the difference between the effect of this invention and other algorithms, this invention selects the Blip algorithm, Blip2 algorithm and InstructBlip algorithm to perform text description on the same image, and finally compares and analyzes the results.
[0162] Examples of image descriptions generated by the model, such as Figure 9A , 9B As shown in 9C.
[0163] The performance of the Blip model was analyzed based on the descriptive sentences generated by different algorithms for the same image. However, this comparison only reflects the performance differences between different algorithms. An ideal image semantic understanding algorithm should generate descriptive sentences that are more consistent with human cognition. Therefore, further comparison with manually annotated sentences is needed to truly reflect the difference between the generated sentences and human cognition. In Figure 9, several models can generate accurate descriptive sentences in almost all cases. Figure 9A The InstructBlip model incorrectly generated "right liane", while other models generated correct descriptions of the traffic signs in the image. Figure 9B All models accurately describe pedestrian crossing generation, but Blip and Blip2 lack the meaning of "fixed intervals," and InstructBlip lacks the meaning of "short, wide." Figure 9C As can be seen, accurate sentence descriptions were generated for various models of the diamond-shaped symbol.
[0164] Figure 10 shows some examples of incomplete descriptions. The marker descriptions generated by the model deviate from the actual markers in the images, either failing to fully describe the markers or describing them incorrectly. Figure 10AThe straight-ahead signs all generated correct descriptions. Blip and InstructBlip generated descriptions for right-turn signs that did not exist in the image, but failed to generate descriptions for straight-ahead left-turn signs that did exist in the image. Only Blip2 generated a description for the left-turn sign, but it was incomplete. Figure 10B All models generated accurate descriptions for the "straight ahead" signs. However, none of the models generated accurate and complete descriptions for the "straight ahead and right turn" signs. The reason for the inaccurate descriptions may be that the signs were not fully displayed or were unclear, leading to incorrect judgments during object detection.
[0165] In conclusion, Blip achieved the best performance in generating image descriptions of road traffic signs.
[0166] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0167] Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. Furthermore, those skilled in the art will recognize that, based on the ideas of this invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A method for recognizing road traffic signs based on image semantic understanding, characterized in that, include: Image acquisition is performed using image acquisition equipment; The collected images are preprocessed to construct a road traffic sign dataset; A road traffic sign detection model is constructed based on the Blip network, and the model is trained based on a road traffic sign dataset to obtain the trained road traffic sign detection model. The performance of the trained road traffic sign detection model was tested based on a road traffic sign dataset; a road traffic sign detection model was constructed based on the Blip network, specifically as follows: A road traffic sign detection model is constructed, comprising a visual encoder, a text encoder, a visual-text encoder, and a visual-text decoder. The visual encoder uses a Vision Transformer structure and a Cross-Attention mechanism to extract feature information of road signs from images, using it as one of the joint representations. The text encoder uses a BERT architecture to extract text features. The visual-text encoder employs a Cross-Attention mechanism, with the attention part using a bidirectional Self-Attention mechanism, introducing an additional [Encode] token for binary classification using image and text features. The visual-text decoder also employs a Cross-Attention mechanism, with the attention part using Casual-Attention, introducing an additional [DEcode] token and an end token for extracting text features of road traffic signs. The road traffic sign detection model combines the image features extracted by the visual encoder and the text features extracted by the visual-text decoder through a cross-attention mechanism to form a new feature representation containing multimodal information of road signs. The road traffic sign detection model uses Blip's downstream task Image Caption to generate a text description for the road traffic sign image, and uses a language model objective function to generate a text description for a given road traffic sign image.
2. The method according to claim 1, characterized in that, Image acquisition is performed using an image acquisition device, specifically as follows: Based on camera equipment installed on the vehicle, image data of the road scene in front of the vehicle is acquired in real time. The camera equipment includes high-definition cameras and night vision cameras, which are used to switch according to different environmental requirements.
3. The method according to claim 1, characterized in that, The acquired images are preprocessed to construct a road traffic sign dataset, specifically as follows: Acquire the collected image data and resize it to unify the image size; Denoise removal is performed on the image data after it has been resized. Perform color correction on the denoised image data; Perform data augmentation processing on the color-corrected image data; Manually annotate the image data after data augmentation. The labeled image data is aggregated to generate a road traffic sign dataset.
4. The method according to claim 3, characterized in that, Denoising is performed on image data after it has been resized using Gaussian filtering or median filtering to remove noise from the image.
5. The method according to claim 4, characterized in that, Color correction is performed on the denoised image data, specifically as follows: The image data after denoising is acquired and inspected. For images with uneven lighting, white balance adjustment and color correction are performed.
6. The method according to claim 5, characterized in that, Data augmentation processing is performed on the color-corrected image data, specifically as follows: Image enhancement is performed on color-corrected image data based on an image enhancement model. The image enhancement model includes a separate feature processing module, a joint feature processing module, and a feature fusion module. The separate feature processing module is used to learn five private feature maps, the joint feature processing module is used to learn a public feature map, and the feature fusion module is used to uniformly represent the private and public feature maps, ultimately achieving effective image enhancement.