Traffic speed limit sign identification method based on multi-modal model and driving assistance system

Through the traffic speed limit sign recognition method based on multimodal model, combined with traffic scene images and auxiliary information, multi-type recognition of speed limit signs is realized, solving the problem of poor recognition effect in the prior art, and improving the robustness and efficiency of recognition.

CN119942498APending Publication Date: 2025-05-06MINGSHANG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510016441.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has poor recognition effects when identifying speed limit signs, especially when dynamic speed limit and implicit speed limit signs, and is not robust enough to change ambient lighting conditions.

Method used

The traffic speed limit sign recognition method based on the multimodal model is adopted. By obtaining traffic scene images and auxiliary information, the image features and semantic information are extracted using the pre-trained model, and multimodal feature fusion is performed, and the speed limit recognition results are finally outputted through the classifier.

Benefits of technology

It realizes one-time recognition of multiple types of speed limit signs of explicit speed limit, dynamic speed limit and implicit speed limit signs, reducing operation complexity and time consumption, improving recognition effect, and adapting to different ambient lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942498A_ABST
    Figure CN119942498A_ABST
Patent Text Reader

Abstract

The invention relates to a traffic speed limit sign recognition method based on a multi-modal model and a driving assistance system, and the method comprises the steps: obtaining a traffic scene image and auxiliary information; extracting image feature information of the traffic scene image and semantic information of the auxiliary information by using a pre-trained model; fusing the image feature information and the semantic information to obtain multi-modal feature information; based on the multi-modal feature information, a traffic speed limit sign recognition result is obtained, and the traffic speed limit sign recognition result comprises one or more of explicit speed limit, dynamic speed limit and implicit speed limit. According to the invention, one-time identification of multiple types of speed limit signs is realized, and the identification effect of the speed limit sign and the implicit speed limit sign is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent traffic technology, and in particular to a traffic speed limit sign recognition method and a driving assistance system based on a multimodal model. Background Art

[0002] Traffic sign recognition technology, as an indispensable part of intelligent driving assistance systems and fully autonomous vehicles, is becoming increasingly important.

[0003] In the modern transportation system, although the GPS navigation system provides accurate navigation services to drivers in most cases with its efficient and convenient path planning capabilities, the reliability of the GPS navigation system is often challenged in some special scenarios, such as road construction, temporary traffic control or signal loss in remote areas. Especially in these cases, temporarily installed traffic signs become a key source of information for guiding driving direction and safety, and the GPS system often cannot update this information in a timely manner. Therefore, it is particularly urgent and important to develop a vision-based system that can autonomously identify various types of traffic signs.

[0004] At present, a variety of methods have emerged in the field of traffic sign recognition, aiming to improve the accuracy and robustness of recognition. Among them, the method based on color space threshold segmentation is a relatively intuitive and widely used technology. This method usually uses the unique color characteristics of traffic signs in a specific color space (such as HSV, HSI or CBH color space) to convert the image into a binary image by setting a reasonable threshold range, thereby highlighting the traffic sign area and facilitating subsequent classification processing. Machine learning algorithms such as support vector machines (SVM) or neural networks are often used for classification tasks at this stage. By learning a large number of samples, accurate recognition of different types of traffic signs can be achieved. However, this type of method is highly dependent on color features. Once the ambient lighting conditions change, such as at night, dusk or strong reflections, the recognition effect may be greatly reduced. Another common method is traffic sign recognition based on edge detection. This type of method uses the characteristics of regular shapes and clear edges of traffic signs. Through image processing techniques such as Fourier descriptors, Hough transforms or distance transforms, the edge information of the signs is extracted, and then shape matching and recognition are performed. This method reduces the reliance on color features to a certain extent, but in the case of complex background or partial occlusion of the logo, the extraction of edge information may become difficult, thus affecting the recognition accuracy. Summary of the invention

[0005] The present invention aims at the problem of how to improve the recognition effect of speed limit signs and implicit speed limit signs.

[0006] In view of the above problems, in a first aspect, the present application provides a traffic speed limit sign recognition method based on a multimodal model, the recognition method comprising:

[0007] Acquire traffic scene images and auxiliary information;

[0008] Extracting image feature information of the traffic scene image and semantic information of the auxiliary information using a pre-trained model;

[0009] fusing the image feature information and the semantic information to obtain multimodal feature information;

[0010] Based on the multimodal feature information, a traffic speed limit sign recognition result is obtained, and the traffic speed limit sign recognition result includes one or more of an explicit speed limit, a dynamic speed limit and an implicit speed limit.

[0011] In some embodiments, the traffic scene image is image data acquired by a camera.

[0012] In some embodiments, the auxiliary information is dynamic text data collected by an OCR module.

[0013] In some embodiments, the extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using a pre-trained model includes:

[0014] A convolutional neural network is used to extract an image feature vector of the traffic scene image.

[0015] In some embodiments, the traffic speed limit sign recognition method further includes:

[0016] The convolutional neural network is trained based on traffic scene training images and label information; wherein the image feature information extracted by the convolutional neural network includes: color, shape and digital content.

[0017] In some embodiments, the extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using a pre-trained model includes:

[0018] A semantic feature vector of the auxiliary information is extracted using a semantic extraction network, wherein the semantic extraction network is a transformer.

[0019] In some embodiments, the fusing the image feature information and the semantic information to obtain multimodal feature information includes:

[0020] Based on the attention weight of the image feature information and the attention weight of the semantic information, a weighted sum of the image feature information and the semantic information is calculated to obtain the multimodal feature information.

[0021] In some embodiments, obtaining a traffic speed limit sign recognition result based on the multimodal feature information includes:

[0022] The multimodal feature information is sent to a classifier, and the classifier outputs a speed limit result;

[0023] In some embodiments, the traffic speed limit sign recognition method further includes:

[0024] The pre-trained model is pruned to reduce the computational complexity of the pre-trained model.

[0025] In a second aspect, the present application provides a driving assistance system, the driving assistance system comprising

[0026] A memory storing a traffic speed limit sign recognition program,

[0027] A processor is connected to the memory, and when the processor executes the traffic speed limit sign recognition program, the traffic speed limit sign recognition method as described above is implemented.

[0028] The present application obtains multimodal feature information by multimodally fusing the image feature information of the traffic scene image and the semantic information of the auxiliary information; then identifies the multimodal feature information to obtain the recognition results of the explicit speed limit, dynamic speed limit and implicit speed limit. The present application realizes the one-time recognition of multiple types of speed limit signs, and does not require multiple processing to separately identify the dynamic speed limit and static speed limit signs, thereby reducing the computational complexity and time consumption, and improving the recognition effect of the speed limit signs and implicit speed limit signs. The recognition results can be directly used by the driving assistance system to limit the speed or convey the speed limit information to the user through indicator devices such as LEDs. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a flowchart of an embodiment of a traffic speed limit sign recognition method based on a multimodal model of the present application.

[0030] Figure 2 A schematic diagram of a scenario of an embodiment of a traffic speed limit sign recognition method based on a multimodal model of the present application;

[0031] Figure 3 A structural diagram of an embodiment of a multimodal model of the present application;

[0032] Figure 4 This is a flowchart of another embodiment of the traffic speed limit sign recognition method based on a multimodal model of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0034] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0035] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in the present invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any technician in the field to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present invention.

[0036] Example 1

[0037] The existing single-modal recognition methods have poor recognition effect on dynamic speed limit signs (such as information on LED display screens) and invisible speed limit signs (such as "school ahead" and "residential area").

[0038] Reference Figure 1 In some embodiments, the traffic speed limit sign recognition method based on the multimodal model of the present application includes:

[0039] S100, acquiring a traffic scene image and auxiliary information;

[0040] In this embodiment, a high-definition camera 10 can be installed in a suitable position, such as at the front windshield of a vehicle or on a fixed bracket beside the road, to ensure that the traffic scene image can be clearly captured. The camera 10 should have sufficient resolution and frame rate to capture rich details and dynamic changes. The camera 10 transmits the collected image data to the processing unit of the system in real time through a wired or wireless connection.

[0041] Auxiliary information can be extracted by the OCR module, which first preprocesses the input image, including image enhancement, binarization, denoising and other operations to improve the accuracy of character recognition. Then, the preprocessed image is recognized using a character recognition algorithm to extract the text content. The extracted dynamic text data is synchronized with the image data collected by the camera 10 to ensure the temporal consistency of the two. For example, synchronization can be achieved by adding timestamps to the image data and text data.

[0042] For example, Figure 2 As shown, a schematic diagram of a traffic speed limit sign recognition scenario of the multimodal model of the present application is shown. In the figure, P30, P40, P50, and P60 are the recognized speed limit results, and the speed limits are 30km / h, 40km / h, 50km / h, and 60km / h.

[0043] S200: Using a pre-trained model to extract image feature information of the traffic scene image and semantic information of the auxiliary information.

[0044] Pre-trained models can include Convolutional Neural Networks (CNN) and Transformer.

[0045] For example, the traffic scene image captured by the camera 10 can be input into the CNN, and after a series of convolution, pooling and activation function operations, the features of the image are gradually extracted. The convolution layer performs convolution operations with the input image through the convolution kernel to extract the local features of the image. The pooling layer downsamples the output of the convolution layer to reduce the size of the feature map while retaining important feature information. Activation functions such as ReLU are used

[0046] It is used to increase the nonlinear expression ability of the network. At different levels of the network, the extracted features have different levels of abstraction. The lower-level features mainly reflect local information such as the edge and texture of the image, while the higher-level features represent more abstract shapes, colors, and semantic information. Finally, the image feature vector is extracted from the last layer or a specific intermediate layer of the CNN for subsequent processing.

[0047] Exemplarily, the dynamic text data extracted by OCR and other sensor-assisted information are input into the Transformer, and semantic features are gradually extracted through a series of self-attention mechanisms and feedforward neural network operations. The self-attention mechanism calculates the attention weights between each element in the input sequence to achieve weighted summation of information at different positions, thereby capturing the global semantic information of the input sequence. The feedforward neural network further processes the output of the self-attention mechanism to increase the nonlinear expression ability of the model. At different levels of the Transformer, the extracted semantic features have different levels of abstraction. The lower-level features mainly reflect the semantic information of a single word or phrase, while the higher-level features represent the semantic information of more complex sentences or paragraphs. Finally, the semantic feature vector is extracted from the last layer or a specific intermediate layer of the Transformer for subsequent processing.

[0048] It should be noted that this embodiment does not limit the network for extracting image feature information and semantic information. In other embodiments, one or more combinations of recurrent neural networks (RNNs), deep neural networks (DNNs), and twin networks can also be used. Similarly, semantic information can also be extracted using feature extractors other than Transformer.

[0049] S300: Fusing the image feature information and the semantic information to obtain multimodal feature information.

[0050] Exemplarily, in this embodiment, the feature vectors output by the visual feature extraction module and the semantic feature extraction module can be input into the attention mechanism, and after attention calculation and weighted summation operations, a fused multimodal feature vector can be obtained. During the attention calculation process, different attention weights are automatically assigned according to the feature importance of different modalities. For example, the correlation between visual features and semantic features can be used as the basis for calculating the attention weight. The higher the correlation, the greater the assigned attention weight. The weighted summation operation weights the feature vectors of different modalities according to the attention weights to obtain a fused multimodal feature vector. This vector contains information such as the shape, color, and digital content of the visual features, as well as the dynamic LED display information of the semantic features and the semantic information of the invisible speed limit sign, realizing the deep fusion of multimodal data.

[0051] In other embodiments, other methods may be used to fuse features, such as concatenating semantic information and image feature information and sending the concatenated information to a fully connected neural network for integration.

[0052] S400. Obtain a traffic speed limit sign recognition result based on the multimodal feature information, wherein the traffic speed limit sign recognition result includes one or more of an explicit speed limit, a dynamic speed limit and an implicit speed limit.

[0053] Exemplarily, the multimodal feature information may be a multimodal feature vector. In this embodiment, the multimodal feature vector output by the multimodal fusion module may be input into the classifier 60. After a series of calculation and decision-making processes, the recognition results of the explicit speed limit, the dynamic speed limit and the invisible speed limit are output. The classifier 60 first further processes the input feature vector, such as fully connected layers, activation functions and other operations, to increase the nonlinear expression ability of the model. Then, according to the decision function of the classifier 60, the probability that the input feature vector belongs to different speed limit types is calculated. Finally, the speed limit type with the highest probability is selected as the output result according to the probability size. For example, if the probability of the explicit speed limit is the highest, the recognition result of the explicit speed limit is output; if the probability of the dynamic speed limit is the highest, the recognition result of the dynamic speed limit is output; if the probability of the invisible speed limit is the highest, the recognition result of the invisible speed limit is output. Exemplarily, the last layer of the classifier 60 may be a softmax function, and the softmax function outputs the probability that the feature vector belongs to different speed limit types.

[0054] The present application obtains multimodal feature information by multimodally fusing the image feature information of the traffic scene image and the semantic information of the auxiliary information; then identifies the multimodal feature information to obtain the recognition results of the explicit speed limit, dynamic speed limit and implicit speed limit. The present application realizes the one-time recognition of multiple types of speed limit signs, and does not require multiple processing to separately identify the dynamic speed limit and static speed limit signs, thereby reducing the computational complexity and time consumption, and improving the recognition effect of the speed limit signs and implicit speed limit signs. The recognition results can be directly used by the driving assistance system to limit the speed or convey the speed limit information to the user through indicator devices such as LEDs.

[0055] Reference Figure 3 , the structure of the multimodal model in an embodiment of the present application is given as an example and the workflow is introduced. First, the camera 10 collects traffic scene images, and the auxiliary sensor 20 collects environmental images; then they are sent to the image preprocessing module 30 for image preprocessing 31 and OCR processing 32 respectively, to obtain preprocessed traffic scene images and dynamic text data; then the preprocessed traffic scene images and dynamic text data are sent to the feature extraction module 40 respectively, the CNN 41 of the feature extraction module 40 extracts information such as the shape, color and digital content of the traffic scene image, and the Transformer 42 of the feature extraction module 40 extracts the dynamic LED display information of the semantic features and the semantic information of the invisible speed limit sign, and then the multimodal feature fusion module 50 fuses the image feature information and the semantic information to obtain multimodal feature information, and finally the multimodal feature information is sent to the classifier 60 for classification processing, and the recognition result of the speed limit sign can be obtained at one time.

[0056] In some embodiments, the traffic scene image is image data collected by the camera 10 .

[0057] In this embodiment, a high-definition camera 10 can be installed at a suitable location, such as at the front windshield of a vehicle or on a fixed bracket beside the road, to ensure that traffic scene images can be clearly captured.

[0058] In some embodiments, the auxiliary information is dynamic text data collected by an OCR module.

[0059] Exemplarily, the OCR module may be just a computing module that recognizes dynamic text data in the traffic scene image.

[0060] Exemplarily, the OCR module further includes a calculation module and an auxiliary sensor 20 . The auxiliary sensor 20 may be a visual sensor, an infrared sensor or other sensor capable of collecting environmental information. The calculation module extracts dynamic text data from the auxiliary information collected by the auxiliary sensor 20 .

[0061] Reference Figure 4 In some embodiments, S200, extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using a pre-trained model, includes:

[0062] S201. Use a convolutional neural network to extract an image feature vector of the traffic scene image.

[0063] In this embodiment, the convolutional neural network may include a convolution layer, a pooling layer, a normalization layer, and an activation function. The convolution layer performs a convolution operation with the input image through a convolution kernel to extract local features of the image. The pooling layer downsamples the output of the convolution layer to reduce the size of the feature map while retaining important feature information. Activation functions such as ReLU are used to increase the nonlinear expression ability of the network. At different levels of the network, the extracted features have different levels of abstraction. The lower-level features mainly reflect local information such as the edges and textures of the image, while the higher-level features represent more abstract shapes, colors, and semantic information. Finally, the image feature vector is extracted from the last layer or a specific intermediate layer of CNN41 for subsequent processing.

[0064] In some embodiments, the traffic speed limit sign recognition method further includes:

[0065] The convolutional neural network is trained based on traffic scene training images and label information; wherein the image feature information extracted by the convolutional neural network includes: color, shape and digital content.

[0066] In the present application, the traffic speed limit sign recognition model may include: CNN41, Transformer42, multimodal feature fusion module 50 and classifier 60. Among them, the information output by CNN41 and Transformer42 is fused by the multimodal feature fusion module 50 to obtain a multimodal feature vector. This vector contains not only information such as the shape, color and digital content of the visual features, but also the dynamic LED display information of the semantic features and the semantic information of the invisible speed limit sign, realizing the deep fusion of multimodal data, improving the input information of the classifier 60, and thus better identifying the traffic speed limit sign.

[0067] Reference Figure 4 Exemplarily, the extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using the pre-trained model includes:

[0068] S202. Use a semantic extraction network to extract a semantic feature vector of the auxiliary information, wherein the semantic extraction network is Transformer42.

[0069] Reference Figure 4 In some embodiments, S300, fusing the image feature information and the semantic information to obtain multimodal feature information includes:

[0070] S301. Based on the attention weight of the image feature information and the attention weight of the semantic information, a weighted sum of the image feature information and the semantic information is calculated to obtain the multimodal feature information.

[0071] Exemplarily, in this embodiment, the feature vectors output by the visual feature extraction module and the semantic feature extraction module can be input into the attention mechanism, and after attention calculation and weighted summation operations, a fused multimodal feature vector can be obtained. During the attention calculation process, different attention weights are automatically assigned according to the importance of features of different modalities.

[0072] Reference Figure 4 In some embodiments, S400, obtaining a traffic speed limit sign recognition result based on the multimodal feature information includes:

[0073] S401, sending the multimodal feature information to the classifier 60, and the classifier 60 outputs a speed limit result.

[0074] Exemplarily, the classifier 60 may include a fully connected layer, an activation function, and an output layer. The classifier 60 first further processes the input feature vector, such as fully connected layers, activation functions, and other operations, to increase the nonlinear expression ability of the model. Then, according to the decision function of the classifier 60, the probability that the input feature vector belongs to different speed limit types is calculated. Finally, the speed limit type with the highest probability is selected as the output result according to the probability size. For example, if the probability of explicit speed limit is the highest, the recognition result of explicit speed limit is output; if the probability of dynamic speed limit is the highest, the recognition result of dynamic speed limit is output; if the probability of invisible speed limit is the highest, the recognition result of invisible speed limit is output.

[0075] In some embodiments, the traffic speed limit sign recognition method further includes: pruning the pre-trained model to reduce the computational complexity of the pre-trained model.

[0076] It should be noted that the model needs to be embedded in the vehicle's driving assistance system, requiring the computational complexity of the model to be small enough. This embodiment applies model pruning and quantization technology to reduce computational complexity; uses mixed precision computing to optimize memory usage efficiency. Combined with the hardware acceleration function of embedded devices (such as NPU and DSP), the running speed is improved; the model inference time is optimized to ensure that real-time requirements are met.

[0077] Example 2

[0078] In a second aspect, the present application provides a driving assistance system, the driving assistance system comprising

[0079] A memory storing a traffic speed limit sign recognition program,

[0080] A processor is connected to the memory, and when the processor executes the traffic speed limit sign recognition program, the following traffic speed limit sign recognition method is implemented:

[0081] S100, acquiring a traffic scene image and auxiliary information;

[0082] In this embodiment, a high-definition camera 10 can be installed in a suitable position, such as at the front windshield of a vehicle or on a fixed bracket beside the road, to ensure that the traffic scene image can be clearly captured. The camera 10 should have sufficient resolution and frame rate to capture rich details and dynamic changes. The camera 10 transmits the collected image data to the processing unit of the system in real time through a wired or wireless connection.

[0083] Auxiliary information can be extracted by the OCR module, which first preprocesses the input image, including image enhancement, binarization, denoising and other operations to improve the accuracy of character recognition. Then, the preprocessed image is recognized using a character recognition algorithm to extract the text content. The extracted dynamic text data is synchronized with the image data collected by the camera 10 to ensure the temporal consistency of the two. For example, synchronization can be achieved by adding timestamps to the image data and text data.

[0084] For example, Figure 2 As shown, a schematic diagram of a traffic speed limit sign recognition scenario of the multimodal model of the present application is shown. In the figure, P30, P40, P50, and P60 are the recognized speed limit results, and the speed limits are 30km / h, 40km / h, 50km / h, and 60km / h.

[0085] S200: Using a pre-trained model to extract image feature information of the traffic scene image and semantic information of the auxiliary information.

[0086] Pre-trained models can include Convolutional Neural Networks (CNN) and Transformer.

[0087] For example, the traffic scene image captured by the camera 10 can be input into the CNN, and after a series of convolution, pooling and activation function operations, the features of the image are gradually extracted. The convolution layer performs convolution operations with the input image through the convolution kernel to extract the local features of the image. The pooling layer downsamples the output of the convolution layer to reduce the size of the feature map while retaining important feature information. Activation functions such as ReLU are used

[0088] It is used to increase the nonlinear expression ability of the network. At different levels of the network, the extracted features have different levels of abstraction. The lower-level features mainly reflect local information such as the edge and texture of the image, while the higher-level features represent more abstract shapes, colors, and semantic information. Finally, the image feature vector is extracted from the last layer or a specific intermediate layer of the CNN for subsequent processing.

[0089] Exemplarily, the dynamic text data extracted by OCR and other sensor-assisted information are input into the Transformer, and semantic features are gradually extracted through a series of self-attention mechanisms and feedforward neural network operations. The self-attention mechanism calculates the attention weights between each element in the input sequence to achieve weighted summation of information at different positions, thereby capturing the global semantic information of the input sequence. The feedforward neural network further processes the output of the self-attention mechanism to increase the nonlinear expression ability of the model. At different levels of the Transformer, the extracted semantic features have different levels of abstraction. The lower-level features mainly reflect the semantic information of a single word or phrase, while the higher-level features represent the semantic information of more complex sentences or paragraphs. Finally, the semantic feature vector is extracted from the last layer or a specific intermediate layer of the Transformer for subsequent processing.

[0090] It should be noted that this embodiment does not limit the network for extracting image feature information and semantic information. In other embodiments, one or more combinations of recurrent neural networks (RNNs), deep neural networks (DNNs), and twin networks can also be used. Similarly, semantic information can also be extracted using feature extractors other than Transformer.

[0091] S300: Fusing the image feature information and the semantic information to obtain multimodal feature information.

[0092] Exemplarily, in this embodiment, the feature vectors output by the visual feature extraction module and the semantic feature extraction module can be input into the attention mechanism, and after attention calculation and weighted summation operations, a fused multimodal feature vector can be obtained. During the attention calculation process, different attention weights are automatically assigned according to the feature importance of different modalities. For example, the correlation between visual features and semantic features can be used as the basis for calculating the attention weight. The higher the correlation, the greater the assigned attention weight. The weighted summation operation weights the feature vectors of different modalities according to the attention weights to obtain a fused multimodal feature vector. This vector contains information such as the shape, color, and digital content of the visual features, as well as the dynamic LED display information of the semantic features and the semantic information of the invisible speed limit sign, realizing the deep fusion of multimodal data.

[0093] In other embodiments, other methods may be used to fuse features, such as concatenating semantic information and image feature information and sending the concatenated information to a fully connected neural network for integration.

[0094] S400. Obtain a traffic speed limit sign recognition result based on the multimodal feature information, wherein the traffic speed limit sign recognition result includes one or more of an explicit speed limit, a dynamic speed limit and an implicit speed limit.

[0095] Exemplarily, the multimodal feature information may be a multimodal feature vector. In this embodiment, the multimodal feature vector output by the multimodal fusion module may be input into the classifier 60. After a series of calculation and decision-making processes, the recognition results of the explicit speed limit, the dynamic speed limit and the invisible speed limit are output. The classifier 60 first further processes the input feature vector, such as fully connected layers, activation functions and other operations, to increase the nonlinear expression ability of the model. Then, according to the decision function of the classifier 60, the probability that the input feature vector belongs to different speed limit types is calculated. Finally, the speed limit type with the highest probability is selected as the output result according to the probability size. For example, if the probability of the explicit speed limit is the highest, the recognition result of the explicit speed limit is output; if the probability of the dynamic speed limit is the highest, the recognition result of the dynamic speed limit is output; if the probability of the invisible speed limit is the highest, the recognition result of the invisible speed limit is output. Exemplarily, the last layer of the classifier 60 may be a softmax function, and the softmax function outputs the probability that the feature vector belongs to different speed limit types.

[0096] The present application obtains multimodal feature information by multimodally fusing the image feature information of the traffic scene image and the semantic information of the auxiliary information; then identifies the multimodal feature information to obtain the recognition results of explicit speed limit, dynamic speed limit and implicit speed limit. The present application realizes the one-time recognition of multiple types of speed limit signs without the need for multiple processing to separately identify dynamic speed limit and static speed limit signs, thereby reducing the computational complexity and time consumption. The recognition results can be directly used by the driving assistance system to limit the speed or convey the speed limit information to the user through indicator devices such as LEDs.

[0097] Reference Figure 3 , the structure of the multimodal model in an embodiment of the present application is given as an example and the workflow is introduced. First, the camera 10 collects traffic scene images, and the auxiliary sensor 20 collects environmental images; then they are sent to the image preprocessing module 30 for image preprocessing 31 and OCR processing 32 respectively, to obtain preprocessed traffic scene images and dynamic text data; then the preprocessed traffic scene images and dynamic text data are sent to the feature extraction module 40 respectively, the CNN 41 of the feature extraction module 40 extracts information such as the shape, color and digital content of the traffic scene image, and the Transformer 42 of the feature extraction module 40 extracts the dynamic LED display information of the semantic features and the semantic information of the invisible speed limit sign, and then the multimodal feature fusion module 50 fuses the image feature information and the semantic information to obtain multimodal feature information, and finally the multimodal feature information is sent to the classifier 60 for classification processing, and the recognition result of the speed limit sign can be obtained at one time.

[0098] In some embodiments, the traffic scene image is image data collected by the camera 10 .

[0099] In this embodiment, a high-definition camera 10 can be installed at a suitable location, such as at the front windshield of a vehicle or on a fixed bracket beside the road, to ensure that traffic scene images can be clearly captured.

[0100] In some embodiments, the auxiliary information is dynamic text data collected by an OCR module.

[0101] Exemplarily, the OCR module may be just a computing module that recognizes dynamic text data in the traffic scene image.

[0102] Exemplarily, the OCR module further includes a calculation module and an auxiliary sensor 20 . The auxiliary sensor 20 may be a visual sensor, an infrared sensor or other sensor capable of collecting environmental information. The calculation module extracts dynamic text data from the auxiliary information collected by the auxiliary sensor 20 .

[0103] Reference Figure 4 In some embodiments, S200, extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using a pre-trained model, includes:

[0104] S201. Use a convolutional neural network to extract an image feature vector of the traffic scene image.

[0105] In this embodiment, the convolutional neural network may include a convolution layer, a pooling layer, a normalization layer, and an activation function. The convolution layer performs a convolution operation with the input image through a convolution kernel to extract local features of the image. The pooling layer downsamples the output of the convolution layer to reduce the size of the feature map while retaining important feature information. Activation functions such as ReLU are used to increase the nonlinear expression ability of the network. At different levels of the network, the extracted features have different levels of abstraction. The lower-level features mainly reflect local information such as the edges and textures of the image, while the higher-level features represent more abstract shapes, colors, and semantic information. Finally, the image feature vector is extracted from the last layer or a specific intermediate layer of CNN41 for subsequent processing.

[0106] In some embodiments, the traffic speed limit sign recognition method further includes:

[0107] The convolutional neural network is trained based on traffic scene training images and label information; wherein the image feature information extracted by the convolutional neural network includes: color, shape and digital content.

[0108] In the present application, the traffic speed limit sign recognition model may include: CNN41, Transformer42, multimodal feature fusion module 50 and classifier 60. Among them, the information output by CNN41 and Transformer42 is fused by the multimodal feature fusion module 50 to obtain a multimodal feature vector. This vector contains not only information such as the shape, color and digital content of the visual features, but also the dynamic LED display information of the semantic features and the semantic information of the invisible speed limit sign, realizing the deep fusion of multimodal data, improving the input information of the classifier 60, and thus better identifying the traffic speed limit sign.

[0109] Reference Figure 4 Exemplarily, the extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using the pre-trained model includes:

[0110] S202. Use a semantic extraction network to extract a semantic feature vector of the auxiliary information, wherein the semantic extraction network is Transformer42.

[0111] Reference Figure 4 In some embodiments, S300, fusing the image feature information and the semantic information to obtain multimodal feature information includes:

[0112] S301. Based on the attention weight of the image feature information and the attention weight of the semantic information, a weighted sum of the image feature information and the semantic information is calculated to obtain the multimodal feature information.

[0113] Exemplarily, in this embodiment, the feature vectors output by the visual feature extraction module and the semantic feature extraction module can be input into the attention mechanism, and after attention calculation and weighted summation operations, a fused multimodal feature vector can be obtained. During the attention calculation process, different attention weights are automatically assigned according to the importance of features of different modalities.

[0114] Reference Figure 4 In some embodiments, S400, obtaining a traffic speed limit sign recognition result based on the multimodal feature information includes:

[0115] S401, sending the multimodal feature information to the classifier 60, and the classifier 60 outputs a speed limit result.

[0116] Exemplarily, the classifier 60 may include a fully connected layer, an activation function, and an output layer. The classifier 60 first further processes the input feature vector, such as fully connected layers, activation functions, and other operations, to increase the nonlinear expression ability of the model. Then, according to the decision function of the classifier 60, the probability that the input feature vector belongs to different speed limit types is calculated. Finally, the speed limit type with the highest probability is selected as the output result according to the probability size. For example, if the probability of explicit speed limit is the highest, the recognition result of explicit speed limit is output; if the probability of dynamic speed limit is the highest, the recognition result of dynamic speed limit is output; if the probability of invisible speed limit is the highest, the recognition result of invisible speed limit is output.

[0117] In some embodiments, the traffic speed limit sign recognition method further includes: pruning the pre-trained model to reduce the computational complexity of the pre-trained model.

[0118] It should be noted that the model needs to be embedded in the vehicle's driving assistance system, requiring the computational complexity of the model to be small enough. This embodiment applies model pruning and quantization technology to reduce computational complexity; uses mixed precision computing to optimize memory usage efficiency. Combined with the hardware acceleration function of embedded devices (such as NPU and DSP), the running speed is improved; the model inference time is optimized to ensure that real-time requirements are met.

[0119] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0120] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0122] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0124] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0125] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A traffic speed limit sign recognition method based on a multimodal model, characterized in that: include: Acquire traffic scene images and auxiliary information; Extracting image feature information of the traffic scene image and semantic information of the auxiliary information using a pre-trained model; fusing the image feature information and the semantic information to obtain multimodal feature information; Based on the multimodal feature information, a traffic speed limit sign recognition result is obtained, and the traffic speed limit sign recognition result includes one or more of an explicit speed limit, a dynamic speed limit and an implicit speed limit.

2. The traffic speed limit sign recognition method according to claim 1, characterized in that: The traffic scene image is image data acquired through a camera.

3. The traffic speed limit sign recognition method according to claim 2, characterized in that: The auxiliary information is dynamic text data collected by the OCR module.

4. The method for identifying a traffic speed limit sign according to claim 3, characterized in that: The extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using the pre-trained model includes: A convolutional neural network is used to extract an image feature vector of the traffic scene image.

5. The method for recognizing a traffic speed limit sign according to claim 4, characterized in that: Also includes: The convolutional neural network is trained based on traffic scene training images and label information; wherein the image feature information extracted by the convolutional neural network includes: color, shape and digital content.

6. The method for recognizing a traffic speed limit sign according to claim 3, characterized in that: The extracting the image feature information of the traffic scene image and the semantic information of the auxiliary information using the pre-trained model includes: A semantic feature vector of the auxiliary information is extracted using a semantic extraction network, wherein the semantic extraction network is a transformer.

7. The method for recognizing a traffic speed limit sign according to any one of claims 1 to 6, characterized in that: The fusing the image feature information and the semantic information to obtain multimodal feature information includes: Based on the attention weight of the image feature information and the attention weight of the semantic information, a weighted sum of the image feature information and the semantic information is calculated to obtain the multimodal feature information.

8. The method for recognizing a traffic speed limit sign according to claim 7, characterized in that: The method of obtaining a traffic speed limit sign recognition result based on the multimodal feature information includes: The multimodal feature information is sent to a classifier, and the classifier outputs a speed limit result.

9. The method for recognizing a traffic speed limit sign according to claim 1, characterized in that: Also includes: The pre-trained model is pruned to reduce the computational complexity of the pre-trained model.

10. A driving assistance system, characterized in that: include A memory storing a traffic speed limit sign recognition program, A processor is connected to the memory, and when the processor executes the traffic speed limit sign recognition program, the traffic speed limit sign recognition method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Lightweight public sign intelligent identification and evaluation system

    CN121938012A