Traffic sign identification method and system based on computer vision

Through deep learning technology, feature extraction and correlation analysis are performed on digital maps and real-time road traffic sign images, and traffic sign category output is generated, which solves the problem of low accuracy in driverless cars when identifying traffic signs and improves the decision-making ability of the autonomous driving system.

CN120220115AInactive Publication Date: 2025-06-27HEBEI JIUJU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510356492.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing driverless cars identify traffic signs set on the road, they have low accuracy and are difficult to meet the decision-making needs of autonomous driving systems.

Method used

By acquiring the digital map traffic sign images collected by the database, real-time road traffic sign images collected by the camera, and real-time road panoramic images collected by the camera, feature extraction and correlation analysis are used using deep learning technology, and traffic sign category output is generated in combination with a classifier.

Benefits of technology

It improves the accuracy of traffic sign recognition and the decision-making ability of the autonomous driving system, and reduces the occurrence of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220115A_ABST
    Figure CN120220115A_ABST
Patent Text Reader

Abstract

The invention relates to the field of traffic sign recognition, and particularly discloses a traffic sign recognition method and system based on computer vision, and the method comprises the steps: firstly obtaining a digital map traffic sign image collected by a database, a real-time road traffic sign image collected by a camera, and a real-time road panoramic image collected by the camera; the deep learning technology is utilized to carry out feature extraction and correlation analysis on the three, a classification result is obtained through a classifier to obtain traffic sign category output, and then a decision instruction of an automatic driving system is output, so that the decision ability of automatic driving is improved, and traffic accidents are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of traffic sign recognition, and more specifically, to a traffic sign recognition method and system based on computer vision. Background Art

[0002] Traffic signs are signs set on roads to convey traffic information, give directions or warnings to drivers and pedestrians. Their main functions are to regulate traffic behavior, ensure safety and provide navigation instructions. Traffic signs are usually divided into types such as warnings, prohibitions, instructions and wayfinding signs. Each sign has clear graphics, colors and symbols to convey information quickly and accurately.

[0003] With the progress of driverless technology, driverless cars will play an increasingly important role in future life. To ensure the safe driving of driverless vehicles, it is necessary to accurately identify traffic signs set on the road and adjust the autonomous driving strategy according to the recognized content. However, currently, driverless cars mainly rely on image recognition technology to recognize traffic sign images, but the recognition accuracy is relatively low.

[0004] Therefore, a more optimized traffic sign recognition method and system based on computer vision is desired. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. Embodiments of this application provide a traffic sign recognition method and system based on computer vision. First, it acquires digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera. Using deep learning technology, it performs feature extraction and correlation analysis on the three, and obtains a classification result through a classifier to obtain a traffic sign category output, so as to output a decision instruction to the autonomous driving system, improve the decision-making ability of autonomous driving, and reduce the occurrence of traffic accidents.

[0006] According to one aspect of this application, there is provided a traffic sign recognition method based on computer vision, which includes: Acquire digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera; Extract traffic sign analysis and recognition feature vectors and real-time road panoramic feature vectors from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera; Based on the traffic sign analysis and recognition feature vectors and the real-time road panoramic feature vectors, obtain a traffic sign category output.

[0007] According to another aspect of the present application, there is provided a traffic sign recognition system based on computer vision, which includes: A traffic sign recognition data acquisition module for acquiring digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera; A traffic sign recognition data extraction module for extracting traffic sign analysis recognition feature vectors and real-time road panoramic feature vectors from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera; A traffic sign category output module for obtaining a traffic sign category output based on the traffic sign analysis recognition feature vector and the real-time road panoramic feature vector.

[0008] Compared with the prior art, a traffic sign recognition method and system based on computer vision provided by the present application first acquire digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera, use deep learning technology to perform feature extraction and correlation analysis on the three, obtain a classification result through a classifier, so as to obtain a traffic sign category output, and thus output a decision instruction to an autonomous driving system to improve the decision-making ability of autonomous driving and reduce the occurrence of traffic accidents. Description of the Drawings

[0009] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 It is a flowchart of a traffic sign recognition method based on computer vision according to an embodiment of the present application.

[0011] Figure 2 It is a flowchart of extracting traffic sign analysis recognition feature vectors and real-time road panoramic feature vectors from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera in the traffic sign recognition method based on computer vision according to an embodiment of the present application.

[0012] Figure 3A flowchart for extracting features from the digital map traffic sign images collected by the database in the computer vision-based traffic sign recognition method according to an embodiment of the present application to obtain digital map traffic sign feature vectors.

[0013] Figure 4 A block diagram of a computer vision-based traffic sign recognition system according to an embodiment of the present application. Detailed implementation manners

[0014] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0015] Figure 1 A flowchart of a computer vision-based traffic sign recognition method according to an embodiment of the present application. As Figure 1 shown, the computer vision-based traffic sign recognition method according to an embodiment of the present application includes: S110, obtaining digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera; S120, extracting traffic sign analysis and recognition feature vectors and real-time road panoramic feature vectors from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera; S130, obtaining a traffic sign category output based on the traffic sign analysis and recognition feature vectors and the real-time road panoramic feature vectors.

[0016] In the above computer vision-based traffic sign recognition system 100, in step S110, digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera are obtained. It should be understood that the digital map traffic sign images provide a standard traffic sign template for comparison with traffic signs in the actual scene; the real-time road traffic sign images capture traffic sign information on the current road; and the real-time road panoramic images provide a broader context environment for the recognition process to help the system better understand the position and meaning of traffic signs on the actual road. Among them, the digital map traffic sign images usually come from databases that contain a set of annotated standard traffic sign images. The real-time road traffic sign images are collected by cameras installed on autonomous vehicles. The real-time road panoramic images are also collected by cameras but focus on different areas. In this way, the accuracy of traffic sign recognition can be improved, and the adaptability of the autonomous driving system to complex traffic environments can be enhanced, thereby effectively reducing the occurrence of traffic accidents.

[0017] Specifically, traffic signs are indicators installed on roads to convey traffic information, guide directions, or issue warnings to drivers and pedestrians. Their main purpose is to regulate traffic behavior, ensure traffic safety, and provide navigation guidance. Traffic signs are generally divided into four types: warning, prohibition, indication, and wayfinding. Each sign has unique graphics, colors, and symbols to facilitate drivers and pedestrians to quickly and accurately understand the information. With the continuous development of autonomous driving technology, autonomous vehicles will play an increasingly important role in future daily life. To ensure the safe driving of autonomous vehicles, they need to accurately identify traffic signs on the road and adjust their driving strategies according to the recognition results. However, currently, autonomous vehicles mainly use image recognition technology to identify traffic signs, but there are still certain challenges in the recognition accuracy. Therefore, in the technical solution of this application, by obtaining digital map traffic sign images collected by the database, real-time road traffic sign images collected by the camera, and real-time road panoramic images collected by the camera, and combining deep learning technology, the types of traffic signs are output, and the results are transmitted to the autonomous driving system, thereby optimizing its decision-making process, enhancing the judgment ability of autonomous driving, and reducing the occurrence of traffic accidents.

[0018] In the above computer vision-based traffic sign recognition system 100, in step S120, traffic sign analysis and recognition feature vectors and real-time road panoramic feature vectors are extracted from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera. In this way, it is possible to realize the full-process automated processing from feature extraction, analysis, and association from multi-source data to final classification, providing a solid foundation for the decision-making of the autonomous driving system. This method not only improves the accuracy of traffic sign recognition but also enhances the system's adaptability to complex traffic environments, thereby effectively reducing the probability of traffic accidents.

[0019] Figure 2 It is a flowchart for extracting traffic sign analysis and recognition feature vectors and real-time road panoramic feature vectors from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera in the computer vision-based traffic sign recognition method according to an embodiment of the present application. As Figure 2As shown, in a specific embodiment of the present application, in step S120, extracting traffic sign analysis and recognition feature vectors and real-time road panorama feature vectors from the digital map traffic sign images collected from the database, the real-time road traffic sign images collected by the camera, and the real-time road panorama images collected by the camera includes: S121, extracting features from the digital map traffic sign images collected from the database to obtain digital map traffic sign feature vectors; S122, extracting features from the real-time road traffic sign images collected by the camera to obtain real-time road traffic sign tracking feature vectors; S123, passing the digital map traffic sign feature vectors and the real-time road traffic sign tracking feature vectors through a traffic sign difference analysis network to obtain the traffic sign analysis and recognition feature vectors; S124, extracting features from the real-time road panorama images collected by the camera to obtain the real-time road panorama feature vectors.

[0020] In step S121, extracting features from the digital map traffic sign images collected from the database to obtain digital map traffic sign feature vectors. It should be understood that the digital map traffic sign images are sourced from the database and usually contain annotated standard traffic sign templates. These template images have clear graphic, color, and symbol features and are important reference bases in the recognition process. However, the original images may be subject to noise interference or resolution limitations, so it is necessary to preprocess and extract features from them to generate more representative feature vectors. This process can not only enhance the image quality but also extract key visual features for subsequent classification and matching operations.

[0021] In step S122, extracting features from the real-time road traffic sign images collected by the camera to obtain real-time road traffic sign tracking feature vectors. It should be understood that the real-time road traffic sign images are collected by cameras installed on autonomous vehicles, and these images reflect the real traffic sign conditions in the current road environment. However, due to the complexity of the actual scenario (such as lighting changes, weather conditions, occlusion, or background interference), the original images may contain a large amount of noise or blurred information, and directly using these images for recognition will result in a high misjudgment rate. Therefore, it is necessary to extract features from the real-time road traffic sign images to highlight key information and remove redundant data. In this way, it can dynamically reflect the changes in traffic signs on the current road.

[0022] In the step S123, the digital map traffic sign feature vector and the real-time road traffic sign tracking feature vector are input into a traffic sign difference analysis network to obtain the traffic sign analysis and recognition feature vector. It should be understood that the digital map traffic sign feature vector is derived from the standard template images in the database and has the characteristics of high normalization, which can provide reliable reference information. The real-time road traffic sign tracking feature vector, on the other hand, is extracted from the actual scene and reflects the real traffic sign situation on the current road. However, due to factors such as lighting changes, occlusion, and weather conditions in the actual scene, the real-time feature vector may deviate from the standard template to some extent. Therefore, it is necessary to use the traffic sign difference analysis network to quantify the similarity or difference between the two and generate the final traffic sign analysis and recognition feature vector. Through the traffic sign difference analysis network, the system can generate a comprehensive traffic sign analysis and recognition feature vector. This feature vector not only retains the key information of the original feature vectors but also incorporates the comparison results between the two, thus more comprehensively describing the essential attributes of traffic signs. For example, when the real-time feature vector is highly similar to the standard template, the analysis and recognition feature vector will tend to a specific category; while when there are significant differences between the two, it may prompt the system to re-detect or adjust the recognition strategy. In addition, the design of the traffic sign difference analysis network also takes into account the dynamic characteristics in the autonomous driving scenario. For example, during vehicle driving, traffic signs may present different appearances due to angle changes or partial occlusion. Through comparative analysis, the system can better adapt to these changes and ensure the consistency and stability of the recognition results. At the same time, this network structure can also be extended to various types of traffic signs and is applicable to the recognition requirements in different driving environments. In this way, by quantifying the differences between the standard template and the actual scene, a more representative feature representation is generated, providing a reliable basis for subsequent classification and decision-making.

[0023] In a specific embodiment of the present application, the traffic sign difference analysis network is a traffic sign difference siamese network. It should be understood that the traffic sign difference siamese network is a deep learning model for comparing the similarity of two input data, which can efficiently evaluate the similarity or difference between the traffic sign feature vectors of the digital map and the real-time road traffic sign tracking feature vectors. Specifically, the core idea of the traffic sign difference siamese network is to process the two input data respectively through a convolutional neural network (CNN) with shared weights and generate corresponding feature representations. In the traffic sign recognition task, the two input branches of the network respectively receive the traffic sign feature vectors of the digital map and the real-time road traffic sign tracking feature vectors. After these two feature vectors are processed by the same convolutional layer, pooling layer and fully connected layer, high-level abstract feature representations are generated. Since the weights of the network are shared, it can ensure that the two branches extract features in the same way, so that the final outputs are comparable. Next, the traffic sign difference siamese network quantifies their differences by calculating the distance or similarity between the two feature representations. Commonly used similarity measurement methods include Euclidean distance, cosine similarity or contrast loss function, etc. For example, when the two feature vectors are very close, the Euclidean distance between them will be small, indicating that the traffic sign in the digital map is highly similar to the real-time road traffic sign; on the contrary, a larger distance may indicate a significant difference between the two. Based on this quantization result, the system can further determine the category of the traffic sign on the current road and use it as the decision-making basis for the autonomous driving system. It should be noted that traditional methods usually rely on a single feature extraction method and are easily affected by environmental interference. By introducing a siamese network for comparative analysis, the system can maintain stable performance under a wider range of conditions. At the same time, the training process of this network can be optimized through a large amount of labeled data to gradually adapt to different types of traffic signs and their variants.

[0024] In step S124, feature extraction is performed on the real-time road panoramic image collected by the camera to obtain the real-time road panoramic feature vector. It should be understood that compared with a single traffic sign image, the panoramic image can reflect the position, direction of the traffic sign in the actual scene and its relationship with the surrounding environment, thereby improving the accuracy and robustness of recognition. Specifically, the real-time road panoramic image is collected by a wide-angle or fisheye camera installed on the autonomous vehicle, and these images cover a wide field of view around the vehicle and contain rich environmental information. However, since panoramic images usually have a high resolution and a complex background, directly using the original images for analysis will result in a waste of computing resources and a reduction in processing efficiency. Therefore, it is necessary to perform feature extraction on the panoramic image to convert it into a compact and meaningful feature representation, that is, the real-time road panoramic feature vector. In this way, it can provide important auxiliary information for traffic sign recognition.

[0025] Figure 3 It is a flowchart for extracting features from the digital map traffic sign images collected by the database in the computer vision-based traffic sign recognition method according to an embodiment of the present application to obtain digital map traffic sign feature vectors. As Figure 3 shown, in a specific embodiment of the present application, the step S121 of extracting features from the digital map traffic sign images collected by the database to obtain digital map traffic sign feature vectors includes: S1211, performing noise reduction on the digital map traffic sign images collected by the database to obtain a denoised digital map traffic sign enhanced image; S1212, passing the denoised digital map traffic sign enhanced image through a digital map traffic sign convolutional neural network model serving as a filter to obtain the digital map traffic sign feature vectors.

[0026] In the step S1211, the digital map traffic sign image collected from the database is denoised to obtain a denoised and enhanced digital map traffic sign image. It should be understood that in practical applications, the digital map traffic sign image may have a certain degree of quality degradation due to noise interference during the data collection process (such as sensor noise, compression artifacts, or transmission errors). These noises will reduce the resolution and contrast of the image, affect the accuracy of feature extraction, and thus weaken the performance of the traffic sign recognition system. Therefore, through denoising processing, useless information in the image can be effectively removed, target features can be highlighted, and the overall performance of the system can be improved. Specifically, in a specific embodiment of the present application, the denoising process of the digital map traffic sign image generally includes the following steps: First, analyze the noise characteristics of the original image to determine its main source and distribution form. Common types of noise include Gaussian noise, salt-and-pepper noise, and blocking effect noise, etc. For different types of noise, corresponding denoising algorithms can be selected for processing. For example, for Gaussian noise, a deep learning model based on convolutional neural network (CNN) can be used for learning and removal; while for salt-and-pepper noise, traditional methods such as median filters can be used for processing. Next, the denoised image is further optimized to generate a denoised and enhanced digital map traffic sign image. The core task of this stage is to improve the visual effect and feature expression ability of the image through image enhancement techniques. Commonly used enhancement methods include histogram equalization, contrast stretching, and edge sharpening, etc. These techniques can significantly improve the brightness, contrast, and detail clarity of the image, making the key features of the traffic signs more prominent. In addition, advanced image processing algorithms (such as wavelet transform or adaptive filtering) can be combined to dynamically adjust the enhancement parameters according to the image content to ensure that no new distortion or blurring phenomenon is introduced while removing noise. The denoised and enhanced digital map traffic sign image has higher quality and stronger robustness, and can better meet the needs of subsequent feature extraction and classification. In a specific embodiment of the present application, denoising the digital map traffic sign image collected from the database to obtain a denoised and enhanced digital map traffic sign image includes: inputting the digital map traffic sign image collected from the database into the encoder of the denoising module, where the encoder uses convolutional layers to perform explicit spatial encoding on the digital map traffic sign image collected from the database to obtain image features; and inputting the image features into the decoder of the denoising module, where the decoder uses deconvolutional layers to perform deconvolution processing on the image features to obtain the denoised and enhanced digital map traffic sign image.

[0027] In the step S1212, the enhanced image of the digital map traffic sign after noise reduction is passed through a digital map traffic sign convolutional neural network model acting as a filter to obtain the digital map traffic sign feature vector. It should be understood that passing the enhanced image of the digital map traffic sign after noise reduction through a digital map traffic sign convolutional neural network (CNN) model acting as a filter can extract highly representative and discriminative feature representations from the image and reflect the key attributes of traffic signs (such as shape, color, and symbol, etc.). Compared with traditional manually designed feature methods, convolutional neural networks can automatically learn complex patterns in images and generate more robust and efficient feature representations, thus significantly improving the performance of traffic sign recognition systems. Specifically, the enhanced image of the digital map traffic sign is a high-quality image after noise reduction and enhancement processing, with its noise effectively removed and key features highlighted. However, the original image still contains a large amount of redundant information, and directly using it for classification or matching may lead to waste of computing resources and reduced efficiency. Therefore, it is necessary to extract features from the image through a convolutional neural network, converting high-dimensional pixel-level data into low-dimensional but information-rich feature vectors. In the implementation process, first, the enhanced image of the digital map traffic sign after noise reduction is input into the convolutional neural network model. The convolutional neural network consists of multiple convolutional layers, pooling layers, and fully connected layers, with each layer responsible for extracting features at different levels. The convolutional layer scans different regions of the image by applying a series of learnable filters (i.e., convolutional kernels) to capture low-level features such as local edges, textures, and shapes. As the depth of the network increases, subsequent convolutional layers gradually combine low-level features to generate higher-level abstract representations, such as the overall contour or specific symbol of a traffic sign. The pooling layer is used to reduce the spatial dimension of features while retaining the most important information. Common pooling operations include max pooling and average pooling, which can maintain the invariance of features (such as robustness to transformations like translation and scaling) while reducing computational complexity. Through multiple convolutional and pooling operations, the network gradually constructs a compact and expressive feature space. Finally, the feature map after convolutional and pooling processing is flattened and input into the fully connected layer to further integrate global information and generate the final digital map traffic sign feature vector. This feature vector not only contains the essential attributes of traffic signs but also has strong generalization ability and can adapt to different types of traffic signs and their variants. In this way, important features in the image are automatically extracted through deep learning technology, generating a structured and information-rich feature representation.In a specific embodiment of the present application, the enhanced image of the digital map traffic signs after noise reduction is passed through a digital map traffic sign convolutional neural network model serving as a filter to obtain the digital map traffic sign feature vector, including: using each layer of the digital map traffic sign convolutional neural network model serving as a filter to perform convolution processing, mean pooling processing based on the local feature matrix, and non-linear activation processing on the input data respectively during the forward pass of the layer, so as to output the digital map traffic sign feature vector by the last layer of the digital map traffic sign convolutional neural network model serving as a filter, wherein the input of the digital map traffic sign convolutional neural network model serving as a filter is the enhanced image of the digital map traffic signs after noise reduction.

[0028] In a specific embodiment of the present application, in step S122, feature extraction is performed on the real-time road traffic sign image collected by the camera to obtain a real-time road traffic sign tracking feature vector, including: S1221, passing the real-time road traffic sign image collected by the camera through a real-time road traffic sign target detection network to obtain multiple regions of interest of real-time road traffic signs; S1222, arranging the multiple regions of interest of real-time road traffic signs and passing them through a real-time road traffic sign feature encoder to obtain the real-time road traffic sign tracking feature vector.

[0029] In the step S1221, the real-time road traffic sign image collected by the camera is passed through a real-time road traffic sign target detection network to obtain multiple regions of interest of real-time road traffic signs. It should be understood that passing the real-time road traffic sign image collected by the camera through the real-time road traffic sign target detection network can quickly locate and screen out potential traffic sign positions from complex actual scenes. This method can significantly reduce the amount of data to be processed subsequently, improve the recognition efficiency, and provide a more accurate target range for feature extraction and classification tasks. Compared with directly processing the entire image, the target detection network can focus on the key regions in the image that may contain traffic signs, thereby improving the overall performance of the system. Specifically, the real-time road traffic sign images are collected by cameras installed on autonomous vehicles. These images usually contain rich background information and various interference factors, such as other road elements, weather conditions, or lighting changes, etc. In such a complex scene, directly analyzing the entire image may lead to waste of computing resources and an increase in the misjudgment rate. Therefore, it is necessary to use the target detection network to initially screen out the regions of interest that may contain traffic signs so that the subsequent steps can focus on these key regions for further processing. In a specific embodiment of the present application, first, after inputting the real-time road traffic sign image, the network extracts high-level feature representations of the image through a convolutional neural network (CNN). This process is similar to feature extraction, but its goal is to generate a global feature map for subsequent candidate region generation and classification operations. Next, the network uses a Region Proposal Network (RPN) to generate a series of candidate regions. The RPN is a lightweight sub-network that can slide a window on the feature map and predict whether each window may contain a traffic sign. The candidate regions it outputs are usually represented in the form of bounding boxes, and each bounding box is also attached with a confidence score to measure the possibility that the region contains a traffic sign. Subsequently, the target detection network performs further classification and regression operations on these candidate regions. The classification task aims to determine whether each candidate region actually contains a traffic sign, while the regression task is used to optimize the position and size of the bounding box to make it more accurately cover the target object. Commonly used real-time object detection algorithms include YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and Faster R-CNN, etc. These algorithms achieve a good balance between speed and accuracy. In this way, the target detection network finally generates multiple regions of interest of real-time road traffic signs. These regions not only define the specific positions of traffic signs but also provide preliminary category information (such as speed limit signs, no-entry signs, etc.). Compared with the original image, these regions of interest have higher relevance and lower redundancy, and can significantly reduce the computational complexity of subsequent feature extraction and classification.In a specific implementation of the present application, the real-time road traffic sign image collected by the camera is passed through a real-time road traffic sign target detection network to obtain multiple regions of interest of real-time road traffic signs, and the target detection formula is: ; where is the real-time road traffic sign image collected by the camera, is the anchor box, is the region of interest part corresponding to the track, represents classification, represents regression.

[0030] In the step S1222, the multiple real-time road traffic sign regions of interest are arranged and passed through a real-time road traffic sign feature encoder to obtain the real-time road traffic sign tracking feature vector. In this way, the key features of traffic signs can be extracted from complex actual scenes and converted into a compact and discriminative digital representation to better understand the distribution law of traffic signs in images, thereby improving the accuracy and efficiency of recognition. Specifically, the multiple real-time road traffic sign regions of interest generated by the object detection network may contain different types of traffic signs or partially overlapping objects. Directly processing these regions may lead to information redundancy or confusion, so they need to be arranged and optimized. The arrangement process usually sorts based on information such as the spatial position, size, and confidence score of the region of interest to ensure that the key regions most likely to contain traffic signs are processed first. For example, the arrangement order can be determined by calculating the center coordinates or bounding box area of each region, or high-probability target regions can be screened according to the confidence score. Next, the arranged regions of interest are input into the real-time road traffic sign feature encoder for processing. The feature encoder is a module specifically used to extract local features and generate a compact feature representation, usually implemented by a convolutional neural network (CNN) or other deep learning models. The encoder analyzes each region of interest one by one, extracts key features such as shape, color, and texture, and converts these features into a fixed-length feature vector. This process gradually constructs a high-level abstract representation through multiple convolutional operations while reducing the data dimension to improve computational efficiency. To further enhance the feature representation ability, the feature encoder can also incorporate an attention mechanism or adaptive pooling technology. For example, the attention mechanism can dynamically allocate weights to highlight the most relevant parts in the region of interest while suppressing the influence of irrelevant backgrounds; adaptive pooling can adjust the spatial resolution of the feature map according to task requirements to ensure that the dimensions of the output feature vectors are consistent. Finally, the feature vectors generated by the encoder are integrated into the real-time road traffic sign tracking feature vector, which not only retains the key information of each region of interest but also reflects the spatial relationship and context dependence between them. In a specific embodiment of the present application, arranging the multiple real-time road traffic sign regions of interest and passing them through the real-time road traffic sign feature encoder to obtain the real-time road traffic sign tracking feature vector includes: using each layer of the real-time road traffic sign feature encoder to perform convolutional processing, mean pooling processing based on the local feature matrix, and non-linear activation processing on the input data respectively during the forward pass of the layer to output the real-time road traffic sign tracking feature vector by the last layer of the real-time road traffic sign feature encoder, where the input of the real-time road traffic sign feature encoder is the multiple real-time road traffic sign regions of interest.

[0031] In a specific embodiment of the present application, step S124 of extracting features from the real-time road panoramic image collected by the camera to obtain the real-time road panoramic feature vector includes: S1241, passing the real-time road panoramic image collected by the camera through a real-time road panoramic feature extractor based on a spatial attention mechanism to obtain a real-time road panoramic feature map; S1242, performing downsampling on the real-time road panoramic feature map to obtain the real-time road panoramic feature vector.

[0032] In the step S1241, the real-time road panoramic image collected by the camera is passed through a real-time road panoramic feature extractor based on a spatial attention mechanism to obtain a real-time road panoramic feature map. It should be understood that passing the real-time road panoramic image collected by the camera through a real-time road panoramic feature extractor based on a spatial attention mechanism can extract the regions most relevant to the traffic sign recognition task from the complex panoramic image, and generate a highly discriminative and representative feature representation, so as to significantly reduce the influence of background noise and highlight key targets (such as traffic signs), thereby improving the accuracy of subsequent classification and decision-making. In addition, by dynamically allocating weights through the spatial attention mechanism, the system can better adapt to the complex changes in different driving scenarios. Specifically, the real-time road panoramic image is collected by a wide-angle or fisheye camera installed on an autonomous vehicle, and these images cover a wide field of view around the vehicle and contain rich environmental information. However, since panoramic images usually contain a large amount of irrelevant background (such as the sky, ground or other non-traffic sign objects), directly analyzing them will lead to waste of computing resources and an increase in the misjudgment rate. Therefore, a feature extractor based on a spatial attention mechanism is needed to filter out the key regions relevant to the task. The spatial attention mechanism is a deep learning technology, and its core idea is to automatically focus on the most important parts of the image by dynamically allocating weights, while suppressing the influence of irrelevant background. In this process, the real-time road panoramic image is first input into a convolutional neural network (CNN) for feature extraction to generate a multi-channel feature map. Each channel corresponds to a specific feature in the image (such as edges, textures or colors). Subsequently, the spatial attention module will perform weighted processing on these feature maps and calculate the importance score of each position. This score reflects whether this position is highly relevant to the current task (such as traffic sign recognition). To achieve this goal, the spatial attention mechanism usually adopts the following steps: First, a global pooling operation is performed on different dimensions (such as the channel dimension or the spatial dimension) of the feature map to generate a set of descriptive statistics. Then, a fully connected layer or a convolutional layer is used to model these statistics to predict the weight value of each position. Finally, these weight values are multiplied by the original feature map to generate an enhanced real-time road panoramic feature map. In this feature map, the regions related to traffic signs will be significantly enlarged, while the irrelevant background will be weakened. In this way, the key information in the original image can be retained, and the feature expression ability related to the task can be enhanced.In a specific embodiment of the present application, the real-time road panoramic image collected by the camera is passed through a real-time road panoramic feature extractor based on a spatial attention mechanism to obtain a real-time road panoramic feature map, including: using the convolutional encoding part of the real-time road panoramic feature extractor based on the spatial attention mechanism to perform deep convolutional encoding on the real-time road panoramic image collected by the camera to obtain an initial convolutional feature map; inputting the initial convolutional feature map into the real-time road panoramic feature extractor based on the spatial attention mechanism to obtain a spatial attention map; passing the spatial attention map through a Softmax activation function to obtain a spatial attention feature map; calculating the element-wise multiplication of the spatial attention feature map and the initial convolutional feature map to obtain the real-time road panoramic feature map. More specifically, inputting the initial convolutional feature map into the real-time road panoramic feature extractor based on the spatial attention mechanism to obtain a spatial attention map includes: performing average pooling and max pooling on the initial convolutional feature map along the channel dimension respectively to obtain an average feature matrix and a max feature matrix; concatenating and adjusting the channels of the average feature matrix and the max feature matrix to obtain a channel feature matrix; using the convolutional layer of the spatial attention feature map to perform convolutional encoding on the channel feature matrix to obtain a spatial attention map.

[0033] In step S1242, downsampling the real-time road panoramic feature map to obtain the real-time road panoramic feature vector. It should be understood that in the technical solution of the present application, through methods such as global pooling (such as global average pooling or max pooling), convolutional operations, or dimensionality reduction techniques (such as PCA), a high-resolution feature map is compressed into a low-dimensional feature vector. Among them, global average pooling extracts global statistical features by calculating the average value of each channel, and convolutional operations use a convolutional kernel with a stride greater than 1 to gradually reduce the size of the feature map, so as to remove redundant information (such as noise or background interference) and highlight the core features of traffic signs (such as color, shape, or symbol), thereby enhancing the robustness of the recognition system and ensuring accuracy and stability in complex scenarios.

[0034] In the above traffic sign recognition system 100 based on computer vision, step S130, based on the traffic sign analysis and recognition feature vector and the real-time road panoramic feature vector, obtaining a traffic sign category output, includes: S131, fusing the traffic sign analysis and recognition feature vector and the real-time road panoramic feature vector to obtain a traffic sign category judgment and classification feature vector; S132, performing an improvement on the traffic sign category judgment and classification feature vector based on the core prior information feature energy threshold truncation of the hidden space construction to obtain an optimized traffic sign category judgment and classification feature vector; S133, passing the optimized traffic sign category judgment and classification feature vector through a classifier to obtain a classification result, and the classification result is used to represent the traffic sign category.

[0035] In the step S131, the traffic sign analysis and recognition feature vector and the real-time road panoramic feature vector are fused to obtain a traffic sign category judgment and classification feature vector. In this way, by combining multi-source information, the accuracy and robustness of traffic sign recognition can be improved. Among them, the traffic sign analysis and recognition feature vector is derived from the comparative analysis of digital map traffic sign images and real-time road traffic sign images, and can reflect the attributes of traffic signs themselves (such as shape, color, symbol, etc.), while the real-time road panoramic feature vector extracts context information through panoramic images, providing the position, direction of traffic signs in the actual scene and their relationship with the surrounding environment. Fusing the two can make up for the limitations of a single feature vector and enable the system to better adapt to complex driving environments. Specifically, the fusion process first requires preprocessing of the two feature vectors to ensure that they have the same dimension or format. If the dimensions are different, they can be adjusted through dimensionality reduction techniques (such as PCA) or dimensionality increase operations (such as zero padding). Then, a suitable fusion method is selected to integrate the two feature vectors. Common methods include the concatenation method (connecting the two feature vectors in sequence into a higher-dimensional feature vector), the weighted summation method (generating a new feature vector through linear combination by assigning different weights to the two feature vectors, and the weights can be dynamically adjusted according to experimental results), and the fusion network based on deep learning (designing a dedicated neural network structure, such as a multi-input fully connected layer or an attention mechanism module, to let the model automatically learn the relationship between the two features and generate an optimal comprehensive feature representation).

[0036] In the step S132, the classification feature vector for traffic sign category judgment is improved by truncating the core prior information feature energy threshold based on the construction of the latent space to obtain an optimized classification feature vector for traffic sign category judgment. It should be understood that in the technical solution of the present application, considering that in the processing process at different stages, multiple feature extraction modules (such as the convolutional neural network for digital map traffic signs, the real-time road traffic sign target detection network, and the real-time road panoramic feature extractor) extract some similar or duplicate features from different image data. These features are not fully de-duplicated or screened during the final fusion process, resulting in multiple redundant features being incorporated into the same feature vector, thus affecting the classification performance. Especially when fusing images from different sources (such as digital maps collected from databases and real-time images collected by cameras), the same road scene or sign may be repeatedly described, increasing the computational burden without improving the recognition effect. Moreover, in the feature fusion process, different types of features (such as traffic sign features and road panoramic features) often have large differences in information volume and dimension. Directly fusing these features may lead to an imbalance in the dimension of the feature vector, causing some important features to be ignored, or some features to occupy too large or too small weights in the calculation, thus affecting the judgment ability of the classifier. Therefore, in the technical solution of the present application, the classification feature vector for traffic sign category judgment is improved by truncating the core prior information feature energy threshold based on the construction of the latent space to obtain an optimized classification feature vector for traffic sign category judgment.

[0037] Among them, improving the classification feature vector for traffic sign category judgment by truncating the core prior information feature energy threshold based on the construction of the latent space to obtain an optimized classification feature vector for traffic sign category judgment includes: First, extract the prior basis matrix of the traffic sign structure. It should be understood that extracting the prior basis matrix of the traffic sign structure is not merely an operation of obtaining information, but actually a key operation for structuring and computabilizing prior knowledge. The principle is to condense domain expertise, model design concepts, or macroscopic laws inverted from data into the mathematical form of the prior basis matrix of the traffic sign structure for the effective utilization of subsequent algorithms.

[0038] Then, perform core prior information feature energy threshold truncation on the prior basis matrix of the traffic sign structure to obtain a set of anchor points for the core prior information feature subspace of the traffic sign model, which is represented by the energy threshold truncation formula as; ; where represents the set of anchor points for the core prior information feature subspace of the traffic sign model, respectively represent the first, second, th anchor points for the core prior information feature subspace of the traffic sign model, represents the transpose operation, Indicates the truncation of the core prior information feature energy threshold, Indicates the prior basis matrix of the traffic sign structure, Indicates a diagonal matrix, Respectively represent the values at the first and the th positions on the diagonal of the diagonal matrix. It should be understood that the prior basis matrix of the traffic sign structure may contain redundant or high-dimensional information, and direct application may lead to an excessive computational burden and interference from non-critical information to the optimization process. Therefore, the principle of truncating the core prior information feature energy threshold is to denoise and refine the prior knowledge, similar to the filtering process in signal processing, retaining the main components and filtering out the redundancy.

[0039] Next, construct the traffic sign model core prior information feature subspace anchor hidden variable interaction matrix between the traffic sign category judgment classification feature vector and each traffic sign model core prior information feature subspace anchor in the set of traffic sign model core prior information feature subspace anchors to obtain a set of traffic sign model core prior information feature subspace anchor hidden variable interaction matrices, which is represented by the feature subspace anchor formula as; ; where, Represents the traffic sign category judgment classification feature vector, Represents the Linear transformation is performed on, and the feature vector after linear transformation has the same feature scale as the corresponding traffic sign model core prior information feature subspace anchor, Represents the th traffic sign model core prior information feature subspace anchor, Represents matrix multiplication, Represents the length of the traffic sign model core prior information feature subspace anchor, Represents the th traffic sign model core prior information feature subspace anchor hidden variable interaction matrix. It should be understood that the core of this step is to construct the interaction relationship between the traffic sign category judgment classification feature vector and the refined prior knowledge. Specifically, through hidden space mapping, the non-linear response mode of the traffic sign category judgment classification feature vector to different aspects of prior knowledge is learned. Essentially, the traffic sign model core prior information feature subspace anchor hidden variable interaction matrix is a re-encoding of the traffic sign category judgment classification feature vector from the perspective of prior knowledge, integrating the interpretation and processing of knowledge, and its function exceeds information association, realizing the directional enhancement of feature representation and the extraction of multi-perspective feature information.

[0040] Next, calculate the traffic sign prior information activation saliency factor of each traffic sign model core prior information feature subspace anchor point latent variable interaction matrix in the set of traffic sign model core prior information feature subspace anchor point latent variable interaction matrices to obtain a set of traffic sign prior information activation saliency factors, which is expressed by the activation saliency formula as; ; where represents the F-norm of the matrix, represents the th traffic sign prior information activation saliency factor. It should be understood that the principle of this step is to refine key information and streamline feature representation for each traffic sign model core prior information feature subspace anchor point latent variable interaction matrix. Just like generating an information summary, it extracts the summary or feature vector that can best represent the core information. The traffic sign prior information activation saliency factor needs to be representative and discriminative. It embodies the idea of feature selection and feature aggregation, retains the most informative part, and aggregates it into a concise activation saliency factor to achieve deeper compression and refinement. The function of the traffic sign prior information activation saliency factor is not only information compression, but more importantly, to improve the efficiency and robustness of the subsequent fusion process, reduce the data dimension, and reduce the computational burden, especially in the case of high-dimensional matrices.

[0041] Subsequently, based on the set of traffic sign prior information activation saliency factors, perform adaptive sparse aggregation on the set of traffic sign model core prior information feature subspace anchor point latent variable interaction matrices to obtain a traffic sign model prior information response projection coding matrix, which is expressed by the adaptive sparse aggregation formula as; ; where represents the normalized exponential function, represents the traffic sign model prior information response projection coding matrix. It should be understood that the core principle of the adaptive sparse aggregation lies in emphasizing an adaptive and selective prior information integration strategy. Sparsity reflects that not all prior information responses are equally important, and the fusion should be selective, focusing on more important responses, weakening or ignoring unimportant responses, improving the feature selection ability and generalization ability, and avoiding overfitting. Dynamics means that the weight or method of adaptive sparse aggregation is not fixed, but is adaptively adjusted according to the input data or model state, improving the flexibility and adaptability of the model. The essence of performing adaptive sparse aggregation on the set of traffic sign model core prior information feature subspace anchor point latent variable interaction matrices is to optimally combine the response information from different prior knowledge perspectives to form a traffic sign model prior information response projection coding matrix that comprehensively reflects the model prior response.

[0042] Finally, map the traffic sign category judgment classification feature vector to the feature space of the prior information response projection coding matrix of the traffic sign model to obtain the optimized traffic sign category judgment classification feature vector, which is represented by the mapping formula: ; where represents the optimized traffic sign category judgment classification feature vector. It should be understood that finally, by mapping the traffic sign category judgment classification feature vector to the feature space of the prior information response projection coding matrix of the traffic sign model, the optimized traffic sign category judgment classification feature vector is obtained, and the entire adaptation process is completed. The principle of this step is to use the feature space defined by the prior information response projection coding matrix of the traffic sign model to project the traffic sign category judgment classification feature vector into this space. Substantially, the prior information response projection coding matrix of the traffic sign model acts as a transformation matrix to perform a linear or non-linear mapping transformation on the traffic sign category judgment classification feature vector, so that it is embedded in the feature space integrating the model prior information. The optimized traffic sign category judgment classification feature vector better conforms to the constraints and guidance of the model prior, and realizes boundary adaptation on the feature manifold. Compared with the traffic sign category judgment classification feature vector, the optimized traffic sign category judgment classification feature vector is usually significantly improved in terms of expression ability, discriminability, and generalization, providing a better feature representation for subsequent machine learning tasks.

[0043] In the step S133, the optimized classification feature vector for traffic sign category judgment is input into a classifier to obtain a classification result, which is used to represent the traffic sign category. It should be understood that after the previous steps, the feature vector has integrated the attribute information of the traffic sign itself (such as shape, color, symbol) and the context information (such as position, direction, and surrounding scene). These optimized feature vectors provide structured and discriminative input data for the classifier. By processing these feature vectors through the classifier, they can be mapped to specific traffic sign categories (such as speed limit signs, no-entry signs, etc.), thereby completing the final recognition task. Specifically, the role of the classifier is to learn the boundaries between different categories based on the input feature vectors and accurately predict the category to which an unknown sample belongs in practical applications. Common classifiers include Support Vector Machine (SVM), Random Forest, Logistic Regression, and the Fully Connected Neural Network in deep learning. In the traffic sign recognition task, classifiers based on deep learning are usually used because they can better handle high-dimensional feature vectors and automatically learn complex non-linear relationships. The process of inputting the optimized classification feature vector for traffic sign category judgment into the classifier can be divided into the following steps: First, the feature vector is standardized or normalized to ensure that its numerical range meets the requirements of the classifier. Then, the feature vector is input into the input layer of the classifier. If it is a classifier based on deep learning, the input layer will pass the feature vector to the hidden layers (such as fully connected layers or multi-layer perceptrons), and these hidden layers gradually extract high-level abstract features through activation functions and weight matrices. Subsequently, the output layer calculates the probability values for each category through the softmax function or other classification functions, generating a probability distribution vector. Finally, the category with the highest probability is selected as the final classification result. This classification result not only provides clear input for the decision-making module of the autonomous driving system but also maintains high accuracy in complex environments.

[0044] In summary, in the embodiment of the present application, first, a digital map traffic sign image collected by a database, a real-time road traffic sign image collected by a camera, and a real-time road panoramic image collected by a camera are obtained. Using deep learning technology, feature extraction and correlation analysis are performed on the three, and a classification result is obtained through a classifier to obtain a traffic sign category output, so as to output a decision-making instruction to the autonomous driving system, improve the decision-making ability of autonomous driving, and reduce the occurrence of traffic accidents.

[0045] Figure 4 For the block diagram of the traffic sign recognition system based on computer vision according to the embodiment of the present application. As Figure 4As shown, the computer vision-based traffic sign recognition system 100 according to an embodiment of the present application includes: a traffic sign recognition data acquisition module 110, configured to acquire digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera; a traffic sign recognition data extraction module 120, configured to extract traffic sign analysis and recognition feature vectors and real-time road panoramic feature vectors from the digital map traffic sign images collected by the database, the real-time road traffic sign images collected by the camera, and the real-time road panoramic images collected by the camera; a traffic sign category output module 130, configured to obtain a traffic sign category output based on the traffic sign analysis and recognition feature vectors and the real-time road panoramic feature vectors.

[0046] Here, those skilled in the art can understand that the specific operations of each step in the above computer vision-based traffic sign recognition system have been described in detail in the description of the computer vision-based traffic sign recognition method referred to above, and thus, the repeated description thereof will be omitted. Figures 1 to 3 As described above, the computer vision-based traffic sign recognition system 100 according to an embodiment of the present application can be implemented in various terminal devices. In one example, the computer vision-based traffic sign recognition system 100 can be integrated into a terminal device as a software module and / or a hardware module. For example, the computer vision-based traffic sign recognition system 100 can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the computer vision-based traffic sign recognition system 100 can also be one of many hardware modules of the terminal device.

[0047] Alternatively, in another example, the computer vision-based traffic sign recognition system 100 and the terminal device can also be separate devices, and the computer vision-based traffic sign recognition system 100 can be connected to the terminal device through a wired and / or wireless network and transmit and interact information in accordance with a predefined data format.

[0048] Alternatively, in another example, the computer vision-based traffic sign recognition system 100 and the terminal device can also be separate devices, and the computer vision-based traffic sign recognition system 100 can be connected to the terminal device through a wired and / or wireless network and transmit and interact information in accordance with a predefined data format.

Claims

1. A traffic sign recognition method based on computer vision, characterized in that: include: Acquire digital map traffic sign images collected by a database, real-time road traffic sign images collected by a camera, and real-time road panoramic images collected by a camera; Extracting a traffic sign analysis and recognition feature vector and a real-time road panorama feature vector from the digital map traffic sign image collected by the database, the real-time road traffic sign image collected by the camera, and the real-time road panorama image collected by the camera; Based on the traffic sign analysis and recognition feature vector and the real-time road panorama feature vector, a traffic sign category output is obtained.

2. The traffic sign recognition method based on computer vision according to claim 1, characterized in that: Extracting a traffic sign analysis and recognition feature vector and a real-time road panorama feature vector from the digital map traffic sign image collected by the database, the real-time road traffic sign image collected by the camera, and the real-time road panorama image collected by the camera includes: Extracting features from the digital map traffic sign image collected from the database to obtain a digital map traffic sign feature vector; Performing feature extraction on the real-time road traffic sign image captured by the camera to obtain a real-time road traffic sign tracking feature vector; The digital map traffic sign feature vector and the real-time road traffic sign tracking feature vector are passed through a traffic sign difference analysis network to obtain the traffic sign analysis and recognition feature vector; Feature extraction is performed on the real-time road panoramic image captured by the camera to obtain the real-time road panoramic feature vector.

3. The method for traffic sign recognition based on computer vision according to claim 2, characterized in that: Extracting features from the digital map traffic sign image collected from the database to obtain a digital map traffic sign feature vector includes: Denoising the digital map traffic sign image collected from the database to obtain a denoised digital map traffic sign enhanced image; The denoised digital map traffic sign enhanced image is passed through a digital map traffic sign convolutional neural network model as a filter to obtain the digital map traffic sign feature vector.

4. The method for traffic sign recognition based on computer vision according to claim 3, characterized in that: Extracting features from the real-time road traffic sign image captured by the camera to obtain a real-time road traffic sign tracking feature vector includes: Passing the real-time road traffic sign image collected by the camera through a real-time road traffic sign target detection network to obtain a plurality of real-time road traffic sign regions of interest; The multiple real-time road traffic sign regions of interest are arranged and passed through a real-time road traffic sign feature encoder to obtain the real-time road traffic sign tracking feature vector.

5. The method for traffic sign recognition based on computer vision according to claim 4, characterized in that: The traffic sign difference analysis network is a traffic sign difference twin network.

6. The method for traffic sign recognition based on computer vision according to claim 5, characterized in that: Extracting features from the real-time road panoramic image captured by the camera to obtain the real-time road panoramic feature vector includes: The real-time road panoramic image collected by the camera is passed through a real-time road panoramic feature extractor based on a spatial attention mechanism to obtain a real-time road panoramic feature map; The real-time road panorama feature map is downsampled to obtain the real-time road panorama feature vector.

7. The method for traffic sign recognition based on computer vision according to claim 6, characterized in that: Based on the traffic sign analysis and recognition feature vector and the real-time road panorama feature vector, a traffic sign category output is obtained, including: Fusion of the traffic sign analysis and recognition feature vector and the real-time road panorama feature vector to obtain a traffic sign category determination classification feature vector; Performing a core prior information feature energy threshold truncation improvement based on latent space on the traffic sign category judgment classification feature vector to obtain an optimized traffic sign category judgment classification feature vector; The optimized traffic sign category judgment classification feature vector is passed through a classifier to obtain a classification result, and the classification result is used to represent the traffic sign category.

8. The method for traffic sign recognition based on computer vision according to claim 7, characterized in that: The traffic sign category judgment classification feature vector is improved by truncation of the core prior information feature energy threshold constructed based on the latent space to obtain an optimized traffic sign category judgment classification feature vector, including: Extract traffic sign structure prior basis matrix; Performing a core prior information feature energy threshold truncation on the traffic sign structure prior basis matrix to obtain a set of anchor points of the core prior information feature subspace of the traffic sign model; Constructing a traffic sign model core prior information feature subspace anchor point latent variable interaction matrix between the traffic sign category judgment classification feature vector and each traffic sign model core prior information feature subspace anchor point in the set of traffic sign model core prior information feature subspace anchor points to obtain a set of traffic sign model core prior information feature subspace anchor point latent variable interaction matrices; Calculating the traffic sign prior information activation significance factor of each traffic sign model core prior information feature subspace anchor point latent variable interaction matrix in the set of the traffic sign model core prior information feature subspace anchor point latent variable interaction matrix to obtain a set of traffic sign prior information activation significance factors; Based on the set of traffic sign prior information activation significance factors, adaptively sparsely aggregate the set of anchor point latent variable interaction matrices of the traffic sign model core prior information feature subspace to obtain a traffic sign model prior information response projection coding matrix; The traffic sign category judgment classification feature vector is mapped to the feature space of the traffic sign model prior information response projection coding matrix to obtain the optimized traffic sign category judgment classification feature vector.

9. A traffic sign recognition system based on computer vision, characterized in that: include: Traffic sign recognition data acquisition module, used to acquire digital map traffic sign images collected by the database, real-time road traffic sign images collected by the camera, and real-time road panoramic images collected by the camera; A traffic sign recognition data extraction module, used to extract traffic sign analysis and recognition feature vectors and real-time road panorama feature vectors from the digital map traffic sign image collected by the database, the real-time road traffic sign image collected by the camera, and the real-time road panorama image collected by the camera; The traffic sign category output module is used to obtain a traffic sign category output based on the traffic sign analysis and recognition feature vector and the real-time road panorama feature vector.

10. The computer vision-based traffic sign recognition system according to claim 9, characterized in that: The traffic sign recognition data extraction module is used to: Extracting features from the digital map traffic sign image collected from the database to obtain a digital map traffic sign feature vector; Performing feature extraction on the real-time road traffic sign image captured by the camera to obtain a real-time road traffic sign tracking feature vector; The digital map traffic sign feature vector and the real-time road traffic sign tracking feature vector are passed through a traffic sign difference analysis network to obtain the traffic sign analysis and recognition feature vector; Feature extraction is performed on the real-time road panoramic image captured by the camera to obtain the real-time road panoramic feature vector.