Traffic sign image recognition model construction and recognition method
By combining the mixed network structure and feature pyramid network with sparse self-attention units, the error detection and missed detection problems of traffic sign recognition models in complex backgrounds are solved, and high-precision recognition in complex environments is achieved.
Patent Information
- Application Number
- CN202510447467.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing traffic sign recognition models are difficult to accurately capture traffic sign information under complex backgrounds, especially in the background, reflective or extreme conditions, which are prone to misjudgment or missed inspection.
A hybrid network structure is adopted, combining multi-scale convolutional units and sparse self-attention units for feature extraction, an extended feature pyramid network is used for positioning and classification, and through transfer learning and dynamic learning rate optimization, Focal-EIoU loss function is introduced to improve recognition accuracy.
It significantly reduces mis-detection and missed detection caused by complex light and shadow and reflective interference, and improves the recognition accuracy and robustness of traffic signs in complex environments.
Smart Images

Figure CN120388349A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic sign image recognition, and particularly to a method for constructing and recognizing a traffic sign image recognition model. Background Art
[0002] In urban traffic, traffic sign recognition models mostly rely on classical convolutional networks or artificial feature extraction methods. For traffic signs in areas with cluttered backgrounds, it becomes difficult to extract the edges of the signs. Here, a cluttered background means that there are billboards, walls, or reflective buildings around when the image acquisition device captures images, and the colors are close to those of the signs.
[0003] When traditional methods perform data preprocessing, they mostly strengthen the distinction effect between the target and the background through means such as color space conversion, histogram equalization, and edge enhancement, and cannot effectively solve the problem of highly overlapping features between signs and backgrounds in complex environments. For artificial feature extraction algorithms designed based on HOG, LBP, etc., when facing the lack of details or local distortions caused by reflection and shadow, misjudgment or missed detection is likely to occur. In addition, when traditional convolutional networks capture traffic signs, there are problems such as insufficient scale and insufficiently detailed feature expression. In night, haze, or other extreme conditions, the detailed information of the signs is more likely to be submerged by background noise. Therefore, today with the continuous development of autonomous driving and intelligent transportation systems, how to accurately capture traffic sign information in complex backgrounds has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] The present invention provides a method for constructing and recognizing a traffic sign image recognition model to solve the problem that the current traffic sign image recognition method has poor recognition effect on traffic signs in areas with cluttered backgrounds.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: An embodiment of the present invention provides a method for constructing and recognizing a traffic sign image recognition model, which includes: Step S1, perform image data enhancement on the traffic sign image to obtain an enhanced image including low-contrast, reflective, and complex light and shadow regions; the image data enhancement includes locally adjusting the brightness and performing color separation processing on the reflective regions in the image; Step S2, based on the enhanced image, use a hybrid network structure for multi-scale feature extraction, and the hybrid network structure integrates multi-scale convolutional units and sparse self-attention units; Step S3, use an extended feature pyramid network structure to fuse the extracted multi-scale features, and use an improved detection head to locate and classify the target; The detection head includes a sub-detection module suitable for small target detection. The sub-detection module refines the target boundary information based on the local features obtained by focusing with the sparse self-attention unit. Step S4: Adopt transfer learning technology to transfer the weights pre-trained on a large-scale image dataset to the hybrid network structure, freeze the parameters of the low-level convolutional units, and fine-tune the high-level fully connected units. And introduce a dynamic learning rate decay strategy and adopt a boundary regression loss function based on the focal mechanism to further improve the traffic sign recognition accuracy of the network in complex environments. Step S5: Output the traffic sign recognition result obtained after preprocessing in step S1, feature extraction in step S2, feature fusion and classification in step S3, and model optimization in step S4; the recognition result reflects the true category information of the traffic sign under complex backgrounds, occlusion, and reflection conditions.
[0007] As a preferred scheme of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: the image data augmentation further includes color space conversion, histogram equalization, local contrast enhancement, and reflection suppression; to reduce background interference and retain the inherent color features of traffic signs.
[0008] As a preferred scheme of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: the multi-scale convolutional unit is used to capture local texture and geometric shape information in the image. The sparse self-attention unit is used to focus on local features with high signal-to-noise ratio in the image features. The hybrid network structure effectively suppresses information redundancy caused by the high degree of color fusion between traffic signs and the background and complex light and shadow interference by performing feature fusion at different depth levels.
[0009] As a preferred scheme of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: in the hybrid network structure, the multi-scale convolutional unit and the sparse self-attention unit are arranged alternately to achieve collaborative extraction of traffic sign edge and detail information at different scales.
[0010] As a preferred scheme of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: in step S2, the method of performing multi-scale feature extraction using the hybrid network structure based on the enhanced image is as follows: The multi-scale convolutional unit performs convolutional operations on the input image at different scales, and its output feature is expressed as: , , wherein, represents the output feature of the multi-scale convolutional unit. Indicates the scale index, Indicates the total number of scales, Indicates the weight factor of the th scale, Indicates the th scale's feature obtained by performing a convolution operation on the input image , Indicates the convolution kernel weight of the th scale, Indicates the bias of the convolution kernel of the th scale; Indicates the input enhanced image feature; The sparse self-attention unit focuses on the global information of the input feature, linearly maps the input enhanced image feature to obtain the query, key, and value matrices. The formula is: , where, Indicates the query matrix, Indicates the key matrix, Indicates the value matrix, Indicates the enhanced image feature, Indicates the query mapping weight, Indicates the key mapping weight, Indicates the value mapping weight; Calculate the output feature through the scaled dot-product attention mechanism and the sparse mask . The formula is: , where, Indicates the output feature of the sparse self-attention unit, Indicates the normalization operation, Indicates the transpose of the key matrix, Indicates the dimension of the query vector, Indicates the normalization factor, Indicates element-wise multiplication, Indicates the sparse mask used to eliminate redundant information; In the hybrid network structure, the two types of units achieve feature fusion at different levels in an alternating arrangement. The fusion process is expressed as: , where, Indicates the output feature of the th layer, Indicates the fusion weight of the multi-scale convolution unit at the current layer, Indicates the result of applying the multi-scale convolution operation to the output feature of the previous layer, represents the fusion weight of the sparse self-attention unit in the current layer, represents the output features of the previous layer after applying the sparse self-attention operation, represents the current network level, represents the output features of the previous layer.
[0011] As a preferred solution of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: in step S3, the improved detection module adopted in the detection head is based on the extended feature pyramid network EFPN, and a sub-detection layer for small target recognition is introduced in its low layer.
[0012] As a preferred solution of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: in step S3, an improved detection head based on the extended feature pyramid network EFPN is used to locate and classify the target: EFPN constructs a multi-layer feature map fusion structure, and performs bidirectional fusion of bottom-up and top-down on the extracted multi-scale features. The fusion process is described as: , wherein, represents the th layer feature map after fusion, represents the th layer original feature map, represents the horizontal convolution transformation on the th layer feature map, and the calculation method is: , wherein, represents the th layer horizontal convolution kernel weight, represents its bias represents the upsampling operation on the fusion features of the previous layer , is the layer index, represents the total number of feature layers; A multi-scale enhancement module is introduced to the lowest layer feature in EFPN, and its formula is: , wherein, represents the low layer feature after multi-scale enhancement, represents the original low layer feature, represents the th type of convolution kernel, represents the number of scales in the enhancement module, represents the activation function, is the scale index.
[0013] As a preferred solution of the method for constructing and recognizing a traffic sign image recognition model according to the present invention, wherein: in step S3, the detection head uses an anchor box adjustment strategy to locate the candidate region, and the initial anchor box is set as: , wherein, represents the initial anchor box, and are respectively the horizontal and vertical coordinates of the center of the anchor box, and are respectively the width and height of the anchor box; The detection head dynamically adjusts the initial anchor box by predicting the offset , and the adjusted anchor box is: , wherein, and represent the horizontal and vertical translation offsets, and represent the logarithmic scale offsets of the width and height, is the exponential function; Design a small target detection layer in EFPN. This sub-detection module takes the low-level features enhanced by multi-scale as input, and its processing process is: , wherein, represents the small target detection result output by the sub-detection module, represents the sub-detection module operation, including local feature refinement and boundary regression. At the same time, the boundary regression optimization based on the focus mechanism is adopted inside the module, and its calculation formula is: , wherein, represents the optimized boundary regression result, represents the boundary regression function based on the focus mechanism, which uses the low-level features and anchor box information to adjust the regression loss simultaneously.
[0014] As a preferred solution of the method for constructing and recognizing a traffic sign image recognition model according to the present invention, wherein: in step S4, the loss function based on the focus mechanism used is the Focal-EIoU loss function, which effectively reduces the false detection and missed detection phenomena caused by the overlap of traffic signs and background features while ensuring the accuracy of the bounding box regression.
[0015] As a preferred solution of the traffic sign image recognition model construction and recognition method described in the present invention, wherein: in step S4, when using transfer learning technology, the weights pre-trained on a large-scale image dataset are loaded into the hybrid network structure, where the parameters of the low-level convolutional units are kept frozen, and the parameters of the high-level fully connected units are fine-tuned to adapt to the requirements of the traffic sign recognition task; at the same time, a dynamic learning rate decay strategy is introduced to gradually reduce the learning rate; In step S4, the FocalEIoU loss function based on the focal mechanism is used to optimize the boundary regression. Its design not only ensures the accuracy of the bounding box regression but also effectively reduces the false detection and missed detection phenomena caused by background interference. The FocalEIoU loss function adds a focal modulation term to the traditional IoU loss, assigns a higher loss weight to the low-IoU candidate boxes, and the calculation method is as follows: , where, represents the FocalEIoU loss, represents the balance coefficient for adjusting the focal term, is a positive real number, represents the complement value of the intersection over union (IoU) of the candidate box and the ground truth box, represents the ratio of their intersection to the union, represents the adjustment parameter in the focal mechanism, which is used to amplify the loss of the low-IoU candidate boxes, , represents the center of the candidate box and the center of the ground truth box the square of the Euclidean distance between them, and respectively represent the center coordinates of the candidate box and the ground truth box, represents the diagonal length of the minimum bounding box of the candidate box and the ground truth box, represents the width of the candidate box, represents the width of the ground truth box, represents the width of the minimum bounding box, which is used to normalize the width error, represents the height of the candidate box, represents the height of the ground truth box, represents the height of the minimum bounding box, which is used to normalize the height error; Compared with the traditional CIoU loss function, the FocalEIoU loss introduces a focal modulation term on the basis of the original center distance and ratio difference constraints , which can reduce the attention to easily regressed targets during the training process, enabling the network to focus more on difficult-to-locate candidate bounding boxes; at the same time, the width and height differences are normalized in the form of the squared difference. Compared with the ratio penalty term in CIoU, this design describes the deviation of the bounding box size more intuitively, is more sensitive especially in scenarios with severe background interference, and effectively reduces the risk of false detection and missed detection caused by the overlap between the candidate bounding box and the background features.
[0016] The beneficial effects of the present invention are as follows: In the present invention, multi-scale convolutional units and sparse self-attention units are alternately arranged to collaboratively extract local texture and global semantic information, and can still highlight key features when the signs and the background are highly fused; the detection head realizes the fine positioning of small targets through an extended feature pyramid and a specially designed small target detection layer; the transfer learning technology uses pre-trained weights to ensure the stability of the underlying features, while fine-tuning the high layer, supplemented by a dynamic learning rate and the Focal-EIoU loss function, significantly reducing false detection and missed detection caused by complex light and shadow and reflection interference.
[0017] The present invention effectively addresses the problem of feature loss caused by background and light interference in traditional models in actual traffic scenarios, improves the positioning and classification accuracy of traffic signs, and enhances the robustness of the model in complex and variable environments. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a schematic flowchart of the traffic sign image recognition model construction and recognition method in Embodiment 1. Detailed Embodiments
[0020] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification.
[0021] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0022] Second, the "one embodiment" or "embodiment" mentioned herein refers to specific features, structures, or characteristics that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that are mutually exclusive with other embodiments.
[0023] Embodiment 1, referring to Figure 1 , this embodiment provides a traffic sign image recognition model construction and recognition method, including the following steps: Step S1, perform image data enhancement on the traffic sign image to obtain an enhanced image containing low-contrast, reflective, and complex light and shadow regions; the image data enhancement includes local brightness adjustment and color separation processing on the reflective regions in the image; The image data enhancement also includes color space conversion, histogram equalization, local contrast enhancement, and reflection suppression; to reduce background interference and retain the inherent color characteristics of traffic signs; Step S2, based on the enhanced image, use a hybrid network structure for multi-scale feature extraction, and the hybrid network structure integrates multi-scale convolutional units and sparse self-attention units; The multi-scale convolutional unit is used to capture local texture and geometric shape information in the image, The sparse self-attention unit is used to focus on local features with high signal-to-noise ratio in the image features; The hybrid network structure effectively suppresses information redundancy caused by the high fusion of traffic signs and background colors and complex light and shadow interference through feature fusion at different depth levels; In the hybrid network structure, the multi-scale convolutional units and sparse self-attention units are arranged alternately to achieve collaborative extraction of traffic sign edge and detail information at different scales; In step S2, the method of using a hybrid network structure for multi-scale feature extraction based on the enhanced image is: The multi-scale convolutional unit performs convolutional operations on the input image at different scales, and its output feature is expressed as: , , Wherein, represents the output feature of the multi-scale convolutional unit, represents the scale index, represents the total number of scales, represents the th weight factor, represents the th scale's convolution operation on the input image to obtain the feature, represents the th scale convolution kernel weight, Denote the bias of the convolutional kernel of the th scale, and denote the input enhanced image features; The sparse self-attention unit focuses on the global information of the input features, and linearly maps the input enhanced image features to obtain the query, key, and value matrices. The formula is: , where denotes the query matrix, denotes the key matrix, denotes the value matrix, denotes the enhanced image features, denotes the query mapping weight, denotes the key mapping weight, denotes the value mapping weight; Calculate the output features through the scaled dot-product attention mechanism and the sparse mask . The formula is: , where denotes the output features of the sparse self-attention unit, denotes the normalization operation, denotes the transpose of the key matrix, denotes the dimension of the query vector, denotes the normalization factor, denotes the element-wise multiplication, denotes the sparse mask, which is used to eliminate redundant information; Specifically, the sparse self-attention unit selects the local features with high signal-to-noise ratio in the image through the linear mapping and the scaled dot-product attention mechanism, and at the same time uses the sparse mask to reduce the interference of background noise. This module re-weights the local features globally, highlights the important regions, and thus improves the network's ability to recognize the detailed information of traffic signs in complex environments; In the hybrid network structure, the two types of units achieve feature fusion at different levels in an alternating arrangement. The fusion process is expressed as: , where denotes the output features of the th layer, denotes the fusion weight of the multi-scale convolutional unit in the current layer, Represents the output features of the previous layer The result after applying multi-scale convolution operations Represents the fusion weights of the sparse self-attention unit at the current layer Represents the output features of the previous layer The result after applying sparse self-attention operations Represents the current network layer Represents the output features of the previous layer; Specifically, by alternately arranging the two types of units, the hybrid network structure can effectively integrate local details and global semantic information at each layer, allocate fusion weights to balance the advantages of different operations, so that the entire feature extraction process can reduce information redundancy while enhancing the ability to depict traffic sign features in complex backgrounds; Step S3: Use an extended feature pyramid network structure to fuse the extracted multi-scale features, and use an improved detection head to locate and classify the targets; The detection head includes a sub-detection module suitable for small target detection. The sub-detection module refines the target boundary information based on the local features obtained by focusing on the sparse self-attention unit; In step S3, the improved detection module used in the detection head is based on the extended feature pyramid network EFPN, and a sub-detection layer for small target recognition is introduced in its lower layer; In step S3, an improved detection head based on the extended feature pyramid network EFPN is used to locate and classify the targets: EFPN constructs a multi-layer feature map fusion structure to perform bidirectional fusion of the extracted multi-scale features from bottom to top and from top to bottom. The fusion process is described as: , Among them, Represents the layer feature map after fusion, Represents the layer original feature map, Represents the horizontal convolution transformation of the layer feature map. The calculation method is: , Among them, Represents the layer horizontal convolution kernel weight, Represents its bias, Represents the upsampling operation on the fusion features of the previous layer , is the layer index, Represents the total number of feature layers; In EFPN, a multi-scale enhancement module is introduced to the lowest layer features. The formula is: , Among them, represents the low-level features after multi-scale enhancement, represents the original low-level features, represents the th kind of convolutional kernel, represents the number of scales in the enhancement module, represents the activation function, is the scale index; In step S3, the detection head uses the anchor box adjustment strategy to locate the candidate regions. Let the initial anchor box be: , Among them, represents the initial anchor box, and are the horizontal and vertical coordinates of the center of the anchor box respectively, and are the width and height of the anchor box respectively; The detection head dynamically adjusts the initial anchor box by predicting the offset . The adjusted anchor box is: , Among them, and represent the horizontal and vertical translation offsets, and represent the logarithmic scale offsets of the width and height, is the exponential function; Design a small object detection layer in EFPN. This sub-detection module takes the low-level features after multi-scale enhancement as input, and its processing process is: , Among them, represents the small object detection result output by the sub-detection module, represents the operations of the sub-detection module, including local feature refinement and boundary regression. At the same time, the module internally adopts boundary regression optimization based on the focus mechanism, and its calculation formula is: , among which, represents the optimized boundary regression result, represents the boundary regression function based on the focus mechanism, which uses both low-level features and anchor box information to adjust the regression loss; Specifically, EFPN is used to achieve multi-level feature fusion, effectively integrating the information of each layer through the bottom-up and top-down bidirectional paths, retaining both high-level semantic information and strengthening the low-level detail description. The introduced multi-scale enhancement module magnifies and refines the low-level features, improving the response ability of small objects; During the anchor box adjustment process, the candidate boxes are finely corrected by dynamically predicting the offsets, enabling the traffic signs to be accurately located even in complex backgrounds. The specially designed small object detection layer combined with the focal mechanism further reduces the influence of background noise and effectively improves the robustness of small object detection. The overall detection head structure fully exploits the complementary advantages of each layer's features and realizes the tight coupling of the localization and classification tasks. Step S4: Adopt transfer learning technology to transfer the weights pre-trained on a large-scale image dataset to the hybrid network structure, freeze the parameters of the lower-level convolutional units, and fine-tune the higher-level fully connected units. And introduce a dynamic learning rate decay strategy and use a boundary regression loss function based on the focal mechanism to further improve the recognition accuracy of traffic signs in complex environments. In step S4, the loss function based on the focal mechanism used is the Focal-EIoU loss function. This loss function ensures the accuracy of the bounding box regression while effectively reducing the false detection and missed detection phenomena caused by the overlap between traffic signs and background features. In step S4, when using transfer learning technology, the weights pre-trained on a large-scale image dataset are loaded into the hybrid network structure. Among them, the parameters of the lower-level convolutional units are kept frozen, and the parameters of the higher-level fully connected units are fine-tuned to adapt to the requirements of the traffic sign recognition task. At the same time, a dynamic learning rate decay strategy is introduced to gradually reduce the learning rate. In step S4, the FocalEIoU loss function based on the focal mechanism is used to optimize the boundary regression. Its design not only ensures the accuracy of the bounding box regression but also can effectively reduce the false detection and missed detection phenomena caused by background interference. The FocalEIoU loss function adds a focal modulation term to the traditional IoU loss, assigns a higher loss weight to the low-IoU candidate boxes, and the calculation method is: , where, represents the FocalEIoU loss, represents the balance coefficient for adjusting the focal term, is a positive real number, represents the complement value of the intersection over union (IoU) between the candidate box and the ground truth box, represents the ratio of their intersection to their union, represents the adjustment parameter in the focal mechanism, which is used to amplify the loss of low-IoU candidate boxes, , represents the center of the candidate box and the center of the ground truth box the square of the Euclidean distance between them, and represent the center coordinates of the candidate box and the ground truth box respectively, represents the diagonal length of the minimum bounding box of the candidate box and the ground truth box. represents the width of the candidate box. represents the width of the ground truth box. represents the width of the minimum bounding box, which is used to normalize the width error. represents the height of the candidate box. represents the height of the ground truth box. represents the height of the minimum bounding box, which is used to normalize the height error. Compared with the traditional CIoU loss function, the FocalEIoU loss introduces a focal modulation term on the basis of the original center distance and ratio difference constraints. It can reduce the attention to easy-to-regress targets during training, so that the network focuses more on candidate boxes that are difficult to locate. At the same time, the width and height differences are normalized in the form of the square difference. Compared with the ratio penalty term in CIoU, this design describes the size deviation of the bounding box more intuitively, especially in scenes with severe background interference, and effectively reduces the risk of false detection and missed detection caused by the overlap between the candidate box and the background features. Specifically, through the transfer learning strategy, the pre-trained weights are applied to the hybrid network structure, which not only ensures the stability of the underlying feature extraction, but also enhances the task adaptability by fine-tuning the high-level parameters. The FocalEIoU loss function uses the focal modulation term to impose a greater penalty on candidate boxes with low IoU, which helps the model to achieve more accurate bounding box regression on difficult samples. The width and height normalization error term further refines the size matching effect, so as to overall improve the accuracy of traffic sign localization and classification in complex backgrounds and severely disturbed scenes. Step S5: Output the traffic sign recognition result obtained after preprocessing in step S1, feature extraction in step S2, feature fusion and classification in step S3, and model optimization in step S4; the recognition result reflects the true category information of the traffic sign under complex backgrounds, occlusion, and reflection conditions.
[0024] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for constructing and recognizing a traffic sign image recognition model, characterized in that: Including, Step S1: Perform image data augmentation on the traffic sign image to obtain an enhanced image; the image data augmentation includes locally adjusting the brightness and performing color separation processing on the reflective area in the image; Step S2: Based on the enhanced image, use a hybrid network structure to perform multi-scale feature extraction, and the hybrid network structure integrates multi-scale convolutional units and sparse self-attention units; Step S3: Use an extended feature pyramid network structure to fuse the extracted multi-scale features, and use an improved detection head to locate and classify the targets; The detection head includes a sub-detection module suitable for small target detection, and the sub-detection module refines the target boundary information based on the local features obtained by focusing through the sparse self-attention unit; Step S4: Adopt transfer learning technology to transfer the weights pre-trained on a large-scale image dataset to the hybrid network structure, freeze the parameters of the low-level convolutional units and fine-tune the high-level fully connected units; And introduce a dynamic learning rate decay strategy and adopt a boundary regression loss function based on the focal mechanism; Step S5: Output the traffic sign recognition result obtained after the preprocessing in Step S1, feature extraction in Step S2, feature fusion and classification in Step S3, and model optimization in Step S4.
2. The traffic sign image recognition model construction and recognition method according to claim 1, characterized in that: The image data augmentation also includes color space conversion, histogram equalization, local contrast enhancement, and reflection suppression.
3. The traffic sign image recognition model construction and recognition method according to claim 1, characterized in that: The multi-scale convolutional unit is used to capture the local texture and geometric shape information in the image, The sparse self-attention unit is used to focus on the local features with high signal-to-noise ratio in the image features.
4. The method for constructing and recognizing a traffic sign image recognition model according to claim 3, wherein: In the hybrid network structure, the multi-scale convolutional units and the sparse self-attention units are arranged alternately.
5. The method for constructing and recognizing a traffic sign image recognition model according to claim 4, characterized in that: In Step S2, the method of using a hybrid network structure to perform multi-scale feature extraction based on the enhanced image is as follows: The multi-scale convolutional unit performs convolutional operations on the input image at different scales, and its output feature representation is: , , Among them, represents the output feature of the multi-scale convolution unit, represents the scale index, represents the total number of scales, represents the weight factor of the th scale, represents the feature obtained by convolving the input image with the th scale convolution kernel, represents the bias of the th scale convolution kernel, and represents the input enhanced image feature; The sparse self-attention unit focuses on the global information of the input features, and enhances the input image features Linear mapping is performed to obtain query, key, and value matrices, and the formula is as follows: , Among them, represents the query matrix, represents the key matrix, represents the value matrix, represents the enhanced image feature, represents the query mapping weight, represents the key mapping weight, represents the value mapping weight; Calculate the output features through the scaled dot - product attention mechanism and the sparse mask The formula is as follows: , where, represents the output feature of the sparse self-attention unit, represents the normalization operation, represents the transpose of the key matrix, represents the dimension of the query vector, represents the normalization factor, represents the element-wise multiplication, represents the sparse mask for removing redundant information; In the hybrid network structure, the two types of units achieve feature fusion at different levels in an alternating arrangement manner, and its fusion process is expressed as: , Among them, represents the output feature of the -th layer, represents the fusion weight of the multi-scale convolution unit in the current layer, represents the result after applying the multi-scale convolution operation to the output feature of the previous layer, represents the fusion weight of the sparse self-attention unit in the current layer, represents the output feature of the previous layer after applying the sparse self-attention operation, represents the current network level, represents the output feature of the previous layer.
6. The traffic sign image recognition model construction and recognition method according to claim 1, characterized in that: In Step S3, the improved detection module adopted in the detection head is based on the extended feature pyramid network EFPN, and a sub-detection layer for small target recognition is introduced in its low layer.
7. The method for constructing and recognizing a traffic sign image recognition model according to claim 6, wherein: In Step S3, use an improved detection head based on the extended feature pyramid network EFPN to locate and classify the targets: EFPN constructs a multi-layer feature map fusion structure to perform two-way fusion of the extracted multi-scale features from bottom to top and from top to bottom, and the fusion process is described as: , Among them, represents the feature map of the th layer after fusion, represents the original feature map of the th layer, represents performing a horizontal convolution transformation on the feature map of the th layer, and the calculation method is as follows: , Among them, represents the weight of the horizontal convolution kernel of the th layer, and represents its bias. represents the upsampling operation on the fused features of the previous layer, is the layer index, and represents the total number of feature layers. In EFPN, a multi-scale enhancement module is introduced to the lowest layer feature, and its formula is: , in, Represents the low-level features after multi-scale enhancement, represents the original low-level features, Indicates Convolution kernels of different scales, represents the number of scales in the enhancement module, represents the activation function, is the scale index.
8. The method for constructing and recognizing a traffic sign image recognition model according to claim 7, characterized in that: In Step S3, the detection head adopts an anchor box adjustment strategy to locate the candidate regions, and the initial anchor box is set as: , Among them, represents the initial anchor box, and are the horizontal and vertical coordinates of the center of the anchor box respectively, and are the width and height of the anchor box respectively; The detection head dynamically adjusts the initial anchor boxes by predicting the offsets The adjusted anchor boxes are as follows: , Among them, and represent the translational offsets in the horizontal and vertical directions, and represent the logarithmic scale offsets of the width and height, is an exponential function; In EFPN, a small target detection layer is designed, and this sub-detection module takes the low-level features enhanced by multi-scale as the input, and its processing process is: , Among them, represents the small object detection result output by the sub-detection module, represents the operation of the sub-detection module, including local feature refinement and boundary regression. At the same time, the boundary regression optimization based on the focus mechanism is adopted inside the module, and its calculation formula is: , where, represents the optimized boundary regression result, represents the boundary regression function based on the focus mechanism, which uses low-level features and anchor box information to adjust the regression loss simultaneously.
9. The traffic sign image recognition model construction and recognition method according to claim 1, characterized in that: In Step S4, the loss function based on the focal mechanism adopted is the Focal-EIoU loss function.
10. The traffic sign image recognition model construction and recognition method according to claim 9, characterized in that: In step S4, when using transfer learning technology, the weights pre-trained on a large-scale image dataset are loaded into the hybrid network structure, where the parameters of the low-level convolutional units are kept frozen, and the parameters of the high-level fully connected units are fine-tuned to meet the requirements of the traffic sign recognition task; at the same time, a dynamic learning rate decay strategy is introduced to gradually reduce the learning rate; In step S4, the FocalEIoU loss function based on the focal mechanism is used to optimize the boundary regression. The FocalEIoU loss function adds a focal modulation term to the traditional IoU loss, assigns higher loss weights to low-IoU candidate boxes, and the calculation method is as follows: , Among them, represents the Focal EIoU loss, represents the balance coefficient for adjusting the focal term, is a positive real number, represents the complement value of the intersection over union (IoU) between the candidate box and the ground truth box, represents the ratio of their intersection to their union, represents the adjustment parameter in the focal mechanism, which is used to amplify the loss of low-IoU candidate boxes, , represents the center of the candidate box and the center of the ground truth box the squared Euclidean distance between them, and respectively represent the center coordinates of the candidate box and the ground truth box, represents the diagonal length of the minimum bounding box of the candidate box and the ground truth box, represents the width of the candidate box, represents the width of the ground truth box, represents the width of the minimum bounding box, which is used to normalize the width error, represents the height of the candidate box, represents the height of the ground truth box, represents the height of the minimum bounding box, which is used to normalize the height error.