Lightweight-based traffic sign detection method, storage medium and electronic equipment

By improving the YOLOv1n model and combining it with StarNet, C2PSA_TSSA modules, ASF feature pyramid network and LADH detection head, the robustness and lightweight issues of traffic sign detection in complex scenarios are solved, and efficient and accurate traffic sign detection is achieved.

CN120913175APending Publication Date: 2025-11-07UNIFORM ENTROPY TECH (WUXI) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511022719.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing traffic sign detection methods struggle to achieve robustness, real-time performance, and generalization capabilities in complex scenarios. Furthermore, it is difficult to balance the number of parameters, computational load, and detection accuracy in lightweight devices.

Method used

A lightweight traffic sign detection method is adopted, which performs layer-by-layer feature extraction through the backbone network, performs feature fusion and optimizes attention weight allocation in the neck, and performs feature separation prediction in the prediction module. The YOLOv1n model is used for improvement, and StarNet, C2PSA_TSSA module, ASF feature pyramid network and LADH detection head are introduced. The detection effect is optimized by combining the PIOU loss function.

Benefits of technology

It enables efficient traffic sign detection in lightweight equipment, improving detection accuracy and computational efficiency, adapting to complex scenarios, reducing computational load, and increasing detection speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913175A_ABST
    Figure CN120913175A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of road traffic sign detection, and particularly discloses a lightweight-based traffic sign detection method, a storage medium and electronic equipment, and the method comprises the steps: obtaining road traffic image information; preprocessing the road traffic image information; the preprocessed image information is input into a traffic sign detection model for feature extraction, feature fusion and classification detection in sequence, the traffic sign detection model at least comprises a backbone network, a neck and a prediction module, the backbone network carries out feature extraction on the preprocessed image information layer by layer according to a multi-kernel convolution strategy and element-by-element multiplication, and the feature extraction is carried out; the neck part carries out feature fusion on the image features of different levels and obtains a plurality of feature maps of different scales on the basis of optimizing attention weight distribution, and the prediction module carries out classification according to the feature maps of different scales to obtain a traffic sign detection result. The light-weight-based traffic sign detection method provided by the invention has the characteristic of light weight, and integration of the traffic sign detection methods can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of road traffic sign detection, and in particular to a traffic sign detection method based on light weight, a computer storage medium and an electronic device. BACKGROUND

[0002] Target detection, as a core task in the field of computer vision, is one of the objects of focus in academia and industry. Target detection technology has undergone tremendous evolution since its inception. From the advent of traditional machine learning models to the emergence of the two-stage representative model R-CNN, from the continuous improvement of Faster R-CNN to the sudden rise of the one-stage detection model represented by YOLO. Today, target detection models based on the Transformer architecture are in harmony with the YOLO series of models, and almost every year new improved models appear, providing endless research and creation motivation for scholars.

[0003] Intelligent transportation systems play a crucial role in optimizing autonomous driving technology, promoting car-cloud integration, and building smart cities. Traffic sign detection algorithms, as one of the key tasks of target detection algorithms, also play an important role in intelligent transportation systems, conveying rules, warnings, and directions to drivers and pedestrians. Traffic sign detection and recognition play an important role in avoiding traffic congestion and ensuring traffic safety. However, in the complex real world, current traffic sign detection generally faces challenges in scene complexity, such as environmental dynamics - differences in day and night lighting, weather interference with image quality; size diversity - image size in the data set varies from 80% to 5%; dynamic blur - shooting fast-moving objects causes target image blur with ghosting; interference density - target and background texture are similar, and there is a lot of overlap, occlusion, etc. This poses a serious challenge to the robustness, real-time performance, and generalization ability of traffic sign detection.

[0004] In addition, current traffic sign detection methods are usually integrated into small chips, such as mobile devices such as vehicles, drones, etc. This requires traffic sign detection to be implemented with light weight characteristics, thereby saving costs, avoiding excessive device space, and reducing device burden. Therefore, the chip cost should be as low as possible, the size should be as small as possible, and the computing power should be as high as possible. However, it is difficult to find a chip that balances these three at the moment. In addition, traditional light weight improvement methods are difficult to balance the balance between parameter size, computation size and detection accuracy. Reducing parameter size and computation size will inevitably cause a significant drop in model accuracy.

[0005] Therefore, how to provide a light weight traffic sign detection method has become a technical problem to be solved by those skilled in the art. SUMMARY

[0006] The application provides a traffic sign detection method based on light weight, a computer storage medium and an electronic device, and solves the problem that traffic sign detection does not have light weight in related technologies.

[0007] As a first aspect of the application, a traffic sign detection method based on light weight is provided, comprising:

[0008] Obtaining road traffic image information under a vehicle driving angle;

[0009] Preprocessing the road traffic image information to obtain preprocessed image information;

[0010] Inputting the preprocessed image information into a traffic sign detection model to sequentially perform feature extraction, feature fusion and classification detection, and obtaining a traffic sign detection result, wherein the traffic sign detection model at least includes a backbone network, a neck and a prediction module, the backbone network is used to perform layer-by-layer feature extraction on the preprocessed image information according to a multi-core convolution strategy and element-by-element multiplication to obtain image features of different levels, the neck is used to fuse the image features of different levels and obtain a plurality of feature maps of different scales based on optimized attention weight distribution according to a combination of top-down semantic information transmission and bottom-up detail transmission, and the prediction module is used to classify according to the feature maps of different scales to obtain the traffic sign detection result.

[0011] Further, the backbone network is used to perform layer-by-layer feature extraction on the preprocessed image information according to a multi-core convolution strategy and element-by-element multiplication to obtain image features of different levels, comprising:

[0012] Processing the preprocessed image information according to a depth separable convolution to extract local feature information in the preprocessed image information;

[0013] Mapping and projecting the local feature information to obtain mapped and projected local feature information;

[0014] Aligning the resolution of the mapped and projected local feature information to enable the aligned local feature information to be nonlinearly fused according to element-by-element multiplication to obtain image features of different levels.

[0015] Further, the expression of the element-by-element multiplication is:

[0016]

[0017] wherein i and j represent index channels, represents the i-th element of the weight vector w1, represents the j-th element of the weight vector w2, and xi and x j respectively represent the i-th element and the j-th element of the input vector, a represents the coefficient of each item, and

[0018] Further, the local feature information is mapped and projected to obtain mapped and projected local feature information, including:

[0019] According to the linear projection manner, the local feature information is mapped to the same dimension to obtain the mapped feature in the same dimension, and the mapping expression of the same dimension is:

[0020]

[0021] where b represents the batch size, n represents the sequence length, d represents the feature dimension, and x' represents the linearly projected feature.

[0022] According to the multi-head splitting manner, the mapped feature in the current dimension is projected to a target dimension subspace to obtain the mapped and projected head feature, where the dimension of the target dimension subspace is smaller than the current dimension, and the projection expression of the target dimension subspace is:

[0023]

[0024] where w represents the mapped and projected head feature, p represents the dimension of the target dimension subspace, and k represents the number of heads.

[0025] Further, the neck is used to fuse image features of different levels and obtain a plurality of feature maps of different scales based on the combination of top-down semantic information transmission and bottom-up detail transmission in the optimization of attention weight distribution, including:

[0026] The image features of different levels are convolved, and the convolved image features are smoothed and resolution-adjusted to obtain an attention scale sequence fusion feature map.

[0027] The image features of different levels are respectively split and processed, and the split and processed image features are spliced to obtain a triple encoding feature map.

[0028] The attention scale sequence fusion feature map and the triple encoding feature map are subjected to cross-level feature enhancement processing to obtain a plurality of feature maps of different scales.

[0029] Further, the image features of different levels are convolved, and the convolved image features are smoothed and resolution-adjusted to obtain an attention scale sequence fusion feature map, including:

[0030] Convolve image features of different levels according to Gaussian kernels with gradually increasing standard deviations;

[0031] Smooth the convolved image features to obtain feature maps with consistent resolutions;

[0032] Extract cross-scale semantic information from the feature maps with consistent resolutions according to 3D convolution, normalization, and activation functions to obtain an attention scale sequence fusion feature map.

[0033] Further, split the image features of different levels respectively, and splice the split image features to obtain triple-encoding feature maps, including:

[0034] Perform spatial dimension reduction processing on the first scale feature in the image features of different levels through convolution and down-sampling to obtain a first scale feature processing result;

[0035] Perform spatial dimension reduction processing on the third scale feature in the image features of different levels through convolution and up-sampling to obtain a third scale feature processing result;

[0036] Convolve and splice the first scale feature processing result, the second scale feature processing result, and the third scale feature processing result in the image features of different levels to obtain the triple-encoding feature maps;

[0037] Wherein, the first scale is greater than the second scale, and the second scale is greater than the third scale.

[0038] Further, perform cross-level feature enhancement processing on the attention scale sequence fusion feature map and the triple-encoding feature map to obtain a plurality of feature maps of different scales, including:

[0039] Process the attention scale fusion feature map based on a channel attention network to obtain a channel-related feature map;

[0040] Process the triple-encoding feature map based on a position attention network to obtain a position-related feature map;

[0041] Fuse the channel-related feature map and the position-related feature map to obtain a plurality of feature maps of different scales.

[0042] As another aspect of the present application, a computer storage medium is provided for storing a computer program, which is executed by a processor to implement the aforementioned traffic sign detection method based on lightweight.

[0043] As another aspect provided by the present application, an electronic device is provided, which comprises a memory and a processor, the processor is communicatively connected with the memory, the memory is configured to store a computer program, and the processor is configured to load and execute the computer program to implement the aforementioned traffic sign detection method based on light weight.

[0044] The traffic sign detection method based on light weight provided by the present application acquires road traffic image information under a driving angle of a vehicle, and pre-processes the road traffic image information to obtain pre-processed image information, inputs the pre-processed image information into a traffic sign detection model to perform feature extraction, feature fusion and classification detection, and finally obtains a traffic sign detection result. In the embodiment of the present application, the traffic sign detection model used in the present application effectively reduces the calculation amount by performing layer-by-layer feature extraction on the pre-processed image information based on the element-by-element multiplication in the specific structure of the main network, and effectively improves the calculation efficiency by performing feature fusion on the image features of different levels in the neck and optimizing the attention weight distribution in the feature fusion, and finally obtains the traffic sign detection result by performing feature separation prediction on the result after feature fusion by the prediction module. The traffic sign detection model performs light weight calculation in the process of feature extraction, feature fusion and final feature classification prediction, so as to realize light weight detection of traffic signs. BRIEF DESCRIPTION OF DRAWINGS

[0045] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, and together with the specific embodiments described below, serve to explain the present application, but do not constitute a limitation on the present application.

[0046] Figure 1 The flowchart of the traffic sign detection method based on light weight provided by the present application.

[0047] Figure 2 The architecture diagram of the traffic sign detection model provided by the present application.

[0048] Figure 3 The flowchart of the feature extraction performed by the main network provided by the present application.

[0049] Figure 4 The architecture diagram of the improved main network provided by the present application.

[0050] Figure 5 The flowchart of the feature fusion performed by the neck provided by the present application.

[0051] Figure 6 The architecture diagram of the SSFF module of the neck provided by the present application.

[0052] Figure 7 The architecture diagram of the TEE module of the neck provided by the present application.

[0053] Figure 8 The architecture diagram of the CPAM module of the neck provided in the present application.

[0054] Figure 9 The architecture diagram of the LADH module of the prediction module provided in the present application.

[0055] Figure 10 The structure block diagram of the electronic device provided in the present application. DETAILED DESCRIPTION

[0056] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0057] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0058] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device containing a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0059] In the present embodiment, a traffic sign detection method based on light weight is provided, Figure 1 is a flow chart of the traffic sign detection method based on light weight provided according to the embodiments of the present application, as Figure 1 shown, comprising:

[0060] S100, acquiring road traffic image information under a driving angle of a vehicle;

[0061] In the embodiments of the present application, the road information video recorded by the front view angle device of the vehicle's driving recorder can be downloaded and then framed according to a fixed frame rate, and the image content containing the traffic sign is screened.

[0062] S200, preprocessing the road traffic image information to obtain preprocessed image information;

[0063] Specifically, by scaling the original road traffic image to a fixed size, applying Gaussian filter denoising, pixel value normalization, and preprocessing the image information through data enhancement techniques, diversified samples are generated, i.e., preprocessed image information is obtained.

[0064] S300, inputting the preprocessed image information into a traffic sign detection model to sequentially perform feature extraction, feature fusion, and classification detection, and obtaining a traffic sign detection result, wherein the traffic sign detection model at least includes a backbone network, a neck, and a prediction module, the backbone network is used to perform layer-by-layer feature extraction on the preprocessed image information according to a multi-core convolution strategy and element-by-element multiplication to obtain image features of different levels, the neck is used to fuse the image features of different levels and obtain a plurality of feature maps of different scales based on optimized attention weight distribution according to a combination of top-down semantic information transmission and bottom-up detail transmission, and the prediction module is used to classify according to the feature maps of different scales to obtain the traffic sign detection result.

[0065] It should be understood that the preprocessed image information is input into the traffic sign detection model, and the features are extracted layer by layer through convolution, pooling, etc. by the backbone network, then the extracted features of different levels are fused by the neck, and a combination of top-down semantic information transmission and bottom-up fine-grained detail transmission is adopted, finally three feature maps of large, medium and small scales are obtained; finally, through the prediction module, the feature maps of different scales are obtained, the classification task and the regression task are separated through the double-head decoupling structure, and the bounding box regression loss based on the non-monotonic focusing mechanism is used in the regression task, so that the loss function is more focused on the anchor box of medium quality, and the convergence speed and accuracy of the bounding box regression are improved.

[0066] Therefore, the traffic sign detection method based on light weight provided by the application obtains road traffic image information under the driving angle of the vehicle, and pre-processes the road traffic image information to obtain pre-processed image information. The pre-processed image information is input into a traffic sign detection model for feature extraction, feature fusion and classification detection, and finally the traffic sign detection result is obtained. The traffic sign detection model used in the embodiment of the application effectively reduces the calculation amount based on the element-by-element multiplication method due to the layer-by-layer feature extraction of the pre-processed image information by the backbone network in the specific structure of the traffic sign detection model, and effectively improves the calculation efficiency by the feature fusion of the image features of different levels by the neck and the optimization of the attention weight distribution in the feature fusion. The feature separation prediction of the prediction module on the result after the feature fusion finally obtains the traffic sign detection result. The traffic sign detection model is light weight in the process of feature extraction, feature fusion and final feature classification prediction, so as to realize the light weight detection of the traffic sign.

[0067] In the embodiment of the application, the specific basic model of the traffic sign detection model is YOLOv11n, such as Figure 2As shown, the core framework is divided into the following three modules: 1. backbone network; 2. neck; 3. prediction module. Among them, the backbone network introduces the C3K2 module, adopts the multi-core convolution strategy (3x3 and 1x1 convolution in parallel) on the basis of the traditional CSP module, enhances the multi-scale feature extraction capability, and at the same time reduces the calculation amount through the depth separable convolution technology. And, the C2PSA module is added behind the original SPPF layer, which improves the calculation efficiency through segmented fusion and pyramid slice attention module. The neck adopts the top-down feature pyramid network (FPN) and path aggregation network (PANet), which aggregates and fuses shallow and deep semantic features through the upper and lower paths. The prediction module uses a decoupled detection head and uses an Anchor-Free design to directly predict the target center point offset and width and height, avoiding the limitations of preset anchor boxes. The traffic sign detection model of the embodiment of the application not only includes a basic model but also includes a lightweight feature added on the basis of the basic model. Specifically: the original backbone network part is improved using StarNet, which uses element-wise multiplication to reduce the amount of calculation while retaining the detection accuracy. Second, the improved C2PSA_TSSA module is used, which optimizes the attention weight distribution through Token statistical information and improves the calculation efficiency with linear time complexity overhead. Next, the ASF feature pyramid network is introduced in the neck part to realize adaptive spatial fusion and dynamically add multi-scale features to better retain high-level features and improve the performance of the lightweight model in complex scenarios. In addition, the prediction module of the embodiment of the application also uses the LADH lightweight asymmetric detection head to dynamically adjust the feature representation of different categories and reduce redundant calculations while improving the sensitivity to small size targets. Finally, the PIoU loss function based on the non-monotonic focusing mechanism is introduced to replace the original CIoU loss function, which can dynamically adjust the gradient gain and combine the target size adaptive penalty factor to optimize the positioning accuracy of small targets and extreme sizes.

[0068] Specifically, the backbone network is used to perform layer-by-layer feature extraction on the preprocessed image information according to a multi-core convolution strategy and element-wise multiplication to obtain image features of different levels, as shown in Figure 3 As shown, it includes:

[0069] S311, processing the preprocessed image information according to the depth separable convolution to extract local feature information in the preprocessed image information;

[0070] S312, mapping and projecting the local feature information to obtain the local feature information after mapping and projecting;

[0071] In the embodiment of the application, the feature map can be mapped to the implicit space through the fully connected layer to generate a high-dimensional feature representation, and the ReLU6 is activated to reduce the parameter amount.

[0072] It should be noted that the YOLOv11 model newly adds a C2PSA module behind the SPPF module in the backbone network, the module combines the CSP structure and the pyramid slice attention mechanism, and the target detection performance is improved through multi-scale feature fusion and attention optimization. In view of the problem that the calculation complexity of C2PSA is relatively high when processing long sequences or large-scale data, the original C2PSA module is optimized by a Token statistics based attention mechanism module.

[0073] Specifically, the Token Statistics Transformer is a new Transformer architecture based on linear time complexity attention mechanism, and the core is the Token Statistics Self-Attention (TSSA) operator, which reduces the complexity by calculating the statistical features instead of pairwise similarity, and significantly improves the calculation efficiency and memory usage efficiency, as shown in Figure 4 The linear attention mechanism (TSSA) is introduced into the C2PSA module, which can further optimize the feature extraction capability and calculation efficiency of the C2PSA.

[0074] Further specifically, the local feature information is mapped and projected to obtain mapped and projected local feature information, including:

[0075] 1) The local feature information is mapped to the same dimension according to the linear projection mode, and the mapped feature in the same dimension is obtained, and the mapping expression of the same dimension is:

[0076]

[0077] Where b represents the batch size, n represents the sequence length, d represents the feature dimension, and x' represents the feature after linear projection.

[0078] 2) The mapped feature in the current dimension is projected to a target dimension subspace according to the multi-head splitting mode, wherein the dimension of the target dimension subspace is less than the current dimension, and the head feature after mapping and projection is obtained, and the projection expression of the target dimension subspace is:

[0079]

[0080] Where w represents the head feature after mapping and projection, p represents the dimension of the target dimension subspace, and k represents the number of heads.

[0081] In the embodiment of the application, the logic of TSSA is first to project the feature matrix Map to the same dimension, get:

[0082]

[0083] Where b represents the batch size, n represents the sequence length, d represents the feature dimension, and x' represents the linearly transformed feature.

[0084] And then split into k heads, each head has a dimension of That is, map the d-dimensional feature to the p-dimensional subspace (p << d), so as to realize feature projection: Where w represents the head feature of x'.

[0085] After splitting, each head only calculates the p-dimensional feature, which significantly reduces the complexity of attention calculation. Then, the projected feature After normalization and squaring, sum along the feature dimension p to get the second moment σ 2 The expression of the approximate variance is as follows:

[0086] In order to control the range of the second moment, it is necessary to scale the second moment by a learnable parameter:

[0087] σ 2 ← α· σ 2 + β,

[0088] Where α and β are learnable parameters. Then, generate weight π through Softmax(), as a gating signal to suppress unimportant features, so as to realize statistical gating:

[0089] π = Softmax(σ 2 ).

[0090] Finally, weight the feature x' to filter out important features x'', so as to make the attention mechanism pay more attention to important Token, thereby reducing the computational complexity: x'' = π·x'.

[0091] In addition, the design of TSSA is based on MCR2 (Maximum Coding Rate Reduction) theory, aiming to maximize the coding rate reduction and optimize data representation. The MCR2 goal is specifically manifested as: ΔR = R(Z) - R(Z|Y), where R(Z) is the coding rate of all data, and R(Z) is the coding rate under the category condition. Y represents a random "component allocation" matrix.

[0092] Therefore, maximizing ΔR makes the same features more compact and different classes more separated. Using TSSA operator in the improved C2PSSA_TSS module can significantly improve the computational efficiency.

[0093] S313, performing resolution alignment on the mapped and projected local feature information, so that the aligned local feature information can be nonlinearly fused according to element-wise multiplication to obtain image features of different levels.

[0094] In the embodiments of the present application, the feature maps of different scales are resolution-aligned, and the aligned feature maps are nonlinearly fused using element-wise multiplication after ensuring the size consistency of the feature maps.

[0095] Specifically, the expression of the element-wise multiplication is:

[0096]

[0097] where i and j represent index channels, represents the i-th element of the weight vector w1, represents the j-th element of the weight vector w2, x i and x j respectively represent the i-th element and the j-th element of the input vector, and a represents the coefficient of each item.

[0098] It should be noted that the above describes the rewriting and expansion process of the element-wise multiplication in detail. Through the above formula, it can be inferred that the element-wise multiplication can be expanded to different items. Thus, it can be illustrated that although the element-wise multiplication is calculated in a d-dimensional space, through these nonlinear combination items, the feature is actually implicitly mapped to a higher dimension, which is:

[0099]

[0100] It should be understood that the implicit feature dimension expansion can not bring additional computational overhead. The implicit feature dimension expansion of single-layer element-wise multiplication is described in detail below. Assuming that the initial layer width of the network is d, after one element-wise multiplication, the output feature is changed to This expression implicitly maps the input feature to a new feature space where represents the number of implicit dimensions. Then, extending to stack multiple element-wise multiplications, assuming that each layer output continues to implicitly expand the feature through element-wise multiplication based on the output of the previous layer, the first layer output is obtained:

[0101]

[0102] The second layer output is:

[0103]

[0104] The third layer output is:

[0105]

[0106] Thus, recursively to the l-th layer output:

[0107]

[0108] As can be seen from to Stacking multiple element-wise multiplication operations can significantly amplify the implicit dimension in exponential form, so adding element-wise operations in the previous StarBlock can significantly improve the computational efficiency of the model while maintaining excellent performance of the model. This is the reason why the original backbone network is replaced by the backbone network of StarNet in the embodiments of the present application.

[0109] In the embodiments of the present application, the pyramid network of the YOLOv11 model usually adopts a fixed weight feature fusion method (FPN+PAN), wherein the FPN is responsible for passing down semantic information from top to bottom, and the PAN is responsible for passing up fine-grained details from bottom to top. In order to improve the up-sampling, down-sampling and splicing operations to fuse high-level semantic information in feature maps of different scales, the embodiments of the present application integrate the attention scale sequence fusion (SSFF module), triple feature encoder (TFE module) and cross-level feature enhancement (CPAM module) in the ASF feature pyramid network into the neck part of YOLOv11n, and automatically adjust the fusion weights of each level of features according to the input image content through a dynamic learning mechanism, so as to replace the original network structure and improve the adaptability of the model to complex scenes.

[0110] Specifically, the neck part is used for feature fusion of image features of different levels and obtains a plurality of feature maps of different scales in a manner of combining top-down semantic information transmission and bottom-up detail transmission on the basis of optimizing attention weight distribution, as shown in Figure 5 , including:

[0111] S321, convolve the image features of different levels, and smooth and adjust the resolution of the convolved image features to obtain an attention scale sequence fusion feature map;

[0112] It should be noted that, as shown in Figure 6 , the SSFF module is designed around the P3 level, and P3 is the feature map with the highest resolution in the pyramid. The SSFF module can preserve fine-grained information of small targets, motion blur and occlusion boundaries.

[0113] Specifically, the image features of different levels are convolved, and the convolved image features are smoothed and adjusted in resolution to obtain an attention scale sequence fusion feature map, including:

[0114] 1) Convolve image features of different levels with Gaussian kernels of gradually increasing standard deviation;

[0115] 2) Smooth the convolved image features to obtain feature maps with consistent resolution;

[0116] 3) Extract cross-scale semantic information from the feature maps with consistent resolution according to 3D convolution, normalization, and activation functions to obtain an attention scale sequence fusion feature map.

[0117] The SSFF module first convolves the feature maps from P3, P4, and P5 using a series of Gaussian kernels with gradually increasing standard deviation. This allows each feature map to be smoothed, which can be expressed as follows:

[0118]

[0119] where F σ (i,j) represents a two-dimensional feature map, o represents the standard deviation of the Gaussian kernel, and G σ (x,y) represents a Gaussian filter.

[0120] As the standard deviation increases, the feature map is gradually smoothed, which allows information of different scales to be captured. To ensure that the smoothed feature map has consistent resolution, the resolution of the feature map is adjusted to that of the P3 layer using nearest-neighbor interpolation. Figure 1 Finally, scale-related information is extracted through 3D convolution, normalization, and activation functions to extract cross-scale semantic information.

[0121] S322, split the image features of different levels respectively, and splice the split image features to obtain triple-encoding feature maps;

[0122] In an embodiment of the present application, as shown in FIG. 3, the TFE module can split the features of large, medium, and small scales. The added large-size feature map can help the small-size feature map to refer to the changes in its shape or appearance, which is more friendly to the occlusion scene. The processing procedure of the TFE module for the feature map will be described in detail below. Figure 7

[0123] Specifically, the image features of different levels are split respectively, and the split image features are spliced to obtain triple-encoding feature maps, including:

[0124] 1) The first scale feature in the image features of different levels is processed by convolution and down-sampling to reduce the spatial dimension to obtain a first scale feature processing result;

[0125] ​2) For the third scale feature in the image feature of different levels, a third scale feature processing result is obtained through convolution and up-sampling processing;

[0126] 3) The first scale feature processing result, the second feature processing result and the third feature processing result in the image feature of different levels are convolved and spliced to obtain a triple encoding feature map;

[0127] Wherein, the first scale is greater than the second scale, and the second scale is greater than the third scale.

[0128] It should be understood that the TFE module needs to process feature maps from three sizes. First, the large size feature map processing process is as follows: after the large scale feature map is processed by the convolution module, the channel number is adjusted to 1C. A mixed processing method of maximum pooling and average pooling is used for down-sampling to reduce the spatial dimension of the feature and realize the translational invariance. Secondly, the small size feature map processing process is as follows: after the small size feature map is adjusted to 1C using the convolution module, the nearest neighbor interpolation method is used for up-sampling to retain local features, retain small object feature information and avoid the interference of complex background. Finally, the processed large, medium and small feature maps are convolved and spliced to obtain the output feature map.

[0129] S323, cross-level feature enhancement processing is performed on the attention scale sequence fusion feature map and the triple encoding feature map to obtain a plurality of feature maps of different scales.

[0130] Specifically, as shown in Figure 8 The CPAM module is used to integrate detailed and multi-scale feature information from the SSFF module and the TFE module. The module mainly consists of a channel attention network and a position attention network. The channel attention network receives the output from the TFE module, while the position attention network receives the output from the SSFF module and the superposition of the output from the channel attention network. The channel attention network and the position attention network in the CPAM module will be introduced below.

[0131] In the embodiment of the application, the cross-level feature enhancement processing is performed on the attention scale sequence fusion feature map and the triple encoding feature map to obtain a plurality of feature maps of different scales, including:

[0132] 1) The channel attention network is used to process the attention scale fusion feature map to obtain a channel related feature map;

[0133] 2) The position attention network is used to process the triple encoding feature map to obtain a position related feature map;

[0134] 3) fusing the channel-related feature map and the position-related feature map to obtain a plurality of feature maps of different scales.

[0135] Specifically, the channel attention network first processes each channel independently through global average pooling, and generates channel weights using two fully connected layers and a nonlinear Sigmoid function. The purpose of the two fully connected layers is to capture the nonlinear interaction between channels. In order to avoid the two fully connected layers reducing the operation efficiency, the embodiment of the application captures the interaction across channels by using each channel and its k nearest neighbors. This way is realized by 1D convolution with a convolution kernel size of k, where k represents how many neighbors participate in the attention prediction of a channel. K is adjustable and can be adjusted according to different network structures and the number of convolution modules. In this way, the expression of the channel dimension C is as follows:

[0136] C = ψ(k) = 2 (γ×k-b) ,

[0137] Where γ and b control the scaling parameters between the convolution kernel size k and the channel dimension C, respectively.

[0138] Then, k can be mapped as: Where || represents the nearest odd number. odd The value of γ is set to 2 and b is set to 1. The formula shows that the higher the channel value, the longer the exchange distance, so that multiple channel features can be mined at a deeper level.

[0139] In addition, the position attention network is different from the channel attention network. It first divides the feature map into two parts according to the width and height, and then performs average pooling on the width and height axes respectively, and finally merges the output. The calculation process of the average pooling on the width and height axes is as follows:

[0140]

[0141] Where W and H represent the width and height of the input feature map respectively, and E(i,j) represents the value of the position coordinate (i,j) where the input feature value is located.

[0142] After completing the average pooling operation on the width and height axes, the features are spliced and convolved, and then the features are split to generate a pair of position-related feature maps: s w = Split(a w ), s h = Split(a h ), and finally the output of the CPAM module is obtained: F CPAM = E x sw x sh, where E represents the weight matrix of the channel and position attention.

[0143] In the embodiments of the present application, as shown in Figure 9 The LADH detection head in the prediction module is used for dynamically adjusting the feature representation of different categories of targets, and the classification and regression tasks are separated through a double-head decoupling structure, which reduces redundant parameters and significantly reduces the computational complexity. The LADH detection head has two convolution modes, 1x1 convolution and 3x3 depth separable convolution. Among them, the classification task uses two layers of 1x1 convolution, and the regression task uses multiple layers of depth separable convolution for feature extraction of the bounding box, and then uses 1x1 convolution to compress the channel. In addition, different detection heads use different compression ratios. The small target head uses a higher compression ratio to reduce redundant calculation, and the large target head uses a lower compression ratio to retain more semantic information.

[0144] In the embodiments of the present application, in view of the problem that the convergence speed may be slow caused by the default use of CIOU loss in YOLOv11 model, Powerful-IoU (hereinafter referred to as PIoU) is introduced in the embodiments of the present application. As an improved loss function, it is used for the bounding box regression task in target detection. It significantly improves the convergence speed and accuracy of the bounding box regression by introducing a target size adaptive penalty factor and a gradient adjustment function based on anchor box quality. In addition, a non-monotonic attention mechanism is also introduced based on the original loss function, which further enhances the focusing ability of the model on the anchor box of medium quality. In the embodiments of the present application, PIoU is used to replace CIoU to obtain more efficient detection results.

[0145] Specifically, an adaptive penalty factor is defined in the PIoU loss function, which can dynamically adjust the regression multiplication strength according to the actual size of the target, ensuring that targets of different sizes can converge efficiently. The formula of the penalty factor P is as follows:

[0146]

[0147] Where dw1, dw2, dh1 and dh2 represent the absolute values of the distances between the corresponding edges of the prediction box and the target box, w gt and h gt represent the width and height of the target box, respectively.

[0148] It is proved by many experimental results that using P as the penalty factor will not cause the anchor box to expand. Because the denominator of P only depends on the size of the target box, and is independent of the size of the smallest enclosing box of the anchor box and the target box. In this way, the target size can be dynamically adapted to ensure that the positioning error of small targets will be fully punished, while the shape deviation of large targets will not be excessively enlarged. As can be seen, PIoU can well cope with the situation of different sizes in the data set.

[0149] It should be understood that the ideal loss function should have the following characteristics: (1) when the regression error is close to 0, the gradient amplitude should be close to 0; (2) for smaller regression errors, the gradient increment should increase rapidly, and vice versa, for larger regression errors, the gradient increment should gradually decrease; (3) the anchor frame with poor quality or high quality should have a smaller gradient; (3) the maximum gradient should be at the anchor frame with medium quality. With these characteristics, the convergence can be accelerated and the redundant frame expansion can be reduced. In order to meet the above requirements, the gradient adjustment function f(x) and the loss function L PIoU The design is as follows:

[0150]

[0151] PIoU=IoU-f(P),-1≤PIoU≤1,

[0152] L PIoU =1-PIoU=L IoU +f(P),0≤L PIoU ≤2。

[0153] In order to further optimize the loss function, a non-monotonic attention layer can be added to the original loss function to dynamically adjust the gradient contribution of anchor frames of different quality in the loss function. The loss function is more focused on anchor frames of medium quality, which can accelerate convergence and suppress noise. Its expression is as follows:

[0154] q=e -P ,q∈(0,1],

[0155]

[0156] Where P represents the penalty factor, q reflects the quality of the anchor frame. When P=0, q=1, indicating that the anchor frame quality is the highest. When P becomes larger and larger, q gradually tends to 0, indicating that the anchor frame quality decreases, and u(x) is an attention function that dynamically adjusts the gradient contribution of anchor frames of different quality in the loss function.

[0157] Then the new version of the loss function L PIoU-v2 The expression is as follows:

[0158]

[0159] Where lambda represents a hyperparameter that controls the behavior of the attention function, and the focus range of the attention function is controlled by the hyperparameter lambda, which adapts to different data sets and task requirements.

[0160] The beneficial effects of the traffic sign detection method based on lightweight of the embodiments of the application are verified as follows.

[0161] Firstly, three traffic sign detection datasets are used in the embodiments of the present application. Respectively: 1. GTSDB, which is annotated into three categories (prohibition signs, indication signs, danger warning signs), which conforms to the typical categories of European traffic signs. 2. CCTSDB2021, which is annotated similarly to GTSDB, but contains more images and covers various lighting and weather conditions, focusing on the diversity of domestic local traffic signs. 3. TT100K, which contains a total of 221 different traffic signs, of which 128 categories have been labeled. The categories are as shown in Figures 3-8 Compared with the above two datasets, more categories can better verify the detection performance of the model in multiple scenes and multiple targets. However, some categories have too small a number, in order to avoid affecting the accuracy of the model, categories with less than 100 are excluded, and finally a new dataset composed of 5621 images and 45 categories is obtained. These traffic sign images are affected by complex scenes, such as dark scenes, being blocked by tree branches and leaves, traffic signs being too small in size in the image, complex backgrounds, etc. From P (precision), R (recall), Parameters (parameter amount), GFLOPS (computational amount), mAP@50 (representing the average precision when IoU is 0.5), mAP@50:95 (average precision when IoU is from 0.5 to 0.95 (step size is 0.05)).

[0162] Secondly, the effect of the embodiments of the present application is evaluated, and compared with other top models on the above three datasets (one-stage real-time detection YOLO series models: YOLOv5n, YOLOv8n, YOLOv9-tiny, YOLOv10n, YOLOv11n, hyper-YOLO), as shown in Tables 1 to 3 respectively. The results show that the model used by the lightweight-based traffic sign detection method of the embodiments of the present application decreases by 1MParams, about 39%, and the computational amount decreases by 2.2GFlops, about 34%, compared with the YOLOv11n benchmark model. On the GTSDB dataset, mAP@50 is improved by 0.7%, and mAP@50:95 is improved by 0.5%; on the CCTSDB2021 dataset, mAP@50 decreases by 0.8%, and mAP@50:95 decreases by 0.4%; on the TT100K dataset, mAP@50 is improved by 0.3%, and mAP@50:95 is improved by 0.7%.

[0163] Table 1 Comparison experiment on GTSDB dataset

[0164]

[0165] Table 2 Comparison experiment on CCTSDB2021 dataset

[0166]

[0167] Table 3 Comparative experiments on the TT100K dataset

[0168]

[0169] In summary, the traffic sign detection method based on light weight provided by the present application uses a traffic sign detection model based on light weight YOLOv11n for traffic sign detection, which is improved, and the model parameter quantity is reduced by 39% and the calculation quantity is reduced by 34% based on YOLOv11n. Good results are achieved in traffic sign detection in various complex scenes. Therefore, the traffic sign detection method based on light weight of the present application uses a light weight module to greatly reduce the model size and calculation overhead; while the model is light weight, the accuracy of traffic sign detection in complex scenes is also considered.

[0170] As another embodiment of the present application, a computer storage medium is provided for storing a computer program, which is executed by a processor to implement the traffic sign detection method based on light weight described above.

[0171] In an embodiment of the present application, a non-transitory computer readable storage medium is provided, which stores computer executable instructions, and the computer executable instructions can execute the traffic sign detection method based on light weight in any method embodiment described above. The storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD), etc. The storage medium can also include a combination of the above types of memories.

[0172] As another embodiment of the present application, an electronic device is provided, which includes a memory and a processor, the processor is in communication connection with the memory, the memory is used to store a computer program, and the processor is used to load and execute the computer program to implement the traffic sign detection method based on light weight described above.

[0173] As Figure 10As shown, the electronic device 10 can include at least one processor 11, such as a CPU (Central Processing Unit), at least one communication interface 13, a memory 14, and at least one communication bus 12. The communication bus 12 is used to realize the connection and communication between the components. The communication interface 13 can include a display, a keyboard, and can also include a standard wired interface and a wireless interface. The memory 14 can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. The memory 14 can also be at least one storage device located away from the aforementioned processor 11. The memory 14 stores an application program, and the processor 11 calls the program code stored in the memory 14 to execute any of the above method steps.

[0174] The communication bus 12 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 12 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0175] The memory 14 can include a volatile memory such as a RAM (Random-Access Memory), and can also include a non-volatile memory such as a flash memory, a hard disk (HDD) or a solid-state disk (SSD). The memory 14 can also include a combination of the above types of memories.

[0176] The processor 11 can be a CPU (Central Processing Unit), a network processor (NP), or a combination of a CPU and an NP.

[0177] The processor 11 can further include a hardware chip. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0178] Optionally, the memory 14 is further configured to store program instructions. The processor 11 can invoke the program instructions to implement the traffic sign detection method based on light weight shown in the embodiments of the present application. Figure 1 Optionally, the memory 14 is further configured to store program instructions. The processor 11 can invoke the program instructions to implement the traffic sign detection method based on light weight shown in the embodiments of the present application.

[0179] It can be understood that the above embodiments are only exemplary embodiments for illustrating the principles of the present application, and the present application is not limited thereto. Various modifications and improvements can be made by those of ordinary skill in the art without departing from the spirit and essence of the present application, and these modifications and improvements are also considered to be within the protection scope of the present application.

Claims

1. A traffic sign detection method based on lightening, characterized in that, The method comprises the following steps: acquiring road traffic image information under a driving angle of a vehicle; preprocessing the road traffic image information to obtain preprocessed image information; inputting the preprocessed image information into a traffic sign detection model to sequentially perform feature extraction, feature fusion and classification detection, and obtaining a traffic sign detection result, wherein the traffic sign detection model comprises at least a backbone network, a neck and a prediction module, the backbone network is used to perform layer-by-layer feature extraction on the preprocessed image information according to a multi-core convolution strategy and element-by-element multiplication to obtain image features of different levels, the neck is used to perform feature fusion on the image features of different levels and obtain feature maps of multiple different scales based on optimized attention weight distribution according to a combination of top-down semantic information transmission and bottom-up detail transmission, and the prediction module is used to perform classification according to the feature maps of different scales to obtain the traffic sign detection result.

2. The lightweight-based traffic sign detection method according to claim 1, characterized in that, The backbone network is used to perform layer-by-layer feature extraction on the preprocessed image information according to a multi-core convolution strategy and element-by-element multiplication to obtain image features of different levels, comprising: processing the preprocessed image information according to a depth separable convolution to extract local feature information in the preprocessed image information; performing mapping projection processing on the local feature information to obtain mapped and projected local feature information; performing resolution alignment on the mapped and projected local feature information to enable the aligned local feature information to be nonlinearly fused according to element-by-element multiplication to obtain image features of different levels.

3. The lightweight-based traffic sign detection method according to claim 2, characterized in that, The expression of the element-by-element multiplication is: where i and j represent index channels, represents the i-th element of the weight vector w1, represents the j-th element of the weight vector w2, x i and x j represent the i-th and j-th elements of the input vector, respectively, and a represents a coefficient for each item, and 4.The method of claim 2, wherein, The mapping projection processing on the local feature information to obtain mapped and projected local feature information comprises: mapping the local feature information to the same dimension according to a linear projection to obtain mapped features in the same dimension, wherein the mapping expression of the same dimension is: wherein b represents batch size, n represents sequence length, d represents feature dimension, and x' represents the feature after linear projection; projecting the mapped features in the current dimension to a target dimension subspace according to a multi-head splitting manner, wherein the dimension of the target dimension subspace is smaller than the current dimension, to obtain head features after mapping projection, wherein the projection expression of the target dimension subspace is: wherein w represents the mapped projected head feature, p represents the dimension of the target dimensional subspace, and k represents the number of heads. 5.The lightweight-based traffic sign detection method according to claim 1, wherein, The neck is used to perform feature fusion on the image features of different levels and obtain feature maps of multiple different scales based on optimized attention weight distribution according to a combination of top-down semantic information transmission and bottom-up detail transmission, comprising: performing convolution on the image features of different levels, and performing smoothing processing and resolution adjustment on the convolved image features to obtain attention scale sequence fusion feature maps; respectively performing splitting processing on the image features of different levels, and performing splicing on the split image features to obtain triple-encoding feature maps; performing cross-level feature enhancement processing on the attention scale sequence fusion feature maps and the triple-encoding feature maps to obtain the feature maps of multiple different scales. 6.The lightweight-based traffic sign detection method according to claim 5, characterized in that, The image features of different levels are convolved, and the convolved image features are smoothed and resolution-adjusted to obtain an attention scale sequence fusion feature map, including: The image features of different levels are convolved according to a Gaussian kernel with gradually increasing standard deviation; The convolved image features are smoothed to obtain feature maps with consistent resolution; The feature maps with consistent resolution are subjected to cross-scale semantic information extraction according to 3D convolution, normalization and activation function to obtain the attention scale sequence fusion feature map. 7.The lightweight-based traffic sign detection method according to claim 5, characterized in that, The image features of different levels are respectively subjected to splitting processing, and the split image features are spliced to obtain triple-encoding feature maps, including: The first scale feature in the image features of different levels is subjected to spatial dimension reduction processing through convolution and down-sampling to obtain a first scale feature processing result; The third scale feature in the image features of different levels is subjected to spatial dimension reduction processing through convolution and up-sampling to obtain a third scale feature processing result; The first scale feature processing result, the second scale feature processing result and the third scale feature processing result in the image features of different levels are convolved and spliced to obtain triple-encoding feature maps; Wherein, the first scale is greater than the second scale, and the second scale is greater than the third scale. 8.The lightweight-based traffic sign detection method according to claim 5, wherein, The attention scale sequence fusion feature map and the triple-encoding feature map are subjected to cross-level feature enhancement processing to obtain a plurality of feature maps of different scales, including: The attention scale fusion feature map is processed based on a channel attention network to obtain a channel-related feature map; The triple-encoding feature map is processed based on a position attention network to obtain a position-related feature map; The channel-related feature map and the position-related feature map are fused to obtain a plurality of feature maps of different scales.

9. A computer storage medium, characterized in that A computer program is stored, and the computer program is executed by a processor to implement the lightweight-based traffic sign detection method of any one of claims 1-8.

10. An electronic device, comprising: A memory and a processor are included, the processor is in communication connection with the memory, the memory is used to store a computer program, and the processor is used to load and execute the computer program to implement the lightweight-based traffic sign detection method of any one of claims 1-8.

Citation Information

Patent Citations

  • Light unmanned aerial vehicle aerial image small target detection method and system based on TFA-YOLO11

    CN119810409A

  • Multi-scale cultivated land change detection system based on frequency domain conversion

    CN119832439A

  • Target detection method for traffic sign board, electronic equipment and medium

    CN120198884A

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1