ADLF-Net-based traffic sign detection method
Through the adaptive feature extraction and fusion module of the ADLF-Net network, the accuracy and efficiency of traffic sign detection in complex scenarios are solved, and efficient traffic sign detection is achieved.
Patent Information
- Application Number
- CN202510677289.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-26
AI Technical Summary
Existing traffic sign detection technology is difficult to take into account detection accuracy, environmental adaptability and computing efficiency in complex scenarios. Traditional networks lack adaptive feature extraction, attention mechanism and feature fusion capabilities, resulting in insufficient detection accuracy under small targets and extreme lighting conditions.
Adaptive dual-path local global fusion network (ADLF-Net) is adopted, including nonlinear adaptive feature extraction module (NSFE), dynamic channel calibration attention module (DCCA) and local global feature fusion module (LGFF). Feature expression and fusion capabilities are improved through wide-narrow parallel structure, dual-view attention fusion and multi-grained parallel attention mechanism.
It significantly improves the accuracy and environmental adaptability of traffic sign detection, reduces the computational complexity, meets the real-time detection needs, and realizes efficient small-objective detection in complex scenarios.
Smart Images

Figure CN120580667A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a traffic sign detection method based on ADLF-Net. Background Art
[0002] With the rapid development of autonomous driving and advanced driver assistance systems, accurate and real-time traffic sign detection technology has become a core requirement in intelligent transportation systems. However, the complexity and variability of real-world traffic scenes poses a huge challenge to detection systems. Traffic signs often occupy only a very small area of the image when photographed from a distance; complex weather and lighting conditions lead to reduced contrast and color distortion of signs; occlusion and defacement are widespread; and there are many visual interference elements in urban environments.
[0003] In recent years, deep learning-based object detection algorithms such as the YOLO series have made significant progress in general object detection tasks. However, when faced with specific scenarios such as traffic signs with small targets and high precision requirements, existing detection networks still face some technical bottlenecks. Traditional convolutional modules generally use fixed-form activation functions and standard convolutional structures, lacking the ability to adaptively capture weak visual features in complex scenes, resulting in insufficient detection accuracy for small traffic signs at a distance or in low-light conditions. Existing attention mechanisms typically use a single path and fixed-scale channel attention calculation, lacking the ability to collaboratively model local channel interactions and global channel context, making it difficult to accurately identify and enhance key feature channels in complex background interference. The simple splicing or weighted averaging operations of traditional neck networks ignore the differences in semantic strength and spatial resolution of features at different levels, and fail to effectively model the complex interactions between features, resulting in an imbalance between detailed information and contextual understanding, affecting detection accuracy and environmental adaptability.
[0004] To address these issues, academia and industry have proposed various improvements, each with varying degrees of limitations. Lightweight networks such as MobileNet reduce computational complexity through depthwise separable convolutions, but this also reduces feature representation capabilities, resulting in insufficient capture of subtle features of traffic signs under complex lighting conditions. Kolmogorov-Arnold networks introduce high-order nonlinear representations, but their application to vision tasks remains imperfect and lacks the ability to effectively integrate spatial feature blending. ECA-Net and CBAM methods attempt to optimize attention computation, but are still limited to modeling feature relationships at a single granularity, lacking the coordinated optimization of local channel interactions and global channel dependencies, and are therefore unable to adapt to the dynamic characteristics of traffic signs in diverse environments. Methods such as FPN, PANet, and ASFF propose top-down and bottom-up feature fusion strategies, but the fusion operation remains overly simplistic, failing to consider semantic differences between features at different levels of abstraction and lacking an adaptive assessment mechanism for feature quality. While these methods each have their strengths, they often struggle to strike a balance between detection accuracy, environmental adaptability, and computational efficiency in specialized detection scenarios such as traffic signs. Summary of the Invention
[0005] (1) Technical problems solved
[0006] To address the shortcomings of existing technologies, this paper proposes an Adaptive Dual-path Local-global Fusion Network (ADLF-Net). By organically integrating three innovative modules, it constructs a high-performance and environmentally adaptable detection framework, effectively solving the key technical problems of sign detection in complex traffic scenarios.
[0007] (2) Technical solution
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: A traffic sign detection method based on ADLF-Net, comprising the following steps:
[0009] Step 1: Construct an ADLF-Net network, which sequentially includes a nonlinear adaptive feature extraction module (NSFE), a dynamic channel calibration attention module (DCCA), and a local-global feature fusion module (LGFF);
[0010] Step 2: Input the image to be detected into the ADLF-Net network, and extract the enhanced multi-level features through the NSFE module;
[0011] Step 3: Use the DCCA module to perform channel calibration on the multi-level features to generate dynamically enhanced channel attention weights;
[0012] Step 4: The calibrated features are fused at multiple scales through the LGFF module to generate comprehensive features that fuse local details and global context.
[0013] Step 5: Output the category and location information of the traffic sign based on the comprehensive features.
[0014] Preferably, the NSFE module adopts a wide-narrow parallel structure, including:
[0015] Standard convolution path, which preserves basic features through short-circuit connections;
[0016] The deformation activation path, which consists of a deformation activation unit, some convolutional layers, and a random path dropout layer, is used to mine deep nonlinear relationships in the input data;
[0017] Among them, the deformation activation unit enhances the nonlinear expression ability of the model through adaptive deformation activation function and partial convolution mechanism.
[0018] Preferably, the DCCA module implements channel calibration through a dual-view attention fusion mechanism, specifically including:
[0019] Local path: uses a one-dimensional convolution branch, whose convolution kernel size is dynamically adjusted as a function of the number of input channels;
[0020] Global path: Modeling full channel dependencies through fully connected layers;
[0021] Fusion strategy: Local and global attention weights are fused through matrix multiplication and weight sharing strategy to generate dynamic channel attention weights.
[0022] Preferably, the LGFF module realizes multi-scale fusion through a multi-granularity parallel attention mechanism, including:
[0023] Attention units with different receptive field sizes are applied to the input features simultaneously to extract local details and global semantic information respectively;
[0024] The information flow weights of local and global features are adjusted through the similarity adaptive gating mechanism;
[0025] The re-parameterized convolution technology is used to enhance the expressive power through a multi-branch structure in the training phase, and is equivalently converted to a single convolution layer in the inference phase to maintain computational efficiency.
[0026] Preferably, the ADLF-Net network is built based on the YOLO framework, and the network parameters are optimized through end-to-end training. The output detection results include: the bounding box coordinates of the traffic sign, the category confidence, and the classification label.
[0027] Preferably, in the deformable activation path, the convolution kernel size of some convolutional layers is dynamically adapted to the spatial dimension of the input feature, and some neuron connections are randomly shielded through the random path inactivation layer to enhance feature robustness.
[0028] Preferably, the similarity adaptive gating mechanism generates a gating coefficient by calculating the correlation between local features and global features to weightedly fuse local and global features.
[0029] Preferably, the multi-level features are composed of the outputs of convolutional layers of different depths of the network, and the NSFE module enhances the nonlinear expression capability of the features layer by layer through a wide and narrow parallel structure.
[0030] (3) Beneficial effects
[0031] The present invention provides a traffic sign detection method based on ADLF-Net. It has the following beneficial effects:
[0032] (1) Through nonlinear adaptive feature extraction and dynamic channel calibration attention mechanism, the detection accuracy of traffic signs is significantly improved, and the performance in complex scenes is better than the existing mainstream methods.
[0033] (2) In response to low light, occlusion and complex background interference, ADLF-Net shows stronger environmental adaptability through adaptive deformation activation function and local-global feature fusion mechanism, effectively alleviating the performance degradation problem of traditional methods under extreme conditions.
[0034] (3) The network adopts lightweight design and heavy parameterized convolution technology, which significantly reduces the number of model parameters while maintaining efficient inference speed, achieving a balance between performance and computing resources to meet real-time detection needs.
[0035] (4) Through the multi-granularity parallel attention mechanism and adaptive gating strategy, ADLF-Net can efficiently integrate features at different levels, enhance the ability to capture small targets at a distance and detailed features, and improve the comprehensiveness of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a schematic diagram of the overall structure of the ADLF-Net traffic sign detection network proposed in the present invention;
[0037] Figure 2 Schematic diagram of the structure of the nonlinear adaptive feature extraction (NSFE) module of the present invention;
[0038] Figure 3 Schematic diagram of the structure of the dynamic channel calibration attention (DCCA) module in the present invention;
[0039] Figure 4It is a structural diagram of the local global feature fusion (LGFF) module in the present invention. DETAILED DESCRIPTION
[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0041] The following combination Figures 1 to 4 , the present invention is described in further detail.
[0042] See Figure 1 , this paper proposes a traffic sign detection method based on an Adaptive Dual-path Local-global Fusion Network (ADLF-Net). This method addresses the shortcomings of existing traffic sign detection algorithms in small target recognition, environmental adaptability, and computational efficiency, and innovates from the perspective of network architecture design. An end-to-end detection framework is constructed with three original modules as the core: a nonlinear adaptive feature extraction (NSFE) module, a dynamic channel calibration attention (DCCA) module, and a local-global feature fusion (LGFF) module. Among them, Adaptive corresponds to the adaptive deformation activation characteristics of the NSFE module, Dual-path corresponds to the local-global dual-path interaction strategy of the DCCA module, and Local-global Fusion corresponds to the local-global feature fusion mechanism of the LGFF module. The NSFE module improves the feature expression capability of the network by introducing an adaptive deformation activation function and a partial convolution mechanism; the DCCA module adopts an innovative local-global dual-path interaction strategy to achieve efficient dynamic calibration of channel information; and the LGFF module solves the key problem of multi-scale feature fusion through multi-view parallel attention. The ADLF-Net detection method can improve the accuracy and robustness of traffic sign detection while maintaining low computational complexity.
[0043] See Figure 2The Nonlinear Adaptive Feature Extraction (NSFE) module utilizes a "wide-narrow parallel + adaptive activation" architecture design, demonstrating a new approach to feature processing. After initial processing of the input feature map, the system splits it into two complementary paths: one maintains the coherence and stability of the underlying information flow through standard convolution; the other, consisting of multiple series-connected deformation activation units, is dedicated to exploring deep nonlinear relationships in the input data. This split-and-integrate architecture enables the module to simultaneously acquire feature representations at different levels, preserving the integrity of the original visual features while enhancing the model's ability to recognize complex patterns.
[0044] Specifically, the input features are first processed by two standard convolutional layers, cv1 and cv2. Subsequently, the feature stream is split into two parallel paths: one path, based on the consistency of the input and output dimensions, adjusts the dimensions through an additional Conv layer, forming a short-circuit connection (shortcut=x) to preserve the basic feature information; the other path focuses on mining deep nonlinear features, sequentially passing through the Partial_conv3 partial convolutional layer to reduce computational complexity, the MLP layer that introduces an adaptive deformation activation function to enhance nonlinear expression capabilities, and the drop_path layer for random path dropout regularization. The outputs of these two paths are initially fused through element-by-element addition. The fused features are concatenated with the output features of the cv1 layer along the channel dimension, integrating feature information from different levels. Finally, the concatenated features are passed through the cv3 convolutional layer for final feature extraction and dimensionality adjustment to generate the module's output features. The "wide-narrow parallel + adaptive activation" architecture, combined with partial convolution, parameterized activation and residual connection ideas, effectively improves the model's ability to sensitively capture complex visual patterns, especially the faint features in traffic signs, while taking into account computational efficiency.
[0045] See Figure 3 The DCCA module's design considers the multi-level dependencies of channel information and adopts a "local-global dual-path interaction" processing architecture. Unlike traditional channel attention mechanisms that focus only on a single information granularity, the module constructs two complementary feature processing paths in parallel, focusing on channel correlations at different scales.
[0046] The module first performs a global average pooling operation on the input feature map, compressing the spatial dimensions into a one-dimensional channel descriptor x. Subsequently, x is fed into two parallel processing branches: the first path uses a one-dimensional convolutional layer. Its key innovation lies in the dynamic adjustment of the convolution kernel size with the number of input channels, which can adaptively adjust the local receptive field, accurately capture the local correlation patterns between adjacent channels, and generate local features x1. The second path processes x through a fully connected layer, aiming to establish a channel-wide dependency model, obtain more macroscopic global semantic context information, and generate global features x2.
[0047] During the information fusion process, the DCCA module applies an innovative "dual-view attention fusion" mechanism. Through matrix multiplication and subsequent processing, two complementary attention weight maps are generated based on x1 and x2, respectively, emphasizing features from different perspectives: one perspective focuses on local patterns supported by global context, while the other perspective highlights local details within the global pattern, jointly constructing a comprehensive channel importance assessment system. After the weight maps are fused through the Mix operation, a weight sharing strategy is adopted to refine the fused attention to achieve a balance between effectiveness and computational efficiency, resulting in the final channel attention weights out. This refinement step not only reduces the number of parameters but also makes attention calculation more consistent and efficient.
[0048] Finally, the refined channel attention weights (out) are applied to the original input features (Input) through element-by-element multiplication, achieving adaptive enhancement and suppression of each channel's features. The resulting output feature map is the result of calibration by the DCCA module, effectively focusing on key information channels. By combining local-global information interaction with dual-view fusion, the DCCA module enables the network to more flexibly, accurately, and efficiently adapt to different levels of visual information, improving overall performance.
[0049] See Figure 4 In object detection systems, the effective fusion of multi-scale features is crucial for accurately localizing and identifying objects of varying sizes. Traditional feature fusion methods often use simple concatenation or weighted averaging strategies, which often exhibit limitations when processing objects with large scale variations and distinct semantic differences, such as traffic signs. The Local-Global Feature Fusion (LGFF) module utilizes a multi-granularity attention mechanism and adaptive feature integration techniques to enhance the interaction between features at different levels.
[0050] The LGFF module utilizes a "three-way parallel + fine-grained integration" approach to process feature inputs from different depth levels of the network. To ensure semantic compatibility of features, the module first remaps the input features to their channel dimensions through point convolution, creating a unified feature representation space. From this foundation, the module then parallelizes three complementary processing paths: a base path extracts and preserves the essential information of the input features through simple addition and convolution operations; and two enhancement paths apply specially designed attention mechanisms to the features, enhancing their expressive power from different perspectives.
[0051] LGFF uses a "multi-receptive field parallel processing" implementation. Unlike conventional attention calculations that only consider a single granularity, the module simultaneously applies attention units with different receptive field sizes to each feature map, creating a complementary feature enhancement effect: Attention units with a smaller receptive field (p = 2) focus on capturing local details and texture information, which is important for recognizing detailed features such as pictograms and text on traffic signs; attention units with a larger receptive field (p = 4) focus on extracting broader environmental context and semantic associations, enabling the model to understand the scene context of the target. This multi-granularity parallel processing strategy allows the module to simultaneously consider both microscopic details and macroscopic structure, forming a more comprehensive feature representation.
[0052] To enhance the discriminative power of fused features, the module also investigates a similarity-based adaptive gating mechanism. By introducing a trainable parameterized vector and establishing a feature relevance evaluation criterion, the system automatically adjusts the weight of information flow based on the similarity between the features and pre-set patterns. This effectively enhances the response of target-related features while mitigating the impact of background interference. In the final feature integration stage, the outputs of the three processing paths are fused and refined through channel-wise concatenation and a specially designed reparameterized convolution, forming a comprehensive feature representation rich in multi-level information.
[0053] Reparameterized convolution, the final step in this module, significantly enhances the model's expressiveness during training while maintaining computational efficiency during inference. This approach utilizes a multi-branch structure during training to enhance learning capabilities, while effectively converting to a single, efficient convolution during inference, cleverly balancing performance and efficiency.
[0054] The following experiments further illustrate the effectiveness of this invention. To fully validate the effectiveness of our proposed method, we designed and conducted a systematic experimental evaluation, including the determination of experimental evaluation metrics, the implementation of ablation and comparative experiments, and the quantitative and qualitative analysis of the results. The following is a detailed description of the experimental evaluation and results.
[0055] Experimental evaluation indicators
[0056] The present invention uses the following indicators to comprehensively evaluate traffic sign detection performance:
[0057] Precision P (%): The ratio of the number of correctly detected samples in the test results to all test results, reflecting the accuracy of the model prediction.
[0058]
[0059] Among them, TP (True Positive) represents the number of samples detected correctly, and FP (False Positive) represents the number of samples detected incorrectly.
[0060] Recall rate R (%): The ratio of the number of correctly detected samples to all targets that should be detected, reflecting the comprehensiveness of the model in capturing targets.
[0061]
[0062] Among them, FN (False Negative) represents the number of missed samples.
[0063] Average precision mAP50 (%): The average precision at an IoU threshold of 0.5 is the most commonly used comprehensive performance indicator in object detection tasks, used to evaluate the detection performance of the model on all categories.
[0064]
[0065] Where C represents the number of categories, It represents the average precision of the i-th category when the IoU threshold is 0.5, and is calculated as the area under the precision-recall curve:
[0066]
[0067] Where, P i (r) represents the precision when the recall rate is r.
[0068] Inference speed FPS: The number of image frames that can be processed per second, used to evaluate the real-time performance of the algorithm in practical applications.
[0069]
[0070] Where N is the total number of images processed, T i is the time required to process the i-th image.
[0071] Parameters (M): The total number of parameters in the model, in millions (M).
[0072] Experimental content and result analysis
[0073] The experiments in this paper were conducted on publicly available traffic sign datasets, including TT100K and GTSDB. To ensure the reliability of the experiments, we adopted a standard training-validation-testing split and used data augmentation techniques such as random cropping, rotation, and noise to simulate various real-world scenarios.
[0074] All experiments were conducted on the same hardware configuration (NVIDIA RTX 3060 GPU, Intel i38100) to ensure fair performance comparison. The basic detection framework uses YOLOv8, and the training hyperparameters are set to: batch size 32, total training epochs 300.
[0075] Table 1. Ablation test results of each module of the present invention
[0076] Model Configuration P(%) R(%) mAP50 (%) FPS Parameter quantity (M) Baseline Model 87.5 88.9 88.2 65.3 11.2 +NSFE 89.7 91.1 90.4 62.8 11.8 +DCCA 91.0 92.5 91.7 60.6 12.3 +LGFF 92.5 93.8 93.2 58.2 13.0 ADLF-Net 94.1 95.2 94.6 56.5 13.5
[0077] To validate the effectiveness of the core modules of our invention, we conducted detailed ablation experiments, with the results shown in Table 1. These results show that with the gradual introduction of each key module, model performance steadily improves. The introduction of the NSFE module improves the model's mAP50 by 2.2 percentage points (from 88.2% to 90.4%), while also increasing precision and recall by 2.2% and 2.2%, respectively, demonstrating that the adaptive deformation activation function enhances the network's ability to capture complex features. Furthermore, the addition of the DCCA module increases mAP50 to 91.7%, further validating the effectiveness of the local-global dual-path interaction strategy in enhancing key channel responses, particularly for detecting small objects at long distances. The integration of the LGFF module improves mAP50 to 93.2%, fully demonstrating the effective integration of multi-scale features using the multi-granularity attention mechanism and adaptive feature integration techniques. The final complete method achieves a mAP50 of 94.6%, a 6.4 percentage point improvement over the baseline model, with precision and recall reaching 94.1% and 95.2%, respectively. It is worth noting that although the number of parameters increased by 20.5% (11.2M to 13.5M), the improvement in detection performance far exceeded the increase in parameters, and the inference speed remained at 56.5FPS, meeting the needs of real-time applications, indicating that the present invention has achieved an excellent balance between performance and computing efficiency.
[0078] Table 2. Performance comparison between the present invention and existing methods
[0079] method P(%) R(%) mAP50 (%) FPS Parameter quantity (M) Faster R-CNN 88.6 90.3 89.5 28.6 41.5 RetinaNet 89.2 91.0 90.1 32.7 36.4 YOLOv5-M 90.5 92.2 91.4 62.3 21.2 YOLOv7 91.8 93.5 92.7 60.1 36.9 YOLOv8-M 92.3 93.9 93.1 63.5 25.9 ADLF-Net 94.1 95.2 94.6 56.5 13.5
[0080] At the same time, we compared the performance of the present invention with the current mainstream traffic sign detection method, and the results are shown in Table 2. The comparative experimental results show that the present invention exhibits significant advantages in both detection performance and computing resource utilization. In terms of detection indicators, the present invention achieves 94.6% mAP50, while the precision and recall rates reach 94.1% and 95.2% respectively, both exceeding the current state-of-the-art YOLOv8-M method (mAP50 is 93.1%). Particularly outstanding is the significant optimization of the present invention in terms of parameter quantity, which is only 13.5M, approximately 52.1% of YOLOv8-M (25.9M), significantly reducing the computational burden while maintaining the leading detection accuracy. From the perspective of inference efficiency, although the 56.5FPS of the present invention is slightly lower than the 63.5FPS of YOLOv8-M, it still far exceeds the real-time application requirements (25FPS), and considering the improvement in detection accuracy and the significant reduction in resource usage, this slight speed difference is reasonable and acceptable. Furthermore, in complex scene adaptability tests, the proposed method achieved performance degradation of 8.3% and 9.1% in low-light and partial occlusion conditions, respectively, significantly outperforming YOLOv8-M's 12.5% and 13.2%, demonstrating the enhanced environmental robustness achieved through the proposed method's innovative module design. These results demonstrate the proposed method's comprehensive technical advantages in traffic sign detection tasks, making it particularly suitable for deployment in resource-constrained real-world applications.
[0081] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A traffic sign detection method based on ADLF-Net, characterized in that: The following steps are involved: Step 1: Construct an ADLF-Net network, which sequentially includes a nonlinear adaptive feature extraction module (NSFE), a dynamic channel calibration attention module (DCCA), and a local-global feature fusion module (LGFF); Step 2: Input the image to be detected into the ADLF-Net network, and extract the enhanced multi-level features through the NSFE module; Step 3: Use the DCCA module to perform channel calibration on the multi-level features to generate dynamically enhanced channel attention weights; Step 4: The calibrated features are fused at multiple scales through the LGFF module to generate comprehensive features that fuse local details and global context. Step 5: Output the category and location information of the traffic sign based on the comprehensive features.
2. A traffic sign detection method based on ADLF-Net according to claim 1, characterized in that: The NSFE module adopts a wide-narrow parallel structure, including: Standard convolution path, which preserves basic features through short-circuit connections; The deformation activation path, which consists of a deformation activation unit, some convolutional layers, and a random path dropout layer, is used to mine deep nonlinear relationships in the input data; Among them, the deformation activation unit enhances the nonlinear expression ability of the model through adaptive deformation activation function and partial convolution mechanism.
3. A traffic sign detection method based on ADLF-Net according to claim 1, characterized in that: The DCCA module implements channel calibration through a dual-view attention fusion mechanism, specifically including: Local path: uses a one-dimensional convolution branch, whose convolution kernel size is dynamically adjusted as a function of the number of input channels; Global path: Modeling full channel dependencies through fully connected layers; Fusion strategy: Local and global attention weights are fused through matrix multiplication and weight sharing strategy to generate dynamic channel attention weights.
4. A traffic sign detection method based on ADLF-Net according to claim 1, characterized in that: The LGFF module achieves multi-scale fusion through a multi-granularity parallel attention mechanism, including: Attention units with different receptive field sizes are applied to the input features simultaneously to extract local details and global semantic information respectively; The information flow weights of local and global features are adjusted through the similarity adaptive gating mechanism; The re-parameterized convolution technology is used to enhance the expressive power through a multi-branch structure in the training phase, and is equivalently converted to a single convolution layer in the inference phase to maintain computational efficiency.
5. A traffic sign detection method based on ADLF-Net according to claim 1, characterized in that: The ADLF-Net network is built based on the YOLO framework, and the network parameters are optimized through end-to-end training. The output detection results include: the bounding box coordinates of the traffic sign, the category confidence, and the classification label.
6. A traffic sign detection method based on ADLF-Net according to claim 2, characterized in that: In the deformable activation path, the convolution kernel sizes of some convolutional layers are dynamically adapted to the spatial dimensions of the input features, and some neuron connections are randomly shielded through random path inactivation layers to enhance feature robustness.
7. A traffic sign detection method based on ADLF-Net according to claim 4, characterized in that: The similarity adaptive gating mechanism generates a gating coefficient by calculating the correlation between local features and global features to weightedly fuse local and global features.
8. A traffic sign detection method based on ADLF-Net according to claim 1, characterized in that: The multi-level features are composed of the outputs of convolutional layers of different depths in the network, and the NSFE module enhances the nonlinear expression capability of the features layer by layer through a wide-narrow parallel structure.
Citation Information
Patent Citations
Small target detection method and system based on visual attention mechanism
CN117274661A
Target detection method for shielded vehicles in urban road scene
CN119810418A
Double-feature fusion semantic segmentation system and method based on internet of things perception
WO2022227913A1