Target detection method for traffic sign board, electronic equipment and medium

By using the spatial channel collaborative attention mechanism and full-dimensional dynamic convolution method in traffic sign target detection, the problem of inefficiency of traditional manual patrols is solved, and more efficient and accurate traffic sign detection is achieved.

CN120198884APending Publication Date: 2025-06-24SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510163695.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional manual patrol methods are inefficient and easily missed, resulting in potential risks in road traffic safety.

Method used

The traffic sign target detection method based on the spatial channel synergistic attention mechanism and full-dimensional dynamic convolution is adopted. By constructing a target detection model, the full-dimensional dynamic convolution module is used to extract low-fine-grained information, and combined with the spatial channel synergistic attention mechanism to strengthen feature expression, realizing multi-scale feature fusion and target detection.

Benefits of technology

It improves the accuracy and efficiency of traffic sign target detection, reduces the fuzzy impact caused by motion, retains more detailed information, and enhances the model's feature extraction ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198884A_ABST
    Figure CN120198884A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic sign target detection method, electronic equipment and a medium, and the method comprises the steps: carrying out the feature extraction of a traffic sign image through a backbone network which comprises a full-dimension dynamic convolution module and a space channel collaborative attention module, and obtaining a preliminary feature; a spatial pyramid pooling fast module is used for obtaining features of different receptive fields, channel aggregation is carried out to obtain deep features, feature enhancement is carried out through a spatial channel collaborative attention module, and then the obtained deep features are input into a bidirectional fusion network module for feature fusion. And finally, carrying out target classification and bounding box regression on the feature representations of different scales through a convolution module and a nonlinear activation function. According to the method, the space channel collaborative attention module and the full-dimensional dynamic convolution module are adopted, the model reasoning speed and accuracy are improved while multi-scale feature extraction is achieved, and low-fine-granularity information is fully extracted to help detection of traffic signs in images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to an object detection method, an electronic device, and a medium for traffic signs. Background Art

[0002] Object recognition technology plays a core role in the field of computer vision. It is responsible for identifying and determining the location of the detected object and is widely used in multiple scenarios such as driverless, medical imaging analysis, and intelligent video surveillance. With the current high attention to road safety and the increasing urban traffic flow, traffic signs, as a key element to ensure road traffic safety, play an important role. However, due to the large number and diverse types of traffic signs, the traditional manual inspection method is not only inefficient but also prone to omission, which brings potential risks to road safety. Summary of the Invention

[0003] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide an object detection method, an electronic device, and a medium for traffic signs based on a spatial-channel collaborative attention mechanism and a full-dimensional dynamic convolution.

[0004] The first technical solution adopted by the present invention is as follows:

[0005] An object detection method for traffic signs includes the following steps:

[0006] Obtain a traffic sign image, preprocess the obtained traffic sign image, and construct a training set;

[0007] Construct an object detection model;

[0008] Use the training set to train the object detection model, and use the trained object detection model for traffic sign image object detection;

[0009] Among them, the object detection model works as follows:

[0010] A backbone network composed of a full-dimensional dynamic convolution and a normal convolution module extracts features from the image to obtain shallow-layer features;

[0011] Use a spatial pyramid pooling fast module to process the input features;

[0012] Use a spatial-channel collaborative attention module to strengthen the feature expression to obtain deep-layer features;

[0013] Input the deep-layer features into a bottom-up fusion network module for multi-scale feature fusion to obtain feature representations of different scales;

[0014] Classify the objective function and regress the bounding box for features of different scales through a convolutional module and a non - linear activation function.

[0015] Furthermore, the step of preprocessing the obtained traffic sign image includes performing enhancement processing on the image data:

[0016] Use the Mosaic method, Mixup method, random horizontal flipping, adding noise or translation method to perform data enhancement processing on the traffic sign image;

[0017] Among them, in the Mosaic method, multiple images are randomly cropped and then stitched into one picture; in the Mixup method, two samples and label data are added proportionally to generate new samples and label data.

[0018] Furthermore, in the backbone network, the full - dimensional dynamic convolution module ODConv and the C3K2 module are fused to form the C3K2_ODConv module;

[0019] The C3K2 module introduces a multi - scale convolutional kernel C3K. This design expands the receptive field, enabling the model to have more extensive context information;

[0020] The full - dimensional dynamic convolution module ODConv calculates four types of attention along all four dimensions of the convolutional kernel space. Such a design allows ODConv to perform fine - grained dynamic adjustment in four dimensions: spatial size, number of input channels, number of filters (number of output channels), and number of convolutional kernels.

[0021] Furthermore, the calculation process expression of the spatial pyramid pooling fast module is as follows:

[0022] X1 = CBS(X)

[0023] X2 = MP(X1)

[0024] X3 = MP(X2)

[0025] X4 = MP(X3)

[0026] f = Concat(X1,X2,X3,X4)

[0027] Y = CBS(f)

[0028] Among them, X represents the input feature map, CBS represents convolution, batch normalization, and SiLU is a non - linear activation function. X1, X2, X3, X4, f are intermediate - layer features, Concat represents the feature concatenation operation, and Y is the output feature.

[0029] Furthermore, strengthening the feature expression by using the spatial - channel co - attention module includes:

[0030] Fuse spatial attention and channel attention, and then fuse them with the C2 module to form the C2_SCSA module, which serves as the spatial-channel collaborative attention module;

[0031] The C2_SCSA module contains a C2 module. Inside the C2 module, add the SCSA module to the bottleneck convolutional module in the C2 module and add a residual connection;

[0032] The calculation process expression of the bottleneck convolutional module is as follows:

[0033] Bottleneck = CBS(CBS(X)) + X

[0034] CBS = SiLU(BN(Conv 3×3,1 (X)))

[0035] Among them, X represents the input feature map, BN represents batch normalization, SiLU is a non-linear activation function, and Conv 3×3,1 is a one-dimensional 3×3 convolution.

[0036] Furthermore, the SCSA module includes a multi-semantic spatial attention module SMSA and a progressive channel attention module PCSA;

[0037] The multi-semantic spatial attention module SMSA uses multi-scale depth-shared 1D convolutions to capture multi-semantic spatial information and enhance local and global feature representations;

[0038] The progressive channel attention module PCSA adopts input-aware self-attention to refine channel features, reduce semantic differences, and ensure strong feature integration between channels.

[0039] Furthermore, the multi-semantic spatial attention module SMSA operates on the input feature map B×C×H×W, where B represents the batch size, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map; this module first decomposes the feature map along the height and width dimensions and performs average pooling operations on each dimension; subsequently, divides the feature set into 4 equally sized and independent sub-features; then, applies depthwise separable one-dimensional convolutions (DWConv1d) with kernel sizes of 3, 5, 7, and 9 to these 4 sub-features respectively; after that, performs batch normalization (BN) operations on the convolved sub-features; finally, uses the Sigmoid activation function to generate a spatial attention feature map to activate and suppress specific spatial regions, and multiplies the one-dimensional feature vectors of the two dimensions to obtain the updated feature map;

[0040] The progressive channel attention module PCSA takes the feature map output by the multi-semantic space attention module as input; this module adopts a single-head self-attention (SHSA) mechanism, enabling the model to dynamically focus on other parts of the sequence when processing each input element.

[0041] Further, inputting the deep features into the bottom-up fusion network module for multi-scale feature fusion to obtain feature representations of different scales, including:

[0042] Upsample the deep features and continuously fuse them with the shallow features. Use the C3K2 module to extract features from the fused shallow feature map, and combine the extracted features with the deep features, thereby generating three feature representations with different scales and rich semantic information.

[0043] The second technical solution adopted by the present invention is:

[0044] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the target detection method of a traffic signboard as described above.

[0045] The third technical solution adopted by the present invention is:

[0046] A computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the target detection method of a traffic signboard as described above.

[0047] The fourth technical solution adopted by the present invention is:

[0048] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.

[0049] The beneficial effects of the present invention are as follows: By adopting the full-dimensional dynamic convolution module, the present invention fully extracts low-level and fine-grained information, reduces the blurring effect caused by motion, and retains more detailed information to assist in the detection of small targets in images. Additionally, in order to improve the model inference speed and accuracy while effectively extracting features, the cross-stage local network structure and the efficient multi-scale attention are combined together to reduce the computational amount while achieving a richer gradient combination. By adopting the spatial-channel collaborative attention mechanism module, the spatial information and semantic information in the feature map are extracted and enhanced to weaken irrelevant interference factors, enhance the feature extraction ability of the model, and improve the target detection ability of the model for traffic signs. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings below only conveniently and clearly illustrate some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a schematic diagram of the traffic sign target detection model based on the spatial-channel collaborative attention mechanism and the full-dimensional dynamic convolution in the embodiment of the present invention;

[0052] Figure 2 It is a schematic diagram of the structure of the C3K2 module in the embodiment of the present invention;

[0053] Figure 3 It is a schematic diagram of the structure of the C3K module in the embodiment of the present invention;

[0054] Figure 4 It is a schematic diagram of the structure of the full-dimensional dynamic convolution module in the embodiment of the present invention;

[0055] Figure 5 It is a schematic diagram of the structure of the bottleneck convolution module in the embodiment of the present invention;

[0056] Figure 6 It is a schematic diagram of the structure of the spatial pyramid pooling fast module in the embodiment of the present invention;

[0057] Figure 7 It is a schematic diagram of the structure of a kind of spatial-channel collaborative attention module in the embodiment of the present invention;

[0058] Figure 8 It is a step flow chart of a kind of traffic sign target detection method based on the spatial-channel collaborative attention mechanism and the full-dimensional dynamic convolution in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0060] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as up, down, front, back, left, right, etc., which relate to orientation description, is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0061] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is more than two, and understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0062] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.

[0063] Based on the existing technical problems, the present invention develops a traffic sign detection technology based on the YOLOV11 algorithm, aiming to quickly and accurately identify traffic signs on the road, which is of great significance for improving road safety and reducing traffic accidents. The present invention adopts the addition of a spatial-channel collaborative attention mechanism, fuses the C2 module and the SCSA module, and performs residual connection to design a new C2_SCSA module, which can improve the feature extraction ability of the model while reducing the computational complexity. In addition, when the vehicle travels from a distance to a close distance from the signboard, dynamic blur will occur. Using full-dimensional dynamic convolution in the backbone network, there are multiple convolutional kernels, which can dynamically adjust the corresponding convolutional kernels according to the characteristics of each input sample to fully extract low-level fine-grained information. In addition, in the neck part, feature fusion is adopted to fuse the features between the underlying and deep layers with different semantics.

[0064] Embodiment 1

[0065] As Figure 8 shown, this embodiment provides an object detection method for traffic signs based on a spatial channel collaborative attention mechanism and full-dimensional dynamic convolution, including the following steps:

[0066] S1. Obtain traffic sign images and preprocess the obtained traffic sign images.

[0067] Obtain a traffic sign dataset, label and classify the images, divide the traffic sign image dataset into a training set, a test set and a validation set, and preprocess and augment the data of the traffic sign images.

[0068] Specifically, data augmentation for the dataset includes:

[0069] Adopt the Mosaic method, the Mixup method, random horizontal flipping, adding noise or translation methods to perform data augmentation processing on the images in the dataset;

[0070] Among them, in the Mosaic method, multiple pictures are randomly cropped and then stitched into one picture; in the Mixup method, two samples and label data are added proportionally to generate new samples and label data.

[0071] In some embodiments, in order to simulate different weather conditions, we classify the dataset according to different weather conditions. These weather conditions include but are not limited to:

[0072] Sunny day: Sufficient sunlight and high visibility.

[0073] Cloudy day: The sky is covered by clouds and the light is weak.

[0074] Rainy day: There is precipitation, low visibility and slippery road surface.

[0075] Foggy day: Extremely low visibility and a large amount of water vapor in the air.

[0076] Snowy day: There is snowfall, the road surface is icy and the visibility is low.

[0077] Through this classification, we can better understand the performance of the model under different weather conditions and perform targeted enhancement processing, such as Gaussian blur and dynamic blur. In addition, to improve the performance of the traffic sign detection model, we performed preprocessing operations on the input images, including Mosaic data augmentation, Mixup data augmentation, random horizontal flipping, and translation. The Mosaic method increases sample diversity by randomly cropping four images and then stitching them together into a new image. Mixup adds two samples and their labels in proportion to generate new sample and label data to enhance the generalization ability of the model. At the same time, all input images are adjusted to a size of 640×640×3 before augmentation to ensure the consistency of model input.

[0078] S2. The backbone network composed of the full-dimensional dynamic convolution and ordinary convolution modules extracts features from the image to obtain shallow features.

[0079] Such as Figure 1 As shown, the backbone network includes a series of ordinary convolutions, full-dimensional dynamic convolution modules, and C3K2_ODConv modules to obtain image features at different levels. The C3K2 module and the ODConv module are fused, and this design endows ODConv with the ability to perform refined dynamic adjustment in four dimensions: spatial size, number of input channels, number of output channels, and number of convolutional kernels.

[0080] The so-called full-dimensional dynamic convolution module adopts a new type of multi-dimensional attention mechanism, which breaks through the limitation of traditional methods that only calculate a single attention scalar for each convolutional kernel, and instead calculates four different types of attention along all four dimensions of the convolutional kernel space: spatial size (asi), number of input channels (aci), number of output channels (afi), and number of convolutional kernels (awi). This design endows ODConv with the ability to perform refined dynamic adjustment in four dimensions: spatial size, number of input channels, number of output channels, and number of convolutional kernels. The so-called C3K2 module is an important feature extraction component. Compared with the standard C3 module, C3K2 introduces a multi-scale convolutional kernel C3K, and this design can expand the receptive field, enabling the model to have more extensive context information. The so-called C3K2_ODConv module modifies the bottleneck module in the C3K module, replaces the convolution with the ODConv module, and dynamically adjusts the size of the receptive field according to the size of the instance to obtain features at different levels.

[0081] Such as Figure 2 As shown, the C3K2 module is an important feature extraction component. Compared with the standard C3 module, C3K2 introduces a multi-scale convolutional kernel C3K, and this design can expand the receptive field, enabling the model to have more extensive context information.

[0082] See Figure 3 For the C3K module, the C3K module used in this embodiment is the C3K module in the case of N = 2, that is, the C3K2 module, which enhances the feature extraction ability.

[0083] As Figure 4 shown, for the all-dimensional dynamic convolution module ODConv, different from only adjusting the number of convolutional kernels, ODConv calculates four types of attention along all four dimensions of the convolutional kernel space. Such a design allows ODConv to perform fine-grained dynamic adjustment in the four dimensions of spatial size, number of input channels, number of filters (number of output channels), and number of convolutional kernels.

[0084] S3. Process the input features using the spatial pyramid pooling fast module.

[0085] With the help of the spatial pyramid pooling fast module SPPF, features with different receptive fields are extracted, thereby obtaining deep features.

[0086] As Figure 6 shown, using the spatial pyramid pooling fast module to process the input features, first pass through a 1×1 convolutional layer to retain more detailed information. Subsequently, the features pass through three pooling layers with different scales, and the pooled features are concatenated with the original features to obtain deep features with different receptive fields. The specific calculation process expression is as follows:

[0087] X1 = CBS(X)

[0088] X2 = MP(X1)

[0089] X3 = MP(X2)

[0090] X4 = MP(X3)

[0091] f = Concat(X1, X2, X3, X4)

[0092] Y = CBS(f)

[0093] Among them, X represents the input feature map, CBS represents convolution, batch normalization, and SiLU is a non-linear activation function. X1, X2, X3, X4, and f are intermediate layer features, Concat represents the feature concatenation operation, and Y is the output feature.

[0094] S4. Strengthen the feature expression using the spatial-channel collaborative attention module to obtain deep features.

[0095] The C2_SCSA module with spatial-channel collaborative attention, as Figure 7As shown, the C2_SCSA module contains a C2 module. Inside the C2 module, SCSA attention is added to the bottleneck convolutional module in the C2 module, and a residual connection is added. The SCSA module is composed of a multi-semantic space attention module (SMSA) and a progressive channel attention module (PCSA). Among them, SMSA uses multi-scale depth-shared 1D convolution to capture rich multi-semantic space information and strengthen the expression of local and global features; while PCSA uses an input-aware self-attention mechanism to optimize channel features, effectively reducing semantic differences and ensuring the tight integration of features between channels.

[0096] The so-called added spatial-channel attention collaboration module contains a multi-semantic space attention module (SMSA) and a progressive channel attention module (PCSA). SMSA uses multi-scale depth-shared 1D convolution to capture multi-semantic space information and enhance local and global feature representation. PCSA uses input-aware self-attention to refine channel features, effectively reducing semantic differences and ensuring strong feature integration between channels.

[0097] The so-called added multi-semantic space attention module is for the input feature map B×C×H×W, where B represents the batch, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map; the feature map is decomposed into 2 dimensions along the height and width respectively, average pooling is performed on each dimension, and the feature set is decomposed into 4 sub-features of the same size and independent. Then, the 4 sub-features are respectively convolved using DWConv1d with kernel sizes of 3, 5, 7, and 9, followed by BN normalization operation. Finally, the Sigmoid activation function is used to generate the spatial attention feature map, activate and suppress specific spatial regions, and multiply the one-dimensional feature vectors of the 2 dimensions to obtain a new feature map.

[0098] The so-called added progressive channel attention module takes the output feature map of the spatial attention mechanism module as input, first performs average pooling and batch normalization, reducing the subsequent computational cost; in order to enable the model to dynamically focus on other parts of the input sequence when processing each input element, the SHSA self-attention mechanism is adopted. Compared with using common convolutional operations to model channel dependencies, the progressive channel attention module shows stronger input-aware ability and effectively uses the spatial prior provided by the spatial attention module to deepen learning.

[0099] The so-called C2_SCSA module is exactly the fusion of the C2 module and SCSA. SCSA attention is added to the bottleneck convolutional module in the C2 module, and a residual connection is added.

[0100] See Figure 5 , the specific calculation process expression of the bottleneck convolutional module is as follows:

[0101] Bottleneck = CBS(CBS(X)) + X

[0102] CBS = SiLU(BN(Conv 3×3,1 (X)))

[0103] Among them, X represents the input feature map, BN represents batch normalization, and SiLU is a non - linear activation function.

[0104] S5. Input the deep features into the bottom - up fusion network module for multi - scale feature fusion to obtain feature representations of different scales.

[0105] Input the deep features obtained in step S4 into the bottom - up neck module for multi - scale fusion with the shallow - level features to obtain feature representations of different scales. The bidirectional fusion network consists of a series of upsampling and C3K2 modules. Refer to Figure 1 The entire fusion network is that the deep - level features are continuously upsampled and fused with the shallow - level features. After fusion, the C3K2 module is used to obtain more extensive context information, thereby obtaining three feature representations with different scales and rich semantic information.

[0106] The implementation process of the so - called bidirectional fusion network is that the deep - level features are upsampled and continuously fused with the shallow - level features, and then the fused shallow - level feature map is used to extract features through the C3K2 module and fused with the deep - level features, thereby obtaining three feature representations with different scales and rich semantic information.

[0107] S6. Classify the objective function and regress the bounding box for the features of different scales through the convolution module and the non - linear activation function.

[0108] Use the convolution module and the non - linear activation function to process the feature representations of different scales to perform the object classification and bounding box regression tasks. Subsequently, use the collected dataset to train the constructed object detection network, and apply the trained model to the object detection of traffic signs.

[0109] In the training stage of this embodiment, the SGD optimizer is used to update the network parameters, and the polynomial decay strategy is used to adjust the learning rate. The entire training process goes through 200 epochs. After each epoch ends, it is evaluated on the validation set, and the model weights when the loss function value on the validation set is the lowest are saved. When entering the test stage, first pre - process the test dataset, then input it into the previously saved best model for detection, and finally output the detection results.

[0110] Embodiment 2

[0111] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory and is loaded and executed by the processor to implement as Figure 8 a target detection method for traffic signs based on a spatial channel collaborative attention mechanism and full-dimensional dynamic convolution as shown.

[0112] It can be understood that the memory may include a random access memory (RAM) and may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area can store data created according to the use of the server, etc.

[0113] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a single chip.

[0114] Since this electronic device is the electronic device corresponding to the target detection method for traffic signs based on a spatial channel collaborative attention mechanism and full-dimensional dynamic convolution in an embodiment of the present invention, and the principle of this electronic device for solving problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0115] Example 3

[0116] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement as Figure 8 shown in a method for target detection of traffic signs based on a spatial channel collaborative attention mechanism and full-dimensional dynamic convolution.

[0117] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0118] Since this storage medium is the storage medium corresponding to the method for target detection of traffic signs based on a spatial channel collaborative attention mechanism and full-dimensional dynamic convolution in the embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0119] Example 4

[0120] In some possible embodiments, various aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a method for target detection of traffic signs based on a spatial channel collaborative attention mechanism and full-dimensional dynamic convolution according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in high-level programming languages such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0121] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0122] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.

[0123] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for detecting a traffic sign, characterized in that: The following steps are involved: Obtain traffic sign images, pre-process the acquired traffic sign images, and construct a training set; Build an object detection model; The target detection model is trained using the training set, and the trained target detection model is used for traffic sign image target detection; The target detection model works as follows: The backbone network composed of full-dimensional dynamic convolution and ordinary convolution modules extracts features from the image and obtains shallow-level features; Use spatial pyramid pooling fast module to process input features; Use spatial channel collaborative attention module to strengthen feature expression and obtain deep features; The deep features are input into the bottom-up fusion network module for multi-scale feature fusion to obtain feature representations of different scales; Features of different scales are passed through convolution modules and nonlinear activation functions for target function classification and bounding box regression.

2. The method for detecting a traffic sign according to claim 1, characterized in that: The step of preprocessing the acquired traffic sign image includes enhancing the image data: The traffic sign images are enhanced by using Mosaic method, Mixup method, random horizontal flipping, adding noise or translation method. Among them, in the Mosaic method, multiple images are randomly cropped and then spliced ​​into one picture; in the Mixup method, two samples and label data are added in proportion to generate new samples and label data.

3. The target detection method of a traffic sign according to claim 1, characterized in that: In the backbone network, the full-dimensional dynamic convolution module ODConv and the C3K2 module are fused to form the C3K2_ODConv module; the C3K2 module introduces a multi-scale convolution kernel C3K, which expands the receptive field and enables the model to have a wider range of context information; The full-dimensional dynamic convolution module ODConv calculates four types of attention along all four dimensions of the convolution kernel space. This design allows ODConv to perform fine-grained dynamic adjustments in four dimensions: spatial size, number of input channels, number of filters, and number of convolution kernels.

4. The method for detecting a traffic sign according to claim 1, characterized in that: The calculation process expression of the spatial pyramid pooling fast module is as follows: X1=CBS(X) X2=MP(X1) X3=MP(X2) X4=MP(X3) f=Concat(X1,X2,X3,X4) Y=CBS(f) Among them, X represents the input feature map, CBS represents convolution, batch normalization and SiLU is a nonlinear activation function, X1, X2, X3, X4, f are intermediate layer features, Concat represents feature concatenation operation, and Y is the output feature.

5. The method for detecting a traffic sign according to claim 1, characterized in that: The method of using the spatial channel collaborative attention module to enhance feature expression includes: The spatial attention and channel attention are fused and then fused with the C2 module into the C2_SCSA module as the spatial channel collaborative attention module; The C2_SCSA module contains a C2 module. Inside the C2 module, the bottleneck convolution module in the C2 module is added with the SCSA module, and a residual connection is added; The calculation process of the bottleneck convolution module is expressed as follows: Bottleneck=CBS(CBS(X))+X CBS=SiLU(BN(Conv 3×3,1 (X))) Among them, X represents the input feature map, BN represents batch normalization, SiLU is a nonlinear activation function, Conv 3×3,1 It is a one-dimensional 3×3 convolution.

6. The method for detecting a traffic sign according to claim 5, characterized in that: The SCSA module includes a multi-semantic spatial attention module SMSA and a progressive channel attention module PCSA; The multi-semantic spatial attention module SMSA uses multi-scale deep shared 1D convolution to capture multi-semantic spatial information and enhance local and global feature representation; The progressive channel attention module PCSA adopts input-aware self-attention to refine channel features, alleviate semantic differences and ensure strong feature integration across channels.

7. The method for detecting a traffic sign according to claim 6, characterized in that: The multi-semantic spatial attention module SMSA operates on an input feature map B×C×H×W, where B represents the batch size, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map; the module first decomposes the feature map along the height and width dimensions, and performs an average pooling operation on each dimension; then, the feature set is divided into four equal-sized and independent sub-features; then, a depth-separable one-dimensional convolution with kernel sizes of 3, 5, 7, and 9 is applied to the four sub-features respectively; then, a batch normalization operation is performed on the convolved sub-features; finally, a spatial attention feature map is generated using a sigmoid activation function to activate and inhibit specific spatial regions, and the one-dimensional feature vectors of the two dimensions are multiplied to obtain an updated feature map; The progressive channel attention module PCSA takes the feature map output by the multi-semantic spatial attention module as input; the module adopts a single-head self-attention mechanism, which enables the model to dynamically pay attention to other parts in the sequence when processing each input element.

8. The method for detecting a traffic sign according to claim 1, characterized in that: The deep features are input into the bottom-up fusion network module to perform multi-scale feature fusion to obtain feature representations of different scales, including: The deep features are upsampled and continuously fused with the shallow features. The C3K2 module is used to extract features from the fused shallow feature maps, and the extracted features are combined with the deep features to generate three feature representations with different scales and rich semantic information.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Defect detection method and system for engineering signboard based on deep learning

    CN120894282A

  • Defect detection method and system for engineering sign based on deep learning

    CN120894282B

  • Lightweight-based traffic sign detection method, storage medium and electronic equipment

    CN120913175A