Lightweight and small traffic sign recognition method for vehicle-mounted intelligent systems
The local prominent features and global expression information of micro traffic signs are extracted through the information flow logic propagation network, solving the accuracy and real-time problems of micro traffic sign detection in the on-board intelligent system, and achieving efficient detection under the weak computing power of the on-board intelligent chip.
Patent Information
- Application Number
- CN202310838014.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-07-10
AI Technical Summary
When detecting micro traffic signs, existing in-vehicle intelligent systems are difficult to meet the actual needs of few characteristics and high positioning requirements. Due to the insufficient computing power of the in-vehicle intelligent chip, traditional target detection methods cannot effectively detect micro traffic signs.
The information flow logical propagation network is used to extract the local prominent features and global expression information of micro traffic signs. Through compression image processing, edge information optimization and graph convolution flow interaction, feature layers are fused layer by layer to improve detection accuracy.
It realizes accurate detection of micro traffic signs under the conditions of weak computing power of the on-board smart chip, improving the accuracy and real-time detection.
Smart Images

Figure CN116863443B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation and relates to a lightweight micro traffic sign recognition method for an on-board intelligent system. Background Art
[0002] In-vehicle intelligent system target detection aims to identify all targets of interest while the vehicle is in motion through detection algorithms and determine their location and category. In detection tasks, object features serve as a crucial basis for detection results. However, in real-world environments, small traffic signs have limited available features and demanding positioning requirements. Furthermore, due to hardware limitations in the vehicle environment and the limited computing power of onboard intelligent chips, traditional target detection methods struggle to meet the practical needs of this scenario. To address the difficulty in detecting small traffic signs, an effective strategy involves using methods such as feature enhancement and spatial information association to obtain reliable small target features and spatial position relationships to predict the type and location of small traffic signs. Lightweight networks are also employed to reduce the burden on onboard intelligent chips, meeting the real-time detection requirements.
[0003] To improve the accuracy of small traffic sign detection and adapt to the limited computing power of onboard smart chips, this paper proposes a small traffic sign detection network based on information flow logic propagation. This network extracts local salient features and global expression information from small traffic signs to accurately detect their type and location. Our method compresses the acquired environment image into an image file of a specified size. This image is then fed into the information flow logic propagation network, which first extracts multiple layers of low-level semantic information from the image. The ERM (Edge Refinement Module) module extracts boundary information from this low-level semantic information. This boundary information supervises the graph convolutional flow to extract high-level graph semantic information within and between multiple feature layers. Classification prediction is first performed on the obtained high-level feature layer, resulting in a sparse representation of the target region. The multi-layer high-level information extracted by the graph convolutional flow is then split element-by-element into feature layers of equal size. The convolutional flow network then gradually fuses the expanded feature layers from left to right and from top to bottom within the sparse region. Finally, object detection is performed using the fused feature layers. Summary of the Invention
[0004] In view of this, an object of the present invention is to provide a lightweight small traffic sign recognition method for an on-board intelligent system.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A lightweight and small traffic sign recognition method for an on-board intelligent system includes the following steps:
[0007] S1: Acquire the image to be detected;
[0008] S2: network feature extraction;
[0009] S3: Edge information extraction optimization;
[0010] S4: Use graph convolution flow to interact with F1, F2, F3, and F4 feature information to enhance feature extraction;
[0011] S5: Extract F from left to right with convolution flow from top to bottom z1 ,F z2 ,F z3 Feature information, enhance feature extraction.
[0012] Optionally, the S1 specifically includes the following steps:
[0013] S11: Using a professional camera to capture an actual image of the scene, thereby obtaining a required original image;
[0014] S12: Perform compression processing to process the image into a specified size as the network input image.
[0015] Optionally, the step S2 specifically includes the following steps:
[0016] S21: Extract complete irregular features through deformable convolutional networks, while reducing the amount of data required for subsequent processing; improving network speed and detection accuracy;
[0017] S22: Repeat the feature extraction operation to obtain preliminary feature layers F1, F2, F3, and F4.
[0018] Optionally, the S3 specifically includes the following steps:
[0019] S31: Based on the shallow feature F1, the subsequent feature layers F2, F3, and F4 are supplemented with details; the image edge information is gradually optimized through the ERM image edge information extraction module;
[0020] S32: In the ERM module, feature layer F2 is upsampled to the same size as feature layer F1, and then edge information of feature layer F1 is extracted to obtain the approximate edge information of the image. This edge information is fused with information from feature layer F2 and then the convolution features are aggregated through global average pooling. Then, 1D convolution and Sigmoid function are used to obtain the corresponding channel attention weights to optimize the current edge information.
[0021] S33: Repeat the ERM operation to remove irrelevant information from the edge information using deep semantic information, retaining the edge features of the target of interest; and obtain specific edge information E1.
[0022] Optionally, the S4 specifically includes the following steps:
[0023] S41: Split the feature layer F1, aggregate the internal features of the image through the domain graph flow AGF module for the adjacent features in the feature layer, and supervise the feature aggregation with edge information E1; the AGF module uses the adjacent area F1 of the feature layer 11 and F 12 As input, feature information F is mapped through two 1×1 convolutions 1a ,F 1b ,F 2a ,F 2b ; Polymerization F 1b ,F 2b Feature information, after normalizing and aggregating information through the Softmax function, the adjacency feature F is obtained b ; F b Respectively with F 1a ,F 2a Perform graph convolution operation, the operation process is as follows Then obtain the new adjacency feature layer F 11 and F 12 ; Perform the same operation on other adjacent areas of feature layer F1 to aggregate internal features;
[0024]
[0025] Where W is the trainable weight;
[0026] S42: Repeat the adjacent area graph convolution operation of F1 for feature layers F2, F3, and F4 to obtain node aggregation features F x1 ,F x2 ,F x3 ,F x4 ;
[0027] S43: For F x1 The same-position features between feature layers are aggregated through the spatial graph flow SGF convolution module, and the edge information E1 is used to supervise the feature aggregation; the AGF module is based on the same-position area F1 and F2 of the feature layer. 11 and F 21 As input, feature information F is mapped through two 1×1 convolutions 1a ,F 1b ,F 2a ,F 2b ; Polymerization F 1b ,F 2b Feature information, after normalizing and aggregating information through the Softmax function, the adjacency feature F is obtained b ; F b Respectively with F 1a ,F 2a Perform graph convolution operation, the operation process is as follows Then obtain the new adjacency feature layer F 11and F 12 , and finally aggregate F through convolution operation 11 and F 12 Feature acquisition new feature layer F 11 ; Perform the same operation on other adjacent areas of feature layer F1 to aggregate internal features;
[0028] S44: Repeat the regional graph convolution operation of F1 in the same position area for feature layers F2, F3, and F4 to obtain node aggregation feature F y1 ,F y2 ,F y3 ;
[0029] S45: Feature layer F y1 ,F y2 ,F y3 The adjacent region features are linked through the AGF graph flow convolution to re-aggregate the feature layer to obtain a new feature layer F z1 ,F z2 ,F z3 .
[0030] Optionally, the S5 specifically includes the following steps:
[0031] S51: Through F z1 Perform simple classification processing to extract the approximate target area of interest to guide subsequent feature fusion;
[0032] S52: Image segmentation, current feature layer F z1 ,F z2 ,F z3 The image size ratio is 4:2:1, so the F z1 Transformed into four-layer features F k1 ,F k2 ,F k3 ,F k4 ,F z2 Converted into the second-level feature F q1 ,F q2 ; At this time F k1 ,F k2 ,F k3 ,F k4 ,F q1 ,F q2 ,F z3 The same size, recorded as
[0033] S53: Graph stream convolution IFC information stream fusion, the current feature layer performs feature interaction in pairs in the IFC1 module; in particular, for the last feature layer F z3 , which is consistent with the feature layer F q2 interact; among them, for F k1 ,Fk2 , first to F k1 ,F k2 Information is stacked in the channel direction, the number of channels is adjusted after convolution activation, and then up-sampling is performed. The feature map size at this time is recorded as The feature layer is F v1 ,F v2 ,F v3 ,F v4 ;
[0034] S54: Graph stream convolution IFC information stream fusion, the current feature layer performs feature interaction in pairs in the IFC module; v1 ,F v2 , first to F v1 ,F v2 Information is stacked in the channel direction, and the number of channels is adjusted after convolution activation; then the feature layer F is upsampled. k1 Then the information is stacked in the channel direction and finally up-sampled. The feature map size is
[0035] S55: Continue feature fusion. At this time, the size of the feature map is (H, W), and the feature map is sent to the detection head for feature detection.
[0036] Optionally, the S5 further includes S6 network training;
[0037] The network is trained as a whole based on the detection results. The overall network training is based on the final result, and regression prediction is performed to obtain the prediction box loss, confidence loss, classification loss and edge loss:
[0038] L c1 =μ1L box1 (b g ,b)+μ2L obj1 (o g )+μ3L cls1 (c g ,c)+μ4L edg1 (e g ,e)
[0039] Among them L c1 is the total loss function; u1, u2, u3, u4 are artificially set weight hyperparameters, L box1 , L obj1 , L cls1 , L edg1 They are intersection-over-union loss, prediction box regression loss, confidence loss, category loss, and margin loss.
[0040] The present invention has the following advantages: Traditional object detection methods generally acquire deep feature maps of image blocks for detection and subsequently filter them to obtain correct classification and location information. However, these methods fall short of the requirements for detecting small traffic signs due to their neglect of spatial information and inability to obtain effective small target features. Furthermore, the limited computing power of onboard intelligent chips makes some commonly used small target detection methods difficult to implement. To improve the accuracy of small target detection and meet the computing power limitations of onboard intelligent chips, we propose a small traffic sign detection method based on an information flow logic propagation network. This method extracts local salient features and global expression information of small traffic signs to accurately detect their type and location. Our method compresses the captured environmental image into an image file of a specified size. This image is then fed into the information flow logic propagation network, which first extracts multiple layers of low-level semantic information from the image. The ERM module then extracts boundary information from this low-level semantic information. This boundary information is then used to supervise the graph convolutional flow to extract high-level graph semantic information within and between multiple feature layers. Classification prediction is first performed on the obtained high-level feature layers, resulting in a sparse representation of the target region. The multi-layer high-level information extracted by the graph convolutional flow is then split element-by-element into feature layers of equal size. The convolutional flow network is then used to gradually fuse the expanded feature layers from left to right and from top to bottom within the sparse region. Finally, the fused feature layers are used for object detection.
[0041] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0043] Figure 1 Flowchart of the small target detection method based on information flow logic propagation network.
[0044] Figure 2 Information flow logic propagation network.
[0045] Figure 3 Edge information reinforcement network ERM network diagram.
[0046] Figure 4 Graph convolutional flow aggregation network diagram.
[0047] Figure 5 Convolutional stream fusion network diagram. DETAILED DESCRIPTION
[0048] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0049] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0050] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0051] like Figures 1 to 5 As shown, the specific implementation details of each part of the present invention are as follows:
[0052] 1. Use a professional camera to capture real-world images of real-world scenes to obtain the required original images. Compress the images and reduce them to a specified size to serve as network input images.
[0053] 2. Use a deformable convolutional network to initially extract complete irregular features while reducing the amount of data required for subsequent processing. This improves network speed and detection accuracy.
[0054] 3. Use the ERM image edge information extraction module, based on the shallow feature F1, and the subsequent feature layers F2, F3, and F4 as detailed supplements. The ERM image edge information extraction module gradually optimizes the image edge information.
[0055] 4. Edge information supervision graph convolutional flow network extracts high-level semantic information within and between layers. Its operation process is as follows
[0056]
[0057] 5. Target regions are sparsely represented using high-level semantic information. The multi-layered high-level information extracted by the graph convolutional flow is then split element-by-element into feature layers of equal size. The convolutional flow network then gradually fuses the expanded feature layers from left to right and top to bottom within the sparse region. Finally, target detection is performed using the fused feature layers.
[0058] 6. During the training process, the entire network is trained based on the detection results. The overall network training is mainly based on the fusion feature layer to perform regression prediction and obtain the prediction box loss, confidence loss, classification loss and edge loss.
[0059] L c1 =μ1L box1 (b g ,b)+μ2L obj1 (o g )+μ3L cls1 (c g ,c)+μ4L edg1 (e g ,e)
[0060] Among them L c1 is the total loss function. u1, u2, u3, u4 are artificially set weight hyperparameters, L box1 , L obj1 , L cls1 , L edg1 They are intersection-over-union loss, prediction box regression loss, confidence loss, category loss, and margin loss.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A lightweight, small traffic sign recognition method for an onboard intelligent system, characterized by: The method comprises the following steps: S1: Acquire the image to be detected; S2: Network feature extraction, specifically including the following steps: S21: Extract complete irregular features through deformable convolutional networks, while reducing the amount of data required for subsequent processing; improving network speed and detection accuracy; S22: Repeat the feature extraction operation to obtain preliminary feature layers F1, F2, F3, F4; S3: Edge information extraction optimization, specifically including the following steps: S31: Based on the shallow feature F1, the subsequent feature layers F2, F3, and F4 are supplemented with details; the image edge information is gradually optimized through the ERM image edge information extraction module; S32: In the ERM module, feature layer F2 is upsampled to the same size as feature layer F1, and the edge information of feature layer F1 is extracted to obtain the edge information of the image. This edge information is fused with the information of feature layer F2 and then the convolution features are aggregated through global average pooling. Then, the corresponding channel attention weights are obtained through 1D convolution and Sigmoid function to optimize the current edge information. S33: Repeat the ERM operation to remove irrelevant information from the edge information with deep semantic information, retaining the edge features of the target of interest; obtain specific edge information E1; S4: Use graph convolution flow to interact with F1, F2, F3, and F4 feature information to enhance feature extraction. This includes the following steps: S41: Split the feature layer F1, aggregate the internal features of the image through the domain graph flow AGF module for the adjacent features in the feature layer, and supervise the feature aggregation with edge information E1; the AGF module uses the adjacent area F1 of the feature layer 11 and F 12 As input, feature information F is mapped through two 1×1 convolutions 1a ,F 1b ,F 2a ,F 2b ; Polymerization F 1b ,F 2b Feature information, after normalizing and aggregating information through the Softmax function, the adjacency feature F is obtained b ; F b Respectively with F 1a ,F 2a Perform graph convolution operation, the operation process is as follows Then obtain the new adjacency feature layer F 11 and F 12 ; Perform the same operation on other adjacent areas of feature layer F1 to aggregate internal features; Where W is the trainable weight; S42: Repeat the adjacent area graph convolution operation of F1 for feature layers F2, F3, and F4 to obtain node aggregation features F x1 ,F x2 ,F x3 ,F x4 ; S43: For F x1 The same-position features between feature layers are aggregated through the spatial graph flow SGF convolution module, and the edge information E1 is used to supervise the feature aggregation; the AGF module is based on the same-position area F1 and F2 of the feature layer. 11 and F 21 As input, feature information F is mapped through two 1×1 convolutions 1a ,F 1b ,F 2a ,F 2b ; Polymerization F 1b ,F 2b Feature information, after normalizing and aggregating information through the Softmax function, the adjacency feature F is obtained b ; F b Respectively with F 1a ,F 2a Perform graph convolution operation, the operation process is as follows Then obtain the new adjacency feature layer F 11 and F 12 , and finally aggregate F through convolution operation 11 and F 12 Feature acquisition new feature layer F 11 ; Perform the same operation on other adjacent areas of feature layer F1 to aggregate internal features; S44: Repeat the regional graph convolution operation of F1 in the same position area for feature layers F2, F3, and F4 to obtain node aggregation feature F y1 ,F y2 ,F y3 ; S45: Feature layer F y1 ,F y2 ,F y3 The adjacent region features are linked through the AGF graph flow convolution to re-aggregate the feature layer to obtain a new feature layer F z1 ,F z2 ,F z3 ; S5: Extract F from left to right with convolution flow from top to bottom z1 ,F z2 ,F z3 Feature information, enhance feature extraction.
2. The lightweight small traffic sign recognition method for an on-board intelligent system according to claim 1 is characterized by: The S1 specifically includes the following steps: S11: Using a professional camera to capture an actual image of the scene, thereby obtaining a required original image; S12: Perform compression processing to process the image into a specified size as the network input image.
3. The lightweight small traffic sign recognition method for an on-board intelligent system according to claim 1 is characterized by: The S5 specifically includes the following steps: S51: Through F z1 Perform simple classification processing and extract the target area of interest to guide subsequent feature fusion; S52: Image segmentation, current feature layer F z1 ,F z2 ,F z3 The image size ratio is 4:2:1, so the F z1 Transformed into four-layer features F k1 ,F k2 ,F k3 ,F k4 ,F z2 Converted into the second-level feature F q1 ,F q2 ; At this time F k1 ,F k2 ,F k3 ,F k4 ,F q1 ,F q2 ,F z3 The same size, recorded as S53: Graph stream convolution IFC information stream fusion, the current feature layer performs feature interaction in pairs in the IFC1 module; in particular, for the last feature layer F z3 , which is consistent with the feature layer F q2 interact; among them, for F k1 ,F k2 , first to F k1 ,F k2 Information is stacked in the channel direction, the number of channels is adjusted after convolution activation, and then up-sampling is performed. The feature map size at this time is recorded as The feature layer is F v1 ,F v2 ,F v3 ,F v4 ; S54: Graph stream convolution IFC information stream fusion, the current feature layer performs feature interaction in pairs in the IFC module; v1 ,F v2 , first to F v1 ,F v2 Information is stacked in the channel direction, and the number of channels is adjusted after convolution activation; then the feature layer F is upsampled. k1 Then the information is stacked in the channel direction and finally up-sampled. The feature map size is S55: Continue feature fusion. At this time, the size of the feature map is (H, W), and the feature map is sent to the detection head for feature detection.
4. The lightweight small traffic sign recognition method for an on-board intelligent system according to claim 3 is characterized by: Said S5 also includes S6 network training; The network is trained as a whole based on the detection results. The overall network training is based on the final result, and regression prediction is performed to obtain the prediction box loss, confidence loss, classification loss and edge loss: L c1 =μ1L box1 (b g ,b)+μ2L obj1 (o g )+μ3L cls1 (c g ,c)+μ4L edg1 (e g ,e) Among them L c1 is the total loss function; u1, u2, u3, u4 are artificially set weight hyperparameters, L box1 , L obj1 , L cls1 , L edg1 They are intersection-over-union loss, prediction box regression loss, confidence loss, category loss, and margin loss.