Multi-scale traffic signal light detection and recognition method, device, equipment and storage medium

By using the feature extraction model of the multi-scale fusion module and the cross-long sequence space attention module in the on-board computing unit, combined with the YOLO series of object detection models, the accuracy and real-time problems of traffic light detection in the on-board computing unit are solved, and efficient and accurate traffic light detection is achieved.

CN119027908BActive Publication Date: 2025-05-23AUTOMOTIVE DATA OF CHINA (TIANJIN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410874816.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2025-05-23
Estimated Expiration
2044-07-02

AI Technical Summary

Technical Problem

In the on-board computing unit, achieving accurate and real-time detection of traffic lights is challenging, especially in complex road environments, where traditional methods perform poorly, while deep learning models are difficult to meet real-time requirements when resource constraints are found.

Method used

A multi-scale traffic light detection and identification method is proposed. Through the multi-scale fusion module and cross-long sequence space attention module in the feature extraction model, multi-scale features are extracted and fused, and combined with the YOLO series object detection model, efficient detection of traffic lights is achieved.

Benefits of technology

This method greatly reduces the computational complexity of the model, ensures that it operates at low resource consumption on the on-board computing unit, maintains high accuracy and meets real-time requirements, and provides a traffic light detection solution that is both accurate and efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119027908B_ABST
    Figure CN119027908B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-scale traffic light detection and recognition method, device, equipment and storage medium, and relates to the field of target detection technology. The technical points of the present invention include: obtaining an image data set containing running traffic lights; preprocessing the image data set; inputting the preprocessed data set into a feature extraction model for feature extraction to obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-order spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features to obtain fusion features; the cross-long-order spatial attention module is used to process the fusion features to obtain features containing pixel relative position information; and the feature set is input into the target detection model for training and testing. The present invention realizes high-accuracy and real-time traffic light detection, greatly reduces the computational complexity of the model, and ensures that the model can run with low resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a multi-scale traffic signal light detection and recognition method, device, equipment and storage medium. Background Art

[0002] With the continuous advancement of science and technology and the continuous development of society, intelligent transportation systems and autonomous driving technologies have gradually entered real life, providing new solutions for road safety and traffic efficiency. In this context, visual perception technology, as a core component of intelligent transportation systems, plays a key role. Among them, traffic lights are important signals for road traffic, and their accurate recognition and real-time detection are crucial to achieving the goals of intelligent driving and traffic management.

[0003] However, achieving accurate detection of traffic lights is not a simple task. The state changes of traffic lights are affected by many factors such as weather, lighting, and traffic scenes, so it is very challenging to perform reliable traffic light detection in complex road environments. Traditional methods often perform poorly in these ever-changing situations, and the rise of deep learning technology has provided new possibilities for traffic light detection.

[0004] In recent years, object detection technology in deep learning has made significant progress, especially object detection methods based on convolutional neural networks (CNN). Among them, YOLO (You Only Look Once) [1] As an end-to-end real-time target detection algorithm, the YOLO series has attracted much attention due to its high efficiency and accuracy. Among them, YOLOv5s, as a lightweight version of the YOLO series, can achieve better detection performance while reducing the demand for computing resources, making it a powerful choice for traffic light detection under resource-constrained conditions.

[0005] However, in vehicle-mounted intelligent systems, due to the limitation of computing resources, deploying efficient target detection models on vehicle-mounted computing units is still a challenging task. Therefore, how to improve the detection accuracy of small targets such as traffic lights while effectively ensuring real-time performance still needs to be explored and studied. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a multi-scale traffic light detection and recognition method, device, equipment and storage medium.

[0007] According to one aspect of the present invention, a multi-scale traffic light detection and recognition method is proposed, the method comprising:

[0008] Acquire a data set, wherein the data set is images of operating traffic lights collected under different conditions;

[0009] Preprocessing the data set;

[0010] The preprocessed data set is input into the feature extraction model for feature extraction to obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-sequence spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features of the preprocessed data set to obtain fused features; the cross-long-sequence spatial attention module is used to process the fused features to obtain features containing pixel relative position information;

[0011] Input the feature set into the target detection model for training, and obtain the trained target detection model;

[0012] After the image to be detected is input into the feature extraction model for feature extraction, the extracted features are then input into the trained target detection model to obtain the detection results.

[0013] In one possible implementation, preprocessing the data set includes data enhancement, and the data enhancement includes: Gaussian blur, median filtering, changing saturation or contrast, random rotation, flipping, scaling, random cropping, splicing, and random erasing.

[0014] In one possible implementation, the multi-scale feature extraction and fusion of the preprocessed data set in the multi-scale fusion module includes: extracting features from images of multiple scales through corresponding pre-trained feature extractors to obtain features of different scales, wherein the images of the multiple scales include original images, images of the original images reduced by a fixed ratio, and images of the original images magnified; processing features of different scales separately to obtain features of the same size; and superimposing and fusing multiple features of the same size to obtain fused features.

[0015] In one possible implementation, the features of different scales are processed separately to obtain features of the same size, including: for an image whose original image is reduced by a fixed ratio, the corresponding feature processing process is: center cropping and then upsampling to the feature size corresponding to the original image; for an image whose original image is enlarged, the corresponding feature processing process is: expansion and then downsampling to the feature size corresponding to the original image.

[0016] In one possible implementation manner, the original image is enlarged by using bilinear interpolation.

[0017] In one possible implementation, the superposition and fusion of multiple features of the same size includes pixel superposition and channel superposition, wherein the pixel superposition is to add and average multiple feature values ​​of the same size, and the channel superposition is to splice multiple features of the same size in the channel dimension; the pixel superposition and channel superposition parts are superimposed again to obtain the final fused feature.

[0018] In one possible implementation, the cross-long-sequence spatial attention module includes two long-sequence spatial attention modules; the cross-long-sequence spatial attention module processes the fusion features to obtain features containing pixel relative position information, including:

[0019] The fused features are processed by the first long-order spatial attention module, and the output features contain the positional relationship between a certain pixel and its horizontal and vertical pixels;

[0020] The output features are transformed as follows: the order of the three dimensions of length, width, and channel is adjusted to channel, length, and width;

[0021] The features after dimensional order adjustment are then processed by the second long-order spatial attention module, and the output features contain the positional relationship between a certain pixel and all other pixels;

[0022] The dimension order of the output features is converted back to length, width, and channel.

[0023] In one possible implementation, the process of the long-order spatial attention module processing its input features includes:

[0024] The input features are passed through the pooling layer to obtain semantic features of different scales;

[0025] Perform 1x1 convolution operations on semantic features of different scales to obtain features with the same number of channels;

[0026] Upsample features with the same number of channels to the same scale, and fuse multiple features of the same scale;

[0027] The fused features are sequentially subjected to 1x1 convolution, ReLU activation function, 3x3 convolution, and Sigmoid activation function, and the generated features are subjected to Hadamard product operation with the fused features;

[0028] Separate the multiple features after the Hadamard product operation, perform matrix addition operation together with the input features, and output the features after the matrix addition operation.

[0029] In one possible implementation, the target detection model includes a YOLO series model.

[0030] According to another aspect of the present invention, a multi-scale traffic light detection and recognition device is provided, the system comprising:

[0031] A data acquisition unit configured to acquire a data set, wherein the data set is an image containing an operating traffic light collected under different conditions;

[0032] A preprocessing unit configured to preprocess the data set;

[0033] A feature extraction unit, configured to input the preprocessed data set into a feature extraction model to extract features and obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-sequence spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features of the preprocessed data set to obtain fused features; the cross-long-sequence spatial attention module is used to process the fused features to obtain features containing pixel relative position information;

[0034] A model training unit, configured to input the feature set into the target detection model for training, and obtain a trained target detection model;

[0035] The target detection unit is configured to input the image to be detected into the feature extraction model for feature extraction, and then input the extracted features into the trained target detection model to obtain the detection result.

[0036] According to another aspect of the present invention, an electronic device is provided, comprising: a memory, a processor and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the multi-scale traffic light detection and recognition method as described above.

[0037] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program; the computer program is executed by a processor to implement the multi-scale traffic light detection and recognition method as described above.

[0038] The embodiments of the present invention have the following technical effects:

[0039] The present invention proposes a multi-scale traffic light detection and recognition method, device, equipment and storage medium, aiming to achieve efficient and real-time traffic light detection on a vehicle-mounted computing unit by adding a multi-scale fusion module and a long-sequence spatial attention module.

[0040] The present invention significantly reduces the computational complexity of the model, ensuring that the model can run on the on-board computing unit with low resource consumption; the optimized model not only maintains high accuracy, but also meets the real-time requirements of the on-board computing unit. The present invention provides an accurate and efficient solution for traffic light detection in on-board intelligent systems, thereby promoting the development of intelligent transportation systems, improving the safety and smoothness of road traffic, and making important contributions to the safe driving and effective traffic management of intelligent connected vehicles in complex traffic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 It is a flow chart of a multi-scale traffic light detection and recognition method provided by an embodiment of the present invention.

[0043] Figure 2 This is an example diagram of an actual road traffic light data set in an embodiment of the present invention.

[0044] Figure 3 This is an example diagram of a data set for a traffic light road test unit in an embodiment of the present invention.

[0045] Figure 4 This is an example diagram after data preprocessing in an embodiment of the present invention.

[0046] Figure 5 It is a structural diagram of a multi-scale fusion module in an embodiment of the present invention.

[0047] Figure 6 Schematic diagram of the fusion operation of the multi-scale fusion module in an embodiment of the present invention.

[0048] Figure 7 It is a schematic diagram of the structure of the cross-long-sequence spatial attention module in an embodiment of the present invention.

[0049] Figure 8 It is a structural diagram of the long-order spatial attention module in an embodiment of the present invention.

[0050] Fig. 9 This is an example diagram of the visualization results of the test set in an embodiment of the present invention.

[0051] Fig.10 It is a structural schematic diagram of a multi-scale traffic signal light detection and recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0053] In order to solve the problem that the on-board computing unit in the prior art has difficulty in realizing accurate and real-time detection of traffic lights, the present invention proposes a multi-scale traffic light detection and recognition method, device, equipment and storage medium. Because traditional traffic light detection methods are often difficult to maintain high accuracy and real-time performance in complex environments (such as changing lighting conditions, different weather conditions and complex traffic backgrounds, etc.), and the existing deep learning models are often difficult to meet real-time requirements when deployed in vehicle systems due to computing resource limitations. The present invention proposes a feature extraction model, which is combined with a deep learning model to improve the detection ability of the target detection model for small targets such as traffic lights.

[0054] The first embodiment of the present invention proposes a multi-scale traffic signal light detection and recognition method, such as Figure 1 As shown, the method includes:

[0055] S110, acquiring a data set, wherein the data set is images of operating traffic lights collected under different conditions;

[0056] S120, preprocessing the data set;

[0057] S130, inputting the preprocessed data set into a feature extraction model to extract features and obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-sequence spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features of the preprocessed data set to obtain fused features; the cross-long-sequence spatial attention module is used to process the fused features to obtain features containing pixel relative position information;

[0058] S140, inputting the feature set into the target detection model for training, and obtaining a trained target detection model;

[0059] S150, input the image to be detected into the feature extraction model for feature extraction, and then input the extracted features into the trained target detection model to obtain the detection result.

[0060] The method starts at S110. In S110, a data set is acquired, wherein the data set is images including operating traffic lights collected under different conditions.

[0061] According to an embodiment of the present invention, the data set comes from two parts, wherein the first part comes from real road traffic light data, including 22654 pictures, such as Figure 2 As shown in the figure, these pictures are labeled with data; the second part is obtained by collecting and labeling data from the traffic light road test unit, with a total of 1810 pictures, such as Figure 3 As shown, the traffic light road test unit is used to perform visual recognition and communication tasks with intelligent connected vehicles.

[0062] The two parts are mixed together to form the data set used in this embodiment, of which the training set has 20792 images and the test set has 3672 images. These data are all images of running traffic lights collected under different conditions, including various viewing angles, various scenes, various time periods, various lighting conditions, and various weather conditions to simulate the real road environment.

[0063] Then, S120 is executed, in which the data set is preprocessed.

[0064] According to an embodiment of the present invention, preprocessing includes: data enhancement of training data, various disturbances of image data, including Gaussian blur, median filtering, changing saturation contrast, random rotation, flipping and scaling, random cropping and splicing, random erasing, etc., for example Figure 4 As shown in the figure, we can see that the training images have undergone a random data augmentation process. These data augmentation operations are used to increase the generalization ability of the model.

[0065] Then execute S130, in which the preprocessed data set is input into the feature extraction model for feature extraction to obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-order spatial attention module, the multi-scale fusion module is used to perform multi-scale feature extraction and fusion on the preprocessed data set to obtain fusion features; the cross-long-order spatial attention module is used to process the fusion features to obtain features containing pixel relative position information.

[0066] The process of extracting and fusing multi-scale features of the preprocessed data set in the multi-scale fusion module includes:

[0067] S1310, extracting features from images of multiple scales using corresponding pre-trained feature extractors to obtain features of different scales, wherein the images of multiple scales include the original image, the image of the original image reduced by a fixed ratio, and the image of the original image enlarged;

[0068] S1320. Process features of different scales separately to obtain features of the same size. Specifically, for the image obtained by reducing the original image by a fixed ratio, the corresponding feature processing process is: after central cropping, upsample to the feature size corresponding to the original image; for the image obtained by enlarging the original image, the corresponding feature processing process is: after expansion, downsample to the feature size corresponding to the original image. Optionally, bilinear interpolation can be used to enlarge the original image.

[0069] S1330. Stack and fuse multiple features of the same size to obtain a fused feature. Specifically, the stack and fusion include pixel stacking and channel stacking. The pixel stacking is to add the feature values of multiple features of the same size and then take the average. The channel stacking is to splice multiple features of the same size in the channel dimension. Stack the parts of pixel stacking and channel stacking again to obtain the final fused feature.

[0070] According to the embodiments of the present invention, based on the fact that traffic signal targets appear in images at different scales, in order to improve the detection effect of the network, a novel multi-scale fusion module (MSF) is proposed in this embodiment. As Figure 5 shown, this module is mainly used to fuse multiple scale feature maps to obtain the model's perception ability for different scale information.

[0071] Specifically, as Figure 5 shown, it can be seen that there are three different feature extractors, which are represented by three colors respectively. The inputs received by these three different feature extractors are different. It can be seen that the original picture is reduced and bilinearly interpolated to obtain X1 and X3 respectively. These two inputs are larger and smaller than the original picture scale, and X2 is the same as the original picture. Then, these three scale data are input into the corresponding three types of feature extractors. It should be noted that these three feature extractors are pre-trained. The specific pre-training is to perform corresponding transformations on the original picture and the corresponding labels respectively, and train the three feature extractors respectively, so that the three feature extractors have the feature extraction ability corresponding to the scale. After that, the pre-trained model is connected to the Figure 5 framework. Back to the model architecture, then, the feature maps output by the three feature extractors need to be processed separately. Among them, the feature map obtained by the blue feature extractor 1 needs to be centrally cropped first and then upsampled to the feature size output by the feature extractor 2. Similarly, for the feature map output by the yellow feature extractor, after bilinear interpolation expansion, it is then downsampled to the feature size output by the feature extractor 2. The three feature maps are then fused with multi-scale features through a feature stacking operation. The fused feature is then passed through a fused feature extractor to obtain the final feature map. This fused feature extractor is trained together with the entire model.

[0072] The process of fusing the above three feature maps with multi-scale features through feature superposition operation is as follows: Figure 6 As shown in the figure. Feature fusion consists of two parts, including pixel superposition and channel superposition. Pixel superposition is equivalent to superposition operation between corresponding pixels of each feature map, specifically, adding three eigenvalues ​​and then taking the average; channel superposition is to splice three feature maps in the dimension of the channel, and the total feature map needs to be superimposed with the pixel superposition part again (that is, the function implemented by the fusion feature extractor). The superposition here is also the superposition on the channel, so as to obtain the final feature map. This superposition operation fusion method can effectively integrate the information of the feature maps of the three scales, and at the same time retain the information of each channel. The fused feature map obtained by the above method can simultaneously receive multi-scale traffic light information, effectively improving the generalization of the network.

[0073] It should be noted that there may be more than three feature extractors, for example, the original image may be reduced in different proportions, or the original image may be enlarged in different proportions; in addition, other methods may be used to enlarge and expand the original image.

[0074] The cross-long-order spatial attention module includes two long-order spatial attention modules; the process of processing the fusion features in the cross-long-order spatial attention module to obtain features containing pixel relative position information includes:

[0075] S1340, the fused features are processed by the first long-order spatial attention module, and the output features include the positional relationship between a certain pixel and its horizontal and vertical pixels;

[0076] S1350, converting the output features as follows: adjusting the order of the three dimensions of length, width, and channel to channel, length, and width;

[0077] S1360, processing the features after the dimensional order is adjusted by a second long-order spatial attention module, and the output features include the positional relationship between a certain pixel and all other pixels;

[0078] S1370. Convert the dimensional order of the output features into length, width, and channel again.

[0079] The process of the long-order spatial attention module processing its input features includes: the input features are passed through a pooling layer to obtain semantic features of different scales; 1x1 convolution operations are performed on the semantic features of different scales to obtain features with the same number of channels; features with the same number of channels are upsampled to the same scale, and multiple features of the same scale are fused; the fused features are sequentially passed through 1x1 convolution, ReLU activation function, 3x3 convolution and Sigmoid activation function, and the generated features are subjected to Hadamard product operation with the fused features; multiple features subjected to the Hadamard product operation are separated and matrix addition operation is performed together with the input features; and the features after the matrix addition operation are output.

[0080] According to an embodiment of the present invention, the current attention mechanism is very effective in deep learning models. The attention mechanism is used to simulate the attention allocation process of humans when processing information. This technology was originally introduced in natural language processing tasks, but was later widely used in computer vision, speech recognition and other fields. The basic idea of ​​the attention mechanism is to dynamically adjust the attention of the model according to the importance of different parts or features of the input data in order to better complete the task. In the traffic light detection task, the traffic light part only occupies a very small part of the entire image. Although image cropping can be used to increase the proportion of the detection target in the entire image, a mechanism is still needed to improve the accuracy of the network, so consider using the attention mechanism to improve the performance of the model. Existing attention mechanisms mainly include channel attention and spatial attention, and there are also mixed channel attention methods, but these methods do not adjust the distribution ratio of attention in different parts of the model very well.

[0081] This embodiment proposes a cross-long-order spatial attention mechanism, the core idea of ​​which is to allow the model to dynamically adjust the attention weights so as to allocate different attention between different time steps or inputs, while allowing each pixel to gain the ability to perceive and establish connections with the entire feature map. This helps the model focus on processing task-related information, thereby improving the performance of the model.

[0082] like Figure 7As shown in the figure, there are two long-order spatial attention modules, where the feature map first passes through the first long-order spatial attention module, and then the output feature map contains the positional relationship between the horizontal and vertical pixels for each pixel. Then the feature map is transformed. The transformation here specifically refers to the sequential adjustment between the length, width, and channel dimensions of the feature map, from length, width, and channel to channel, length, and width. At this time, it is input into the second long-order spatial attention module, and the output feature map contains the positional relationship between a certain pixel and all other pixels. The output is then transformed into length, width, and channel. Through such operations, the model can successfully capture the relationship between channels and channels, and between channels and feature map slices. In addition, after two long-order spatial attention modules, each pixel in the feature map can successfully capture the information and relationship of other pixels in the entire image space, which can effectively improve the model's perception ability.

[0083] The specific operation mechanism of each long-order spatial attention module is as follows Figure 8 As shown in the figure, the feature map first passes through the long-order spatial pooling layer to obtain semantic features of three different scales; then the feature maps of the three different scales are respectively subjected to 1x1 convolution operations to obtain the same number of channels; then they are upsampled to the same scale using bilinear interpolation; then the three features are fused using a feature fusion operation, wherein the feature fusion operation is performed using the feature fusion operation described in the multi-scale fusion module, effectively utilizing channel fusion and pixel fusion; then the feature map is sequentially subjected to 1x1 convolution, ReLU activation, 3x3 convolution, and Sigmoid activation layers, and the generated feature map is subjected to a Hadamard product operation with the previous feature map; then the three features are separated and compared with the initially input feature map. Figure 1 The matrix addition operation is performed at the same time to obtain the final feature result map. The long-order spatial attention module can adaptively learn parameters during the training process, adjust the information of the three parts of the generated feature map, and then effectively fuse them.

[0084] Then, S140 is executed. In S140, the feature set is input into the target detection model for training to obtain a trained target detection model.

[0085] According to an embodiment of the present invention, the target detection model includes a YOLO series model. For example, a YOLOv5s model is used. YOLOv5 mainly includes four versions, namely YOLOv5s, YOLOv5m, YOLOv5l and YOLOv5x. Among them, YOLOv5s has the smallest number of parameters and can also maintain relatively high accuracy. Therefore, the present invention uses YOLOv5s as the backbone network, but if the YOLO network is directly used for the traffic light detection task, it will not get a particularly good effect. There are three main reasons: 1) The traffic light target is small and is mostly in the upper half of the image, and the proportion of the entire image is also very small; 2) When the front-view camera of an intelligent networked car is used to perceive the traffic light, the camera installation position also has a very large impact on the traffic light detection effect. At the same time, in actual driving conditions, weather factors will also affect the detection effect; 3) The computing power on the actual car is limited, and an algorithm that can achieve high precision while requiring very little computing power is required.

[0086] Based on the above three reasons, this embodiment proposes the above-mentioned feature extraction model, in which the multi-scale fusion module can further improve the performance of the model in the multi-scale target detection task: because the scale of the traffic light target in the collected image is random, and the image is preprocessed at the same time, so that the model can focus on identifying and detecting the traffic light part; the cross-long-order spatial attention module can make the network focus more on the traffic light part in the image rather than the entire image, improving the accuracy of the model while improving the detection efficiency of the model. At the same time, the number of channels and scale of the feature extraction model are modified to greatly increase its efficiency, and various data enhancement methods are used to improve the training effect, so as to achieve efficient real-time detection of traffic lights deployed on real vehicles.

[0087] It should be noted that the feature extraction model including the multi-scale fusion module and the cross-long-order spatial attention module can also be added to the feature extraction part of the YOLOv5s model to improve the feature extraction capability of the YOLOv5s model.

[0088] Then, S150 is executed. In S150, the image to be detected is input into the feature extraction model for feature extraction, and then the extracted features are input into the trained target detection model to obtain the detection result.

[0089] The technical effect of the present invention is further verified through experiments.

[0090] The experiment was run on an NVIDIA TITAN RTX graphics card with 32G memory. During the experiment, in order to obtain better experimental results, the batch size was set to 16 during training, 300 rounds of iterative training were performed, the SGD optimizer was used, the initial learning rate was 0.01, and the momentum was set to 0.937.

[0091] The metric parameters used in the experiment are precision, as shown in formula 1; recall, as shown in formula 2, and mAP@0.5 when the intersection over union ratio IoU=0.5, as shown in formula 3, where n represents the number of categories. It represents the area under the precision-recall curve of the i-th category, and IoU represents the ratio between the intersection of two bounding boxes (or regions) and their union. Using the average precision of all categories can comprehensively consider the characteristics of both precision and recall.

[0092]

[0093]

[0094]

[0095] Among them, TP represents the number of both predicted and actual positive examples; FP represents the number of predicted positive examples and actual negative examples; FN represents the number of predicted negative examples and actual positive examples.

[0096] In order to verify the effectiveness of the present invention, the present invention is compared with CenterNet [2] , SSD [3] ,DETR [4] , Yolov5s and other mainstream target detection methods were compared. The experimental results are shown in Table 1.

[0097] Table 1 Comparative experimental results

[0098]

[0099] As can be seen in Table 1, the present invention has significant improvements in accuracy, recall rate and mAP@0.5 compared to existing target detection methods. Compared with Yolov5s, the present invention has improved accuracy by 3.5%, recall rate by 3.4%, and mAP@0.5 by 2.7%. The present invention incorporates information of multiple scales by adding a multi-scale fusion module and a cross-long-order spatial attention module. Combining the attention mechanism with multi-channel and different receptive field information can effectively improve the effect of the model.

[0100] At the same time, in order to enable the model to be deployed on the vehicle side to realize real-time traffic light detection without occupying too many computing resources, the present invention optimizes the algorithm. First, the scale of the input image is reduced to 320x320 when deployed on the actual vehicle, and the number of channels of the network model is reduced by half. Finally, Nvidia's official framework tensorRT is used to accelerate the model. The optimized network model parameters are only 1.7M. Compared with the algorithm before optimization, the detection time of a single image is reduced from 14ms to 8ms. Finally, the algorithm is deployed on the vehicle side. Combined with the team's entire set of autonomous driving algorithms, it can perfectly realize the function of stopping when the vehicle recognizes red and yellow lights, and driving again when the light is green, enriching the visual perception of the whole vehicle.

[0101] Experiments were conducted on 3672 test set images, and the recognition effects of the total and red, green and yellow colors were calculated. As shown in Table 2, the total recognition accuracy can reach 82.4%, the recall rate can reach 82.9%, and the average precision of all categories can reach 85.1%. The trained network can achieve very good results for traffic light targets and traffic light road test units on real roads. At the same time, the model performs well under various weather and lighting conditions, verifying its robustness and reliability.

[0102] Table 2 Classification and recognition results

[0103]

[0104] The loss value of the model in the present invention has been steadily decreasing during the training process, and the accuracy, recall rate and average precision rate have been steadily increasing. The recognition effect of the model on the test set is shown in the figure below. Fig. 9 As shown, it can be seen that the model can recognize traffic lights well, and due to the extensiveness of the experimental dataset, the model can accurately detect and recognize traffic lights in complex scenes.

[0105] Compared with the prior art, the present invention has the following beneficial effects:

[0106] 1) Improved accuracy: By introducing a multi-scale fusion module and a cross-long-order spatial attention module before or in the YOLOv5s network structure, traffic lights in different scales and complex backgrounds can be detected more accurately, effectively improving the detection accuracy.

[0107] 2) Enhanced real-time performance: The optimized model greatly reduces the computational complexity by reducing the number of channels and scale of the network, ensuring real-time detection of traffic lights even on vehicle computing units with limited computing resources, meeting the high real-time requirements of intelligent driving systems.

[0108] 3) Resource utilization optimization: By optimizing and improving the model structure, the dependence on vehicle computing resources is reduced, so that the present invention can operate efficiently even in a resource-constrained environment, thereby improving the resource utilization of the vehicle system.

[0109] 4) Wide range of applications: The present invention is not only suitable for traffic light detection in intelligent connected vehicles, but can also be widely used in other aspects of intelligent transportation systems, such as traffic monitoring, violation detection, etc., and has broad application prospects.

[0110] In summary, the present invention effectively solves the problem of accurate and real-time detection of traffic lights in vehicle-mounted intelligent systems, and provides strong technical support for the development of intelligent transportation systems and road traffic safety.

[0111] Another embodiment of the present invention provides a multi-scale traffic signal light detection and recognition device, such as Fig.10 As shown, the system includes:

[0112] A data acquisition unit 1010 is configured to acquire a data set, wherein the data set is an image containing operating traffic lights collected under different conditions;

[0113] A preprocessing unit 1020, configured to preprocess the data set;

[0114] The feature extraction unit 1030 is configured to input the preprocessed data set into the feature extraction model to extract features and obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-sequence spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features of the preprocessed data set to obtain fused features; the cross-long-sequence spatial attention module is used to process the fused features to obtain features containing pixel relative position information;

[0115] A model training unit 1040 is configured to input the feature set into the target detection model for training, and obtain a trained target detection model;

[0116] The target detection unit 1050 is configured to input the image to be detected into the feature extraction model for feature extraction, and then input the extracted features into the trained target detection model to obtain the detection result.

[0117] For the undetailed parts of the multi-scale traffic signal light detection and recognition device according to the embodiment of the present invention, please refer to the above detailed description of the method embodiment.

[0118] The method of the present invention can be executed in an electronic device. The electronic device can be any device with storage and computing capabilities, which can be implemented as a server, a workstation, etc., or as a personal computer such as a desktop computer or a notebook computer, or as a terminal device such as a mobile phone, a tablet computer, a smart wearable device, an Internet of Things device, but is not limited thereto.

[0119] The electronic device may include: a processor, a memory, an input / output interface, a communication interface and a bus. The processor, the memory, the input / output interface and the communication interface are connected to each other through the bus in the electronic device. The processor may be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided in the embodiments of this specification. The memory may be implemented in the form of ROM, RAM, static storage devices, dynamic storage devices, etc. The memory may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory and called and executed by the processor. The input / output interface is used to connect the input / output module to realize information input and output. The input / output / module may be configured in the electronic device as a component, or it may be externally connected to the electronic device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc. The communication interface is used to connect the communication module to realize the communication interaction between the electronic device and other devices. The communication module may realize communication by wire or by wireless. A bus comprises a pathway that transfers information between the various components of an electronic device.

[0120] The embodiment of the present invention also provides a non-transitory readable storage medium, which stores instructions, and the instructions are used to enable the electronic device to perform a method according to an embodiment of the present invention. The readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be a computer-readable instruction, a data structure, a module of a program, or other data. Examples of readable storage media include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage, etc.

[0121] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

[0123] The documents cited in the present invention are as follows:

[0124] [1]Redmon J, Divvala S, Girshick R, et al. You only look once:Unified, real-time object detection[C] / / Proceedings of the IEEE conference oncomputer vision and pattern recognition. 2016: 779-788.

[0125] [2]Zhou X, Wang D, Krähenbühl P. Objects as points[J]. arXiv preprintarXiv:1904.07850, 2019.

[0126] [3]Liu W, Anguelov D, Erhan D, et al. Ssd: Single shot multiboxdetector[C] / / Computer Vision–ECCV 2016: 14th European Conference, Amsterdam,The Netherlands, October 11–14, 2016, Proceedings, Part I 14. SpringerInternational Publishing, 2016: 21-37.

[0127] [4]Carion N, Massa F, Synnaeve G, et al. End-to-end object detectionwith transformers[C] / / European conference on computer vision. Cham: SpringerInternational Publishing, 2020: 213-229。

Claims

1. A multi-scale traffic light detection and recognition method, characterized in that: include: Acquire a data set, wherein the data set is images of operating traffic lights collected under different conditions; Preprocessing the data set; The preprocessed data set is input into the feature extraction model for feature extraction to obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-sequence spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features of the preprocessed data set to obtain fused features; the cross-long-sequence spatial attention module is used to process the fused features to obtain features containing pixel relative position information; Input the feature set into the target detection model for training, and obtain the trained target detection model; After the image to be detected is input into the feature extraction model for feature extraction, the extracted features are then input into the trained target detection model to obtain the detection results; The cross-long-sequence spatial attention module includes two long-sequence spatial attention modules; the cross-long-sequence spatial attention module processes the fusion features to obtain features containing pixel relative position information, including: The fused features are processed by the first long-order spatial attention module, and the output features contain the positional relationship between a certain pixel and its horizontal and vertical pixels; The output features are transformed as follows: the order of the three dimensions of length, width, and channel is adjusted to channel, length, and width; The features after dimensional order adjustment are processed by the second long-order spatial attention module, and the output features contain the positional relationship between a certain pixel and all other pixels; The dimension order of the output features is converted back to length, width, and channel.

2. The multi-scale traffic light detection and recognition method according to claim 1, characterized in that: Preprocessing the data set includes data enhancement, and the data enhancement includes: Gaussian blur, median filtering, changing saturation or contrast, random rotation, flipping, scaling, random cropping, splicing, and random erasing; the target detection model includes a YOLO series model.

3. The multi-scale traffic light detection and recognition method according to claim 1, characterized in that: The multi-scale fusion module extracts and fuses multi-scale features of the preprocessed data set, including: extracting features of images of multiple scales through corresponding pre-trained feature extractors to obtain features of different scales, wherein the images of multiple scales include original images, images of the original images reduced by a fixed ratio, and images of the original images magnified; processing features of different scales respectively to obtain features of the same size; and superimposing and fusing multiple features of the same size to obtain fused features.

4. The multi-scale traffic signal light detection and recognition method according to claim 3, characterized in that: The features of different scales are processed separately to obtain features of the same size, including: for an image whose original image is reduced by a fixed ratio, the corresponding feature processing process is: up-sampling to the feature size corresponding to the original image after center cropping; for an image whose original image is enlarged, the corresponding feature processing process is: down-sampling to the feature size corresponding to the original image after expansion; wherein the original image is enlarged by using bilinear interpolation.

5. The multi-scale traffic light detection and recognition method according to claim 4, characterized in that: The superposition and fusion of multiple features of the same size includes pixel superposition and channel superposition. The pixel superposition is to add and average multiple feature values ​​of the same size, and the channel superposition is to splice multiple features of the same size in the channel dimension; the pixel superposition and channel superposition parts are superimposed again to obtain the final fusion feature.

6. The multi-scale traffic light detection and recognition method according to claim 1, characterized in that: The process of the long-order spatial attention module processing its input features includes: The input features are passed through the pooling layer to obtain semantic features of different scales; Perform 1x1 convolution operations on semantic features of different scales to obtain features with the same number of channels; Upsample features with the same number of channels to the same scale, and fuse multiple features of the same scale; The fused features are sequentially subjected to 1x1 convolution, ReLU activation function, 3x3 convolution, and Sigmoid activation function, and the generated features are subjected to Hadamard product operation with the fused features; Separate the multiple features after the Hadamard product operation, perform matrix addition operation together with the input features, and output the features after the matrix addition operation.

7. A multi-scale traffic light detection and recognition device, characterized in that: The method for detecting and identifying multi-scale traffic lights according to any one of claims 1 to 6 comprises: A data acquisition unit configured to acquire a data set, wherein the data set is an image containing an operating traffic light collected under different conditions; A preprocessing unit configured to preprocess the data set; A feature extraction unit, configured to input the preprocessed data set into a feature extraction model to extract features and obtain a feature set; wherein the feature extraction model includes a multi-scale fusion module and a cross-long-sequence spatial attention module, the multi-scale fusion module is used to extract and fuse multi-scale features of the preprocessed data set to obtain fused features; the cross-long-sequence spatial attention module is used to process the fused features to obtain features containing pixel relative position information; A model training unit, configured to input the feature set into the target detection model for training, and obtain a trained target detection model; The target detection unit is configured to input the image to be detected into the feature extraction model for feature extraction, and then input the extracted features into the trained target detection model to obtain the detection result.

8. An electronic device, characterized in that: include: A memory, a processor and a computer program; wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the multi-scale traffic light detection and recognition method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The storage medium stores a computer program; the computer program is executed by a processor to implement the multi-scale traffic light detection and recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Target detection method and equipment based on attention mechanism and multi-scale feature fusion, and storage medium

    CN112686304A

  • Target detection method based on full convolutional network

    CN116051899A