A real-time traffic target detection method for low-light-intensity scenes
Patent Information
- Application Number
- CN202411042009.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-07-31
AI Technical Summary
[0004]问题一:现有的公开数据集中大多数样本来自正常光照环境,而低光照样本较少,数据分布不均衡
[0034]本发明使用RT-DETR作为基础检测模型,无需在检测前后进行Anchor设置与NMS处理,加快检测速度;嵌入的图像增强模块PENet能够提取更加有效的图片特征信息,提高检测性能;所替换的主干网络VMamba-tiny具备全局感受野与线性复杂度,不增加计算量的同时使模型获得全局关系建模能力;所使用的Local-FFN单元能够增强局部区域感知能力;为颈部的编码器模块引入的AFPN模块,提升多层级特征融合效果,进一步增强模型对不同尺度目标的检测能力。
Smart Images

Figure CN119007149B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning, computer vision, and object detection technology, and specifically to a real-time traffic object detection method for low-light scenarios. Background Technology
[0002] With the continuous development of the social economy and the gradual increase in residents' income, more and more residents are purchasing cars and using them as their main mode of transportation for daily travel. This has led to a rapid increase in car ownership and an increasingly complex urban traffic environment. Thanks to the rapid development of deep learning and computer vision technologies in recent years, autonomous driving technology has received increasing attention and is widely regarded as an important method for addressing complex urban traffic problems. A key foundational technology for realizing autonomous driving is traffic target detection technology.
[0003] Current object detection methods based on deep learning and computer vision technologies have achieved relatively mature development, giving rise to widely used object detection models such as YOLO and SSD. After training on publicly available traffic object detection datasets for a considerable period, the obtained detection models can usually achieve good detection accuracy on the validation set. However, the above-mentioned methods still have the following problems:
[0004] Problem 1: Most samples in existing public datasets come from normal lighting environments, while low-light samples are scarce, resulting in an imbalanced data distribution. This makes it difficult for the model to fully learn features under low-light conditions during training, leading to decreased detection accuracy in such conditions. Therefore, this traffic target detection method is not reliable in certain practical applications, such as at night, in inclement weather, or inside tunnels.
[0005] Question 2: Achieving high-precision object detection in low-light scenes remains a significant challenge. Images acquired in low-light environments suffer from a range of problems, including blurriness, low sharpness, and poor visibility, making it difficult to effectively detect image details and small targets. Traditional solutions involve simple image enhancement, such as increasing brightness, before feeding the image into the object detection model for training or inference. While this method can improve detection accuracy in low-light environments to some extent, it damages details in bright areas of the image and prevents end-to-end model training, hindering the model from learning more robust features.
[0006] To address the aforementioned challenges in current traffic target detection technology, improve the detection accuracy of models in low-light scenarios, and accelerate the implementation and widespread application of target detection solutions, the following steps are necessary: First, collect a sufficient number of low-light traffic images to establish the target detection dataset required for model training. Second, use more advanced target detection models and select appropriate improvement modules to further optimize the detection model algorithm for low-light scenarios.
[0007] DETR is a novel object detection algorithm that uses a Transformer architecture with global relation modeling capabilities to handle object detection tasks. Unlike traditional anchor-box-based object detection methods, DETR directly predicts bounding boxes and class labels from images, without the need for predefined anchor boxes or region proposal networks, achieving high efficiency and accuracy in object detection tasks.
[0008] In a nighttime vehicle detection method and apparatus (CN113269119A), the inventors disclose a nighttime vehicle detection method, which includes: acquiring an image to be detected; using a trained target detection model to detect the image to obtain a detected target; wherein the target detection model processes the image to be detected by: extracting features from the image to obtain image features; enhancing the image features to obtain enhanced features; inputting the enhanced features into an RPN network to generate candidate boxes; processing the candidate boxes through a ROIPooling layer to obtain a feature map of a fixed size; and performing regression and classification on the feature map to obtain the detected target.
[0009] Existing techniques for processing low-light images generally employ traditional image enhancement methods independent of the model. They typically lack image enhancement modules specifically designed for low-light scenes that are embedded within the model, which hinders end-to-end feature learning.
[0010] Commonly used object detection models, such as YOLO, require complex pre- and post-processing steps. Before model training, bounding boxes in the dataset need to be clustered to determine their distribution, and a set of pre-defined anchor boxes is then used as the basis for prediction. Inappropriate anchor box settings can lead to detection errors. Furthermore, anchor boxes may not be flexible enough when handling targets of different sizes and shapes. After the model outputs its predictions, non-maximum suppression (NMS) is typically applied. NMS helps obtain more accurate results by eliminating highly overlapping prediction boxes. However, NMS is computationally expensive and may reduce detection speed. The effectiveness of NMS is also affected by threshold settings; inappropriate threshold settings can lead to missed or false detections.
[0011] Current object detection models commonly use CNN or ViT structures as their backbone networks. CNN structures have the advantage of low computational complexity, but their drawback is the lack of a global receptive field, making it impossible to model global relationships within the detected image, resulting in lower detection accuracy. While ViT structures possess the ability to model global contextual relationships, their quadratic complexity leads to enormous computational costs, which cannot meet the requirements of lightweight deployment and real-time detection in practical applications.
[0012] In addition, the commonly used structures for the neck of object detection models are generally FPN, PANet, etc., and their multi-level feature fusion effect needs to be further improved. Summary of the Invention
[0013] To overcome the aforementioned technical problems, this invention provides a real-time traffic target detection method for low-light scenarios that achieves a better balance between detection performance and detection speed, resulting in stronger overall performance.
[0014] The present invention is achieved by at least one of the following technical solutions.
[0015] A real-time traffic target detection method for low-light scenes includes the following steps:
[0016] S1. Establish a low-light scene detection dataset;
[0017] S2. Construct a traffic target detection model for low-light scenarios based on RT-DETR improvement;
[0018] S3. Train the traffic target detection model using the detection dataset;
[0019] S4. Use the trained traffic target detection model to detect the target to be detected.
[0020] Furthermore, by obtaining publicly available traffic target detection datasets, image data with low-light scenes are collected, and the collected image data is formatted to form a low-light scene detection dataset.
[0021] Furthermore, the traffic target detection model for low-light scenes based on the RT-DETR improvement is based on the RT-DETR model, which integrates an image enhancement module; the backbone network for feature extraction is upgraded to VMamba-tiny, and the original FFN units in VMamba-tiny are replaced with Local-FFN units. At the same time, the cross-scale feature fusion module (CCFM) in the encoder is updated to the asymptotic feature pyramid network module (AFPN), and the FFN units in the encoder and decoder are also replaced with Local-FFN units.
[0022] Furthermore, the image enhancement module uses a pyramid enhancement network (PENet), which first processes the input image using Gaussian convolution and downsampling to obtain a Gaussian pyramid with four components; then processes the Gaussian pyramid using upsampling and Gaussian convolution to obtain a Laplacian pyramid with four components.
[0023] After obtaining the Laplacian pyramid, the image is enhanced using PENet. Specifically, the detail processing module (DPM) is used to enhance the components of each layer of the Laplacian pyramid, a low-frequency enhancement filter (LEF) is used to extract low-frequency information, and the outputs of the detail processing module (DPM) and the low-frequency enhancement filter (LEF) are stitched and fused together. Finally, the processed Laplacian components are restored to a high-resolution image.
[0024] Furthermore, the Detail Processing Module (DPM) includes context branches and edge branches;
[0025] Edge branch focuses on the edge information of the image. First, the Sobel operator is used to calculate the gradient approximation of the image brightness function in the vertical and horizontal directions to identify the edge information. Then, the convolution operation is used to extract the feature information of the image. Finally, the original image and the processed image are fused by residual connection to achieve information enhancement.
[0026] The context branch is responsible for capturing the remote dependencies of the image. The context branch includes two residual blocks, which expand and restore the number of channels, so that low-frequency information can pass through the residual blocks smoothly and be fused with high-frequency information to capture global information in the scene.
[0027] Furthermore, the Low Frequency Enhancement Filter (LEF) is used to capture low-frequency information of the image. First, the number of channels is expanded by a point convolutional layer, and the input components are divided into four parts along the channel dimension. The features of each part are applied to the corresponding adaptive average pooling layer to obtain low-frequency information at different cutoff frequencies, forming a low-pass filter. Then, bilinear interpolation is used to upsample it back to the original height and width. The four features are concatenated along the channel dimension to form a new feature map, which is then activated by ReLU. Finally, the number of channels is restored by a point convolutional layer.
[0028] Furthermore, the backbone network VMamba-tiny includes a block embedding module and four VMamba layers. The block embedding module scales the size of the input image according to a fixed ratio and expands the channel dimension to a fixed dimension to obtain embedded image features. The VMamba layers include a Visual State Space Block (VSS BLOCK) and a downsampling layer. The VSS BLOCK flattens the two-dimensional image features into one-dimensional sequence data through a four-way scanning strategy, then uses a selective state space model (S6) to model long-distance relationships, and finally uses a two-dimensional aggregation operation to re-aggregate the sequence data. The downsampling layer downsamples the image features to extract multi-scale features at different layers of the backbone network.
[0029] Furthermore, the Local-FFN unit has a 3×3 deep convolutional layer, which is capable of local feature perception of image pixels and their adjacent small areas.
[0030] Furthermore, the encoder module includes an Intra-Scale Feature Interaction Module (AIFI) and an Asymptotic Feature Pyramid Network Module (AFPN); the AIFI structure is a Transformer encoder layer, and the original FFN unit in the Transformer encoder layer is replaced with a Local-FFN unit;
[0031] The AFPN module performs convolution operations on the image features S3, S4, and S5 obtained by the feature extraction module to align the number of channels. It performs progressive feature fusion on S3 and S4 obtained from the backbone network and the fifth-level feature F5 processed by the AIFI module. First, S3 and S4 are weighted and fused, and then S3, S4, and F5 are weighted and fused, which helps to reduce the semantic gap between features at different levels.
[0032] Furthermore, the traffic target detection model training process is as follows: setting hyperparameters for the model training process, inputting training data into the traffic target detection model, performing image enhancement, feature extraction and fusion, outputting predicted values and calculating loss, optimizing the model, and saving the model weights after the model has been iteratively trained to convergence.
[0033] Compared with existing technologies, the beneficial effects of the present invention are as follows:
[0034] This invention uses RT-DETR as the basic detection model, eliminating the need for anchor settings and NMS processing before and after detection, thus accelerating the detection speed. The embedded image enhancement module PENet can extract more effective image feature information, improving detection performance. The replaced backbone network VMamba-tiny has a global receptive field and linear complexity, enabling the model to acquire global relation modeling capabilities without increasing computational cost. The Local-FFN unit used can enhance local region perception capabilities. The AFPN module introduced for the neck encoder module improves the multi-level feature fusion effect, further enhancing the model's ability to detect targets at different scales. Attached Figure Description
[0035] Figure 1 This is a flowchart of a real-time traffic target detection method for low-light scenarios according to an embodiment of the present invention;
[0036] Figure 2 The structural diagram of the traffic target detection model based on RT-DETR improvement for low-light scenarios provided by this invention;
[0037] Figure 3 This is a structural diagram of the Laplace pyramid component in an embodiment;
[0038] Figure 4 This is a structural diagram of the PENet image enhancement module in an embodiment.
[0039] Figure 5 This is a structural diagram of the Detail Processing Module (DPM) in the embodiment.
[0040] Figure 6 This is a structural diagram of a low-frequency enhancement filter (LEF) as an example.
[0041] Figure 7 The following is a structural diagram of the VMamba-tiny feature extraction backbone network for the example implementation;
[0042] Figure 8 This is a structural diagram of the VSS BLOCK example.
[0043] Figure 9 This is a structural diagram of the SS2D module in the embodiment;
[0044] Figure 10 This is a flowchart illustrating the specific process of the four-way scanning operation in the embodiment.
[0045] Figure 11 This is a structural diagram of the Local-FFN unit in the embodiment;
[0046] Figure 12 This is a structural diagram of the improved encoder module in the embodiment. Detailed Implementation
[0047] To clarify the technical content and advantages of the present invention, the specific embodiments of the present invention will be explained in detail below with reference to the accompanying drawings.
[0048] like Figure 1 As shown, the real-time traffic target detection method for low-light scenarios in this embodiment includes the following steps:
[0049] S1. Obtain a large traffic target detection dataset. Collect images and corresponding annotations for low-light scenes such as evening or night in the above dataset, and process the data: modify the collected images to a uniform resolution, convert the corresponding bounding box coordinates of the images after the resolution is modified, and uniformly modify the annotation format to the COCO data format to complete the establishment of the low-light scene traffic target detection dataset.
[0050] As one example, large traffic target detection datasets include BDD100k and KITTI.
[0051] S2. Design a traffic target detection model for low-light scenarios based on RT-DETR improvement, the structure of which is shown in the attached figure. Figure 2 As shown, the improvement process includes: embedding a PENet-based image enhancement module in the RT-DETR framework; replacing the feature extraction backbone network with VMamba-tiny and replacing the original FFN units inside VMamba-tiny with Local-FFN units; replacing the CCFM module in the encoder with the AFPN module; and similarly upgrading the original FFN units in the encoder and decoder modules to Local-FFN units.
[0052] PENet is used as the image enhancement module. Before entering PENet, components of different scales are first obtained through the Laplacian pyramid, as shown in Equation (1). A 5×5 convolution kernel is used to perform Gaussian convolution on the input, and the downsampling method is to discard pixels with even indices. After each Gaussian convolution and downsampling, the width and height of the output component are halved, while the number of channels remains unchanged, resulting in a four-layer Gaussian pyramid {G0,G1,G2,G3}. Next, the generated Gaussian pyramid is upsampled by doubling the image width and height, filling the newly added rows and columns with 0, multiplying the pixel values of the previous Gaussian convolution kernel by 4, and then performing Gaussian convolution. Since the downsampling operation of the Gaussian pyramid is irreversible, the lost information constitutes the components of the Laplacian pyramid, resulting in a four-layer Laplacian pyramid {L0,L1,L2,L3}. The specific formula is shown in Equation (2). The Laplacian pyramid is attached. Figure 3 As shown.
[0053] G i+1=Down(Gaussian(G i )) (1)
[0054] L i =G i -Gaussian(Up(G i+1 (2)
[0055] In the formula, G i Represents the i-th layer component of the Gaussian pyramid, Gaussian() represents the Gaussian convolution operation, Down() represents the downsampling operation, Up() represents the upsampling operation, and L... i This represents the i-th layer component of the Laplace pyramid.
[0056] The PENet structure is shown in the attached figure. Figure 4 As shown, PENet first takes the four layers of the Laplacian pyramid as input, then applies a detail processing module (DPM) to each layer to enhance detail features, and a low-frequency enhancement filter (LEF) to extract low-frequency information. The outputs of the two modules are then concatenated and fused using convolution to obtain the final output. Finally, the output layers are upsampled to restore a high-resolution image.
[0057] The Detail Processing Module (DPM) consists of context branches and edge branches, as shown in the attached diagram. Figure 5 As shown, the edge branch focuses on the edge information of the image. First, it uses the Sobel operator to calculate the gradient approximation of the image brightness function in the vertical and horizontal directions to identify edge information. Then, it uses convolution operations to extract image feature information. Finally, it fuses the original image and the processed image using residual connections to achieve information enhancement. The context branch is responsible for capturing the long-range dependencies of the image. First, it uses residual blocks to perform channel expansion on the input of this branch, obtaining the channel-expanded intermediate features. The intermediate features are then convolved and softmaxed to obtain a feature mask. The feature mask is multiplied point-by-point with the intermediate features, then processed through another convolutional layer, and residual connections are made with the intermediate features. Finally, another residual block is used for channel recovery to obtain the output. The context branch allows low-frequency information to pass smoothly through the residual block and fuse with high-frequency information, which is beneficial for capturing global information in the scene.
[0058] Low-frequency enhancement filters (LEFs) can capture low-frequency information from an image; their structure is shown in the attached figure. Figure 6As shown, the input components are first divided into four parts along the channel dimension by a point convolutional layer to expand the number of channels. Each part's features are then subjected to corresponding adaptive average pooling layers (1×1, 2×2, 3×3, and 6×6) to obtain low-frequency information at different cutoff frequencies, forming a low-pass filter. Bilinear interpolation is then used to upsample the features back to their original height and width. The four features are then concatenated along the channel dimension to form a new feature map, which is activated using ReLU. Finally, a point convolutional layer restores the number of channels.
[0059] The features processed by DPM and LEF are concatenated along the channel dimension, and then feature fusion is performed using a point convolutional layer. After processing the Laplacian components at different resolutions, the inverse operation of formula (2) is executed to recover the high-resolution image.
[0060] The feature extraction module uses VMamba-tiny as its backbone network, and its structure is shown in the attached figure. Figure 7 As shown, the VMamba-tiny network structure can be divided into a block embedding module and four VMamba layers. The training image is first input into the block embedding module. After two convolution operations, the image size is reduced to 1 / 2 and 1 / 4 of its original size, while the number of image channels is expanded to a fixed dimension to obtain embedded image features. The VMamba layers contain a downsampling module and several VSS blocks. The downsampling module performs a 1 / 2 scaling downsampling operation on the input tensor, while simultaneously expanding the number of channels to twice its original size. The network structure of VSSBLOCK is shown in the attached figure. Figure 8 As shown. First, the input feature tensor undergoes layer normalization. Then, the normalized tensor is input into the two-dimensional selective scan module (SS2D). Figure 9 The structure of the SS2D module is demonstrated. The SS2D module uses a Selective Scan Space State Sequential Model (S6) for long-distance relationship modeling. First, a four-way scan (2D Scan) operation is performed on the input tensor to transform the two-dimensional graphic data into one-dimensional sequence data. The specific process of the four-way scan operation is as follows: The two-dimensional graphic feature tensor is scanned alternately row-wise from the top-left pixel to the right, row-wise from the bottom-right pixel to the left, column-wise from the top-left pixel downwards, and column-wise from the bottom-right pixel upwards. This flattens the pixels of the feature map into a one-dimensional sequence according to the scan path, resulting in four four-way scan one-dimensional sequence feature data. The detailed process of the four-way scan operation is attached. Figure 10As shown, the four-way scanning strategy ensures that each feature element can integrate information from all directions to build a global receptive field, while avoiding a quadratic increase in computational complexity, thus guaranteeing the efficiency and accuracy of the model when processing high-resolution data.
[0061] The one-dimensional sequence data is then input into the S6 model for relational modeling. Finally, the output of the S6 model is processed by two-dimensional aggregation (2D Merge) to restore it to two-dimensional graphical data.
[0062] The output of the SS2D module is residually concatenated with the input of the VSS BLOCK to obtain intermediate features. These intermediate features are then normalized before being input into the Local-FFN unit. The structure of the Local-FFN unit is shown in the attached figure. Figure 11 As shown. The Local-FFN module first transforms the input two-dimensional feature matrix, with dimensions (H×W)×C, into a three-dimensional feature tensor of size C×H×W through a reshape operation. Then, it performs 1×1 convolution, batch normalization, and h-swish activation on the feature tensor. The h-swish activation function is:
[0063]
[0064] In the formula, x represents the feature vector activated by h-swish, and ReLU() represents the ReLU activation function.
[0065] Then, a 3×3 depthwise convolution is performed on the feature tensor, batch normalized, and h-swish activated, before being input into the SE attention module. The SE attention module adaptively learns the relationships between feature channels. Next, a 1×1 convolution is performed on the output of the SE attention module, followed by batch normalization. Finally, the 3D feature tensor of size C×H×W is reshaped into a 2D feature matrix of size (H×W)×C, residually concatenated with the input of the Local-FFN module, and then output.
[0066] Compared to the original FFN unit, the Local-FFN unit improves upon the original FFN unit's deficiency of performing linear transformations on each feature individually and lacking local information interaction. By introducing a 3×3 deep convolutional layer, it can perceive local features of image pixels and their adjacent small regions, enabling information exchange between adjacent pixels and thus gaining the ability to model local regional relationships.
[0067] The encoder module mainly consists of an Intra-Scale Feature Interaction Module (AIFI) and an Asymptotic Feature Pyramid Network Module (AFPN), as shown in the attached diagram. Figure 12As shown. First, a 1×1 convolution operation needs to be performed on the last three layers of image features S3, S4, and S5 obtained by the feature extraction module to align the number of channels. Since S5 contains the highest-level semantic features, processing S5 alone can not only maintain the performance of the model but also significantly reduce the amount of computation. Therefore, only the processed S5 is flattened and then fed into the AIFI structure. The AIFI structure is a Transformer encoder layer that replaces the original FFN units in the encoder layer with Local-FFN units. After S5 is processed by AIFI, the feature shape is restored and denoted as F5.
[0068] AFPN is used to perform cross-scale feature fusion of S3, S4, and F5. AFPN uses upsampling and downsampling to align features at different scales, fusing S3 and S4 first, and then fusing S3, S4, and F5. Features at different levels have different weights, and weighted fusion preserves information from key scale features. Fusing S3 and S4 reduces the semantic gap between them, and since S4 and F5 features are adjacent, it also reduces the semantic gap between S3 and F5 features. Therefore, AFPN progressively fuses image features at three scales (S3, S4, and F5), and the feature information at different scales becomes increasingly similar during the progressive fusion process, improving the ability of cross-scale feature fusion.
[0069] The multi-scale features obtained from the encoder module are flattened and concatenated. Based on the query selection scheme with minimal uncertainty, more features with accurate classification and precise location are provided. Noise is added to the ground truth values to obtain the content information of the bounding boxes and the position information of the anchor boxes. These two pieces of information, along with the concatenated multi-scale feature information, are fed into the Transformer decoder for processing to obtain the predicted coordinate and category information. The FFN unit in the Transformer decoder is upgraded to a Local-FFN unit.
[0070] S3. Iteratively train the traffic target detection model using a low-light detection dataset. First, configure the model training hyperparameters, including the input image size, the number of categories to be recognized, the initial learning rate, and the number of training epochs. Then, input the training data into the model for image enhancement, feature extraction, feature fusion, outputting predicted values, calculating the loss, and backpropagation to optimize the model.
[0071] The loss function for iterative optimization is:
[0072]
[0073] In the formula, The feature vector representing the output of the encoder module; Represents the predicted value. Represents the predicted bounding box. The predicted category is represented by y = {b, c}, where b represents the true bounding box and c represents the true category. This represents the difference between the localization and classification prediction distributions; L represents the total loss value, L box L represents the loss value of the bounding box. cls The loss value represents the category.
[0074] As one example, the backpropagation optimization algorithm used is Adam.
[0075] After the network model has been fully trained and converged, the model's weight file is saved, resulting in a fully trained traffic target detection model.
[0076] S4. Use the trained detection model to detect the target.
[0077] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A real-time traffic target detection method for low-light scenarios, characterized in that, Includes the following steps: S1. Establish a low-light scene detection dataset; S2. Construct a traffic target detection model for low-light scenes based on RT-DETR improvement; the traffic target detection model for low-light scenes based on RT-DETR improvement is based on the RT-DETR model, integrating an image enhancement module; the backbone network for feature extraction is upgraded to VMamba-tiny, and the original FFN units in VMamba-tiny are replaced with Local-FFN units. At the same time, the cross-scale feature fusion module CCFM in the encoder is updated to the asymptotic feature pyramid network module AFPN, and the FFN units in the encoder and decoder are also replaced with Local-FFN units; the image enhancement module uses the pyramid enhancement network PENet; S3. Train the traffic target detection model using the detection dataset; S4. Use the trained traffic target detection model to detect the target to be detected.
2. The real-time traffic target detection method for low-light scenarios according to claim 1, characterized in that, By obtaining publicly available traffic target detection datasets, image data of low-light scenes are collected, and the collected image data is formatted to form a low-light scene detection dataset.
3. The real-time traffic target detection method for low-light scenarios according to claim 1, characterized in that, First, the input image is processed using Gaussian convolution and downsampling to obtain a Gaussian pyramid with four components; then, the Gaussian pyramid is processed using upsampling and Gaussian convolution to obtain a Laplacian pyramid with four components. After obtaining the Laplacian pyramid, the image is enhanced using PENet. Specifically, the detail processing module (DPM) is used to enhance the components of each layer of the Laplacian pyramid, a low-frequency enhancement filter (LEF) is used to extract low-frequency information, and the outputs of the detail processing module (DPM) and the low-frequency enhancement filter (LEF) are stitched and fused together. Finally, the processed Laplacian components are restored to a high-resolution image.
4. The real-time traffic target detection method for low-light scenarios according to claim 3, characterized in that, The Detailed Processing Module (DPM) includes context branches and edge branches; Edge branch focuses on the edge information of the image. First, the Sobel operator is used to calculate the gradient approximation of the image brightness function in the vertical and horizontal directions to identify the edge information. Then, the convolution operation is used to extract the feature information of the image. Finally, the original image and the processed image are fused by residual connection to achieve information enhancement. The context branch is responsible for capturing the remote dependencies of the image. The context branch includes two residual blocks, which expand and restore the number of channels, so that low-frequency information can pass through the residual blocks smoothly and be fused with high-frequency information to capture global information in the scene.
5. The real-time traffic target detection method for low-light scenarios according to claim 3, characterized in that, The Low Frequency Enhancement Filter (LEF) is used to capture low-frequency information in an image. First, it expands the number of channels using a point convolutional layer and divides the input components into four parts along the channel dimension. The features of each part are applied to the corresponding adaptive average pooling layer to obtain low-frequency information at different cutoff frequencies, forming a low-pass filter. Then, bilinear interpolation is used to upsample it back to its original height and width. The four features are then concatenated along the channel dimension to form a new feature map, which is activated using ReLU. Finally, a point convolutional layer is used to restore the number of channels.
6. The real-time traffic target detection method for low-light scenarios according to claim 1, characterized in that, The backbone network VMamba-tiny includes a block embedding module and four VMamba layers. The block embedding module scales the input image size according to a fixed ratio and expands the channel dimension to a fixed dimension to obtain embedded image features. The VMamba layers include a visual state space block (VSS BLOCK) and a downsampling layer. The VSS BLOCK flattens the two-dimensional image features into one-dimensional sequence data through a four-way scanning strategy, then uses the selective state space model S6 to model long-distance relationships, and finally uses a two-dimensional aggregation operation to re-aggregate the sequence data. The downsampling layer downsamples the image features to extract multi-scale features at different layers of the backbone network.
7. The real-time traffic target detection method for low-light scenarios according to claim 1, characterized in that, The Local-FFN unit has a 3×3 deep convolutional layer, which is capable of local feature perception of image pixels and their adjacent small areas.
8. The real-time traffic target detection method for low-light scenarios according to claim 1, characterized in that, The encoder module includes an intra-scale feature interaction module AIFI and an asymptotic feature pyramid network module AFPN; the AIFI structure is a Transformer encoder layer, and the original FFN unit in the Transformer encoder layer is replaced with a Local-FFN unit; The AFPN module performs convolution operations on the image features S3, S4, and S5 obtained by the feature extraction module to align the number of channels. It performs progressive feature fusion on S3 and S4 obtained from the backbone network and the fifth-level feature F5 processed by the AIFI module. First, S3 and S4 are weighted and fused, and then S3, S4, and F5 are weighted and fused, which helps to reduce the semantic gap between features at different levels.
9. A real-time traffic target detection method for low-light scenarios according to claim 1, characterized in that, The traffic target detection model training process is as follows: set the hyperparameters for the model training process, input the training data into the traffic target detection model, perform image enhancement, feature extraction and fusion, output the predicted value and calculate the loss, optimize the model, and save the model weights after the model has been iteratively trained to convergence.
Citation Information
Patent Citations
Night vehicle detection method and device
CN113269119A
Container weak and small serial number target detection and identification method based on deep learning
CN117253154A
Artificial intelligence system and method serving electric power robot
WO2022188379A1