Low-illumination pedestrian detection method based on anchor-free lightweight visual Transform
Through the detection method based on anchorless lightweight visual Transformer, combined with LYT-Net image enhancement and mixed data enhancement, the missed detection and accuracy reduction of pedestrian detection under low light conditions is solved, and high accuracy and high robustness detection in night environments is achieved, suitable for autonomous driving and intelligent traffic.
Patent Information
- Application Number
- CN202510439038.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-01
AI Technical Summary
Under low-light conditions, existing pedestrian detection methods have problems of missed detection and reduced accuracy, especially in night and low-light scenarios. Traditional methods are difficult to meet the needs of autonomous driving in real-time and robustness.
The detection method based on anchorless lightweight visual Transformer is adopted, combined with LYT-Net image enhancement technology and hybrid data enhancement, and the architecture combined with lightweight MobileViT-v3 network and convolutional neural network is used to introduce the TSA attention mechanism and reparameterized convolution, and the loss function is optimized to Wise-IOUv3 loss, improving the detection performance of the model under low light conditions.
It significantly improves the pedestrian detection accuracy and recall rate under low light conditions at night, enhances the robustness and detection speed of the model, and is suitable for the fields of autonomous driving and intelligent transportation.
Smart Images

Figure CN120411896A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection in artificial intelligence and computer vision technologies, and particularly relates to a low-light pedestrian detection method based on an anchor-free lightweight vision Transformer. Background Art
[0002] Pedestrian detection is not only a key means to ensure road traffic safety, but also plays an important role in intelligent monitoring systems, especially in avoiding pedestrian traffic accidents. The main task of pedestrian detection is to identify and locate pedestrian targets in images or videos, which is particularly important for realizing the perception function of autonomous vehicles. In a vision-based pedestrian detection system, accurately and real-time detecting pedestrian targets in the road environment not only helps improve driving safety, but also provides key support for the further implementation of unmanned driving technology.
[0003] Traditional pedestrian detection methods mainly rely on manually designed features, such as HOG (Histogram of Oriented Gradients) and LBP (Local Binary Patterns). These methods achieve object detection by combining a sliding window mechanism and a classifier, such as an SVM. Such methods have certain effects in simple backgrounds, but are vulnerable to background interference, target occlusion, and illumination changes in complex scenes and perform poorly; subsequently, detection methods based on ensemble learning, such as AdaBoost and random forests, have been widely applied to pedestrian detection tasks. These methods form a strong classifier by combining multiple weak classifiers, improving the detection performance, but still have deficiencies in detection speed.
[0004] With the rise of deep learning, pedestrian detection methods based on the convolutional neural network (CNN) have gradually become mainstream. Two-stage detection methods such as Faster R-CNN provide high-precision detection performance. By generating candidate regions and classifying and regressing them, they solve the object localization problem to a certain extent. However, this method still has deficiencies in real-time performance. One-stage detection methods such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) directly predict the object class and location in an end-to-end manner, significantly improving the detection speed. These methods show strong robustness in dealing with multi-scale objects and complex backgrounds, but still face challenges in low-light complex scenes. In recent years, anchor-free detection methods have gradually attracted attention. Methods such as FCOS (Fully Convolutional One-Stage Object Detection) and CenterNet directly predict the center point and bounding box size of the object, avoiding the design and matching problems of anchors and performing well in multi-scale detection tasks. In addition, Transformer-based object detection methods such as DETR (DEtection TRansformer) achieve precise detection of objects in complex scenes by modeling long-range dependencies.
[0005] In low-light scenes, the combination of image enhancement technology and object detection has become a research hotspot. The Retinex algorithm and its improved versions are widely used in low-light image enhancement. By enhancing the brightness and contrast of the image, they provide high-quality input for pedestrian detection. In addition, low-light image enhancement methods based on the generative adversarial network (GAN) such as EnlightenGAN can further improve the detection performance under low-light conditions. In night-time and low-light scenes, the difficulty of pedestrian detection increases further. Low-light conditions lead to a decline in image quality and a lower contrast between the object and the background, thus reducing the detection rate and recall rate of the pedestrian detection model. In this case, the problems of missed detection and false detection are more prominent, directly affecting the reliability of the detection system. In recent years, the development of low-light image enhancement technology has provided new solutions to solve this problem. By enhancing the brightness and contrast of the image, the visibility of pedestrians in low-light scenes can be effectively improved, providing higher-quality input data for the detection model. In addition, the rapid development of Transformer technology in the field of computer vision has also brought new opportunities to pedestrian detection research. Summary of the Invention
[0006] The object of the present invention is to provide a low-light pedestrian detection method based on an anchor-free lightweight vision Transformer. Aiming at the influence of low light on pedestrian detection performance in night scenes and low-light scenes, the present invention proposes a low-light anchor-free box pedestrian detection method based on an anchor-free lightweight vision Transformer to overcome the deficiencies existing in the prior art, such as the technical problems of missed detection, as well as the decline in accuracy and speed in pedestrian detection in night and low-light scenes, and to improve the accuracy, speed and robustness of pedestrian detection.
[0007] The inventive concept of the present invention is as follows: First, the LYT-Net algorithm is used to enhance low-light images to reduce the missed detection rate. During the training process, a hybrid data augmentation technique is used to improve the robustness of the model in complex scenes. Then, the backbone network selects a network architecture that combines Transformer and a convolutional neural network. This architecture combines the advantages of Transformer and the convolutional neural network CNN, and can effectively capture long-range dependencies and local details in feature extraction. By introducing a Transformer module into the low-level convolutional feature map, the network can perform deeper information fusion between multi-scale features, thus improving the detection ability for night pedestrian targets. Next, a lightweight self-attention mechanism based on Transformer is introduced to enhance the model's attention to important feature regions, especially in complex backgrounds and low-light conditions. At the same time, through parameterized convolution, the computational amount is reduced while ensuring the detection accuracy, and the inference speed is improved. Finally, the loss function is optimized to further optimize the localization accuracy of the model. In this way, the model can more accurately identify and locate pedestrians, improving the detection accuracy and recall rate in night environments. Finally, through experiments on the night-time autonomous driving dataset, the results show that this method can significantly improve the accuracy and robustness of night detection, especially under low-light conditions. Compared with traditional methods, both the accuracy and speed are significantly improved.
[0008] To achieve the above-mentioned inventive object, the technical solution adopted by the present invention is specifically as follows: A low-light pedestrian detection method based on an anchor-free lightweight vision Transformer, comprising the following steps:
[0009] S1: Obtain a pedestrian dataset, preprocess it into a processable image, perform standardization and normalization processing on the image, and scale it to a fixed size to meet the input requirements of the model;
[0010] S2: Perform data augmentation processing. Perform low-light image enhancement on the obtained data to increase the brightness of the image and enhance the contrast. Then perform hybrid data augmentation processing on the processed data to improve the robustness. Divide the dataset into a training set, a validation set and a test set;
[0011] S3: Construct the YOLOV8 model as an anchor-free base model, and its overall structure includes three parts: the backbone network Backbone, the neck structure Neck, and the detection head Head;
[0012] S4: Construct a lightweight vision Transformer model as the feature extraction network of the model;
[0013] S5: Incorporate a self-attention mechanism that simultaneously focuses on information in both spatial and temporal dimensions into the Backbone and Neck parts;
[0014] S6: Introduce reparameterized convolutions to reduce computational overhead, further optimize the model's feature extraction ability, and improve the model's detection performance in low-light scenarios;
[0015] S7: Use the Wise-IOUv3 loss function to replace the CIOU loss;
[0016] S8: Input the dataset into the model for training iteration until the loss value is stable. At the same time, input the validation set into the model to obtain the model with the optimal training weights, and save the model parameter weights. Test the images in the test set to finally obtain the detection results.
[0017] Furthermore, the publicly available dataset BDD100K in step S1 is an image dataset in the fields of autonomous driving and intelligent transportation. The files are converted into processable image files through preprocessing. The dataset contains 100,000 high-resolution driving scene images, covering various weather, lighting, scene and other different conditions. Among them, there are 70,000 training images, 20,000 test images, and 10,000 validation images. To adapt to the low-light pedestrian detection task, we first process the BDD100K dataset to make it more suitable for model training and validation.
[0018] Furthermore, the step S2: Dataset preprocessing specifically includes the following steps:
[0019] S21: Perform low-light image enhancement on the acquired data to improve the brightness of the image and enhance the contrast. The LYT-Net method is used, which is a low-light image enhancement method based on the combination of Transformer attention mechanism and CNN. It adopts a Y-shaped network architecture, combines the CNN and Transformer mechanisms to achieve global brightness adjustment and local detail enhancement.
[0020] S22: Perform mixed data augmentation on the processed data, including erasing, splicing, rotating, translating, flipping, and color perturbation, etc. These operations help to increase the diversity of the data and improve the model's adaptability to different scenarios and different lighting conditions.
[0021] S23: The dataset is divided into a training set, a validation set, and a test set. Among them, there are 11,340 images in the training set, 1,260 images in the validation set, and 1,400 images in the test set.
[0022] Furthermore, in step S3: Construct the YOLOV8 model as an anchor-free basic model, which specifically includes the following steps:
[0023] S31: YOLOv8 adopts an object detection architecture based on a convolutional neural network (CNN). The overall structure still follows the design of three parts: the backbone network, the neck structure, and the detection head. Among them, the backbone network is used to extract the deep features of the input image, the neck structure further fuses multi-scale information, and the detection head is used for the final object classification and localization prediction. In the backbone network part, the C2f module and SPPF are adopted. Compared with the previous CSPDarkNet structure, this module reduces computational redundancy and enhances the cross-layer feature interaction ability, enabling the network to obtain stronger representation ability at a lower computational cost. In terms of feature fusion, the FPN and PAN structures are combined. Through the top-down and bottom-up feature transfer mechanisms, features of different scales can fully interact. The FPN structure is mainly used to improve the detection ability of small targets, while the PAN structure enhances the information flow transmission between feature layers, enabling the network to more accurately predict targets of different scales.
[0024] Furthermore, in step S4: Construct a lightweight vision Transformer model as the feature extraction network of the model. Specifically, it includes the following:
[0025] Adopt a hybrid network architecture MobileViT-v3 that combines CNN and Transformer as the backbone network. MobileViT-v3 combines the advantages of convolutional neural networks (CNN) and Transformer. While maintaining the powerful representation ability of the Transformer model, it significantly improves the computational efficiency of the model by reducing the model size, computational overhead, and the number of parameters. Effective feature fusion and self-attention mechanisms successfully alleviate the noise impact and improve the clarity and contrast of the image.
[0026] Further, in step S5: A self-attention mechanism that simultaneously focuses on information in both spatial and temporal dimensions is incorporated into the Backbone and Neck parts. TSA attention is introduced. The TSA attention mechanism further improves the performance of the object detection task in low-light scenarios by simultaneously focusing on information in both spatial and temporal dimensions. Different from traditional attention mechanisms, the TSA mechanism can combine the spatial features of images with temporal information, strengthening the network's ability to model local and global information of images. Especially in dynamic low-light image sequences, it can provide more refined and dynamic attention capabilities.
[0027] Further, in step S6: Reparameterized convolution is introduced to reduce computational overhead. By adopting the RCS reparameterized convolution method that combines reparameterized convolution and channel shuffle strategies, the feature expression ability of the model is enhanced during the training phase, and the computational graph is simplified during the inference phase, thereby reducing computational overhead while ensuring more sufficient information interaction between different channels.
[0028] Further, in step S7: The Wise-IOUv3 loss function is used to replace the CIOU loss. The Wise-IOUv3 loss function can more effectively handle the deviation of the model when predicting the boundary, improving the accuracy of object localization. Wise-IOUv3 focuses on the overlapping area between the predicted bounding box and the ground truth bounding box, providing precise gradient updates during the bounding box regression process to accurately fit the boundary of the object.
[0029] Further, in step S8: The dataset is input for training iterations until the loss value stabilizes. The validation set obtains the model with the optimal training weights, and the model parameter weights are saved. The test set is used for testing, and finally the detection results are obtained. A total of 300 epochs are trained, and the SGD optimizer is used. In terms of hyperparameter settings, the maximum learning rate of the model is set to 1e-2, and the minimum learning rate is set to 1e-4. The learning rate is adjusted in a cosine manner. To stabilize the training process, an SGD (Stochastic Gradient Descent) optimizer with a momentum of 0.937 is used, and the weight decay is set to 5e-4. The size of the input image is uniformly 640×640, the batchsize is 8, and a total of 300 epochs are trained.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. A low-light pedestrian detection method based on an anchor-free lightweight vision Transformer, the LYT-Net low-light enhancement technology is used to improve the quality of images under low-light conditions. It is a low-light image enhancement method that combines the Transformer attention mechanism and CNN, effectively enhancing the image brightness, enhancing the contrast, and retaining important texture details, providing clearer and more useful input data for subsequent pedestrian detection tasks, and improving the accuracy and stability of object detection.
[0032] 2. In the present invention, the lightweight MobileViT-v3 network is used as the feature extraction network, which greatly reduces the computational complexity of the model while maintaining high accuracy. The design of the network combines the advantages of the convolutional neural network CNN and Transformer, and can effectively capture long-range dependencies and local details in feature extraction; by introducing the Transformer module into the low-level convolutional feature map, the network can perform deeper information fusion between multi-scale features, thus improving the detection ability for nighttime pedestrian targets.
[0033] 3. The present invention adopts the TSA attention mechanism, which enhances the model's attention to key features by weighted attention to key spatial regions in the image; combined with the RCS (Reparameterized Convolutional Separable) reparameterized convolution, it effectively reduces the overhead pressure brought by the Transformer module and the attention mechanism, enhances the feature expression ability of the model in the training stage, simplifies the computational graph in the inference stage, thereby reducing the computational overhead, and ensuring more sufficient information interaction between different channels.
[0034] 4. In the present invention, the Wise-IOUv3 loss function is used to optimize the loss function, which alleviates the problem that the regression gradient is not balanced when the target scale difference is large or the overlap degree between the predicted box and the ground truth box is small, resulting in slow model convergence or bounding box jitter, and realizes more refined regression optimization. The method of the present invention effectively improves the target localization accuracy and speed, making the model have stronger robustness and accuracy in low-light environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.
[0036] Figure 1 Schematic diagrams of some images of the data set obtained for the present invention.
[0037] Figure 2Schematic diagram of the backbone module of the MobileViT-v3 network of the present invention.
[0038] Figure 3 Schematic diagram of the principle of reparameterized convolution of the present invention.
[0039] Figure 4 Flow chart of the model of the present invention.
[0040] Figure 5 Overall model framework diagram of the present invention.
[0041] Figure 6 Effect diagram of model detection of the present invention.
[0042] Figure 7 Comparison diagram of model detection of the present invention. Specific implementation mode
[0043] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] Embodiment 1
[0045] This embodiment provides a low-light pedestrian detection method based on an anchor-free lightweight vision Transformer. The overall process of the method is as shown in the appendix Figure 1 shown. The technical solutions of the present invention will be further elaborated below with reference to specific embodiments:
[0046] S1: Preprocess the input data, which specifically includes the following steps:
[0047] S1.1: Obtain a pedestrian dataset from an in-vehicle dataset or a public dataset BDD100K, convert the file into an image type file that can be processed, and perform standardization and normalization processing. Uniformly scale the images to a fixed size of 640×640 to meet the model input requirements.
[0048] S1.2: Read the data, use the cvtColor() function to convert the format to RGB, and finally convert it into an Image image for detection. For the dataset, select images in the three time periods of dawn, dusk, and night, and then screen the image files according to the selected time period label files. Convert the file annotation format json to the XML file required for training, and then convert the dataset to the VOC format for training. The processed image files are placed in a separate folder, and the label files are converted into a single file containing the picture name, picture path, picture label name, and target position coordinates. Loop to delete files without targets to reduce the redundancy of the sample training set.
[0049] S2: Data augmentation processing, divided into training set, validation set and test set. The specific method of this step is as follows:
[0050] After converting the image to RGB, use resize_image() to add gray bars to the image, scale the image to 640×640 size and keep it distortion-free. Randomly select different cropping regions to simulate different scenario situations. Normalize the images in the dataset, scale each pixel value from 0 - 255 to 0 - 1. This operation can not only accelerate the training process of the model, but also ensure that the model maintains a consistent input scale among different images, reducing the problem of gradient disappearance or explosion during the training process. Other forms of data augmentation are performed on the dataset, including rotation, translation, flipping and color perturbation, etc. These operations help to increase the diversity of the data and improve the adaptability of the model to different scenarios and different lighting conditions.
[0051] S2.2: After data augmentation, perform dataset division. Divide the dataset into training set, validation set and test set according to a certain proportion. Among them, the training set is 11340, the validation set is 1260, and the test set is 1400. The dataset is processed to generate train.txt and val.txt files containing absolute paths, numbers and specific coordinates of detection targets. Some images are as attached Figure 1 as shown.
[0052] S3: Build the YOLOV8 model as an anchor-free basic model. It includes the design of three parts: the backbone network, the neck structure and the detection head, which are responsible for feature extraction, multi-scale feature fusion and target classification regression localization respectively.
[0053] S4: Build a lightweight vision Transformer model as the feature extraction network of the model, which specifically includes the following steps
[0054] S4.1: Initialize the network weights, initialize the MobileViT-v3 network using pre-trained weights to accelerate convergence and improve generalization ability.
[0055] S4.2: Configure the feature extraction layer, adjust the feature extraction layer of the MobileViT-v3 network to adapt to the input of the Neck, and ensure the change of the height, width and number of channels of the effective feature layer. For the modules in the backbone, as attached Figure 2 as shown.
[0056] S5: Incorporate the Temporal Self-Attention (TSA) mechanism that simultaneously focuses on information in both spatial and temporal dimensions. The TSA attention mechanism first performs weighted processing on the input image in the spatial dimension. Suppose the input feature map is \(X\in\mathbb{R}^{H\times W\times C}\), where \(H\) and \(W\) represent the height and width of the image respectively, and \(C\) represents the number of channels of the feature map. In the calculation of spatial attention, the feature map calculates the similarity between each position through the inner product of the Query, Key, and Value matrices, thereby weighting different regions of the image to enhance the representation of important features. The calculation process of spatial attention can be expressed as:
[0057]
[0058] where \(Q\), \(K\), and \(V\) are the Query, Key, and Value matrices respectively, and \(d_k\) is the dimension of the key. Through this weighted operation, the model can automatically focus on key regions according to the spatial information of the image, thereby effectively suppressing noise and irrelevant backgrounds.
[0059] Then, temporal information is introduced into the attention calculation, enabling the model to not only capture spatial information but also capture dynamic changes in the image sequence by focusing on the relationships between time dimensions. The calculation formula for temporal attention is:
[0060]
[0061] where \(Q\) t , \(K\) t , \(V\) t represent the Query, Key, and Value matrices at the current moment respectively.
[0062] S6: Introduce reparameterized convolution to reduce computational overhead. In the training phase, multiple parallel convolutional branches are adopted, including standard \(3\times3\) convolution, \(1\times1\) convolution, and identity mapping, to enhance the model's ability to capture local and global information. Let the input feature map be \(X\in\mathbb{R}^{C\times H\times W}\), then the calculation formula for the training phase of RCS is as follows:
[0063] \(Y = X*W_3+X*W_1+IX\)
[0064] where \(W_3\) represents the \(3\times3\) convolutional kernel, \(W_1\) is the \(1\times1\) convolutional kernel, and \(I\) represents the identity mapping; this structure can fully utilize feature information at different scales during the training process, thereby improving the richness of features. At the same time, a channel shuffle strategy is further introduced. Specifically, the input features are first divided into \(G\) groups, each group contains \(C / G\) channels, and then rearranged between groups so that features from different groups can be fused. Finally, the rearranged features are merged back to the original channel dimension. This process not only enhances information sharing between different channels but also improves the model's understanding ability of global features, thereby improving detection accuracy. As attachedFigure 3 as shown
[0065] S7: Optimize the loss function, and use the Wise-IOUv3 loss function to replace the CIOU loss. YOLOv8 contains three loss functions, namely classification loss, regression loss, and object confidence loss. Among them, the regression loss uses the CIoU loss. The traditional IoU (Intersection over Union) loss function mainly measures the matching degree between the predicted bounding box and the ground truth bounding box by calculating their intersection over union. However, when the predicted bounding box and the ground truth bounding box do not overlap at all, the IoU value is zero, which cannot provide an effective backpropagation gradient; in addition, IoU only focuses on area overlap and does not fully consider the distance between the center points of the bounding boxes and the aspect ratio, thus limiting the further improvement of the positioning accuracy. Wise-IOUv3, which introduces an adaptive balance mechanism, enables the loss function to dynamically adjust the weights of different error terms according to the scale of the target and the overlap degree of the bounding boxes, so as to achieve more refined regression optimization. This step is as shown in the appendix Figure 4 as shown. After the above steps, the overall framework of the model is as shown in the appendix Figure 5 as shown
[0066] S8: Model training and testing. Input the dataset into the improved model for training iteration until the loss value is stable. It is carried out on the Windows 10 operating system, using an NVIDIA GeForce GTX 4060Ti graphics card with 8GB of GPU video memory and 16GB of system memory. Use the Pytorch deep learning framework for model training and testing. In terms of hyperparameter settings, the maximum learning rate of the model is set to 1e-2, and the minimum learning rate is set to 1e-4. The learning rate is adjusted using cos. To stabilize the training process, an SGD (Stochastic Gradient Descent) optimizer with a momentum of 0.937 is used, and the weight decay is set to 5e-4. The size of the input image is uniformly 640×640, the batchsize is 8, and a total of 300 epochs are trained. During the training process, the validation set is input into the model to obtain the optimal model weights. Save the parameter weights of the optimal model. Input the test set into the model to output the final detection results. After training, the loss and accuracy of the model gradually tend to be stable. The model detection results are as shown in the appendix Figure 6 as shown. The above can improve the detection accuracy and missed detection rate of target pedestrians in low-light scenarios while ensuring the detection speed, and is applicable to fields such as autonomous driving and intelligent transportation
[0067] Example 2
[0068] To comprehensively evaluate the performance of the low-light pedestrian detection method proposed in this embodiment, comparative experiments were conducted between this embodiment and multiple existing models. The comparative models include: SSD, Faster R-CNN, RetinaNet, FCOS, YOLOv6s, DETR, YOLOv8 (baseline model), and the method proposed in this embodiment. The experiments were based on the BDD100K low-light pedestrian detection dataset and used the same preprocessing, input size, and training settings. The following Table 1 shows the comparison results of each model.
[0069] Table 1
[0070]
[0071] As can be seen from the experimental results in Table 1, the model proposed in this embodiment has achieved better performance than other models in multiple evaluation metrics.
[0072] Specifically, in terms of detection accuracy, it reached 78.42% on the mAP@0.5 metric, an increase of 3.52% compared to the baseline model YOLOv8 and 5.38% compared to YOLOv6s, far higher than SSD (an increase of 14.99%) and FCOS (an increase of 8.27%). In addition, on the more stringent mAP@0.5:0.95 metric, YOLO-LFormer reached 44.35%, an increase of 2.56% compared to YOLOv8, indicating that the method in this embodiment can maintain a high detection accuracy under different IoU thresholds.
[0073] In terms of detection speed, the FPS of the model in this embodiment reached 58.30, a decrease of 2.06 compared to YOLOv8 (60.36), but it is relatively high compared to other models and far exceeds Faster R-CNN (19.26), indicating that the method in this embodiment can still maintain a high real-time performance while ensuring accuracy.
[0074] In terms of computational complexity, the GFLOPs of this embodiment is 45.31, which is higher than that of YOLOv8 (28.82), but still lower than that of Faster R-CNN (127.33) and RetinaNet (99.42), indicating that the improved model maintains a good balance in computational resource consumption. In addition, the number of parameters of the model in this embodiment is 36.53M, an increase of 0.82M compared to YOLOv8 (35.71M), indicating that the method in this embodiment can effectively control the scale of model parameters while improving accuracy, which helps to reduce storage and deployment costs.
[0075] Example 3
[0076] To further evaluate the performance of the improved model in actual low-light scenarios, low-light scenarios under dawn, dusk, night, rainy days, and snowy days were separately selected from the dataset for comparative experiments with the baseline model. The experimental results are as shown in the appendix Figure 7 As shown. Through comparative analysis, it can be seen that the model proposed in this embodiment exhibits better detection capabilities than the basic model under various low-light and harsh weather conditions. Whether in low-light scenarios such as dawn, dusk, and night, or in complex weather conditions such as rainy days and snowy days, the method of this embodiment can effectively improve the pedestrian detection accuracy and reduce the phenomena of missed detection and false detection. It has more excellent detection capabilities and shows more stability.
[0077] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A low-light pedestrian detection method based on an anchor-free lightweight vision Transformer, characterized in that It includes the following steps: S1: Obtain a pedestrian dataset, preprocess it into a processable image, perform standardization and normalization on the image, and scale it to a fixed size to meet the input requirements of the model; S2: Perform data augmentation processing. Perform low-light image enhancement on the obtained data to increase the brightness and enhance the contrast of the image. Then perform mixed data augmentation processing on the processed data, and divide the dataset into a training set, a validation set, and a test set; S3: Construct the YOLOV8 model as an anchor-free basic model. The overall structure includes three parts: the backbone network Backbone, the neck structure Neck, and the detection head Head; S4: Construct a lightweight vision Transformer model as the feature extraction network of the model; S5: Incorporate the self-attention mechanism that simultaneously focuses on information in both the spatial and temporal dimensions for the backbone network Backbone and the neck structure Neck parts; S6: Introduce reparameterized convolutions to reduce computational overhead, optimize the feature extraction ability of the model, and improve the detection performance of the model in low-light scenarios; S7: Use the Wise-IOUv3 loss function to replace the CIOU loss; S8: Input the dataset into the model for training iteration until the loss value is stable. At the same time, input the validation set into the model to obtain the model with the optimal training weights, save the model parameter weights, test the images in the test set, and finally obtain the detection results.
2. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, wherein: The specific content of step S1 is as follows: The publicly available dataset BDD100K is an image dataset in the fields of autonomous driving and intelligent transportation. Through preprocessing, the files are converted into processable image files. The dataset contains 100,000 high-resolution driving scene images, covering various weather, lighting, and scene conditions. Among them, there are 70,000 images in the training set, 20,000 images in the test set, and 10,000 images in the validation set; to adapt to the low-light pedestrian detection task, first process the BDD100K dataset to make it suitable for model training and validation.
3. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, wherein: The step S2: Dataset preprocessing includes the following steps: S21: Perform low-light image enhancement on the obtained data. Use the LYT-Net method, which is a low-light image enhancement method based on the combination of Transformer attention mechanism and CNN. It adopts a Y-shaped network architecture and combines the CNN and Transformer mechanisms to achieve global brightness adjustment and local detail enhancement; S22: Perform mixed data augmentation on the processed data, including erasing, splicing, rotating, translating, flipping, and color perturbation; S23: Divide the dataset into a training set, a validation set, and a test set. Among them, there are 11,340 images in the training set, 1,260 images in the validation set, and 1,400 images in the test set.
4. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, characterized in that: The step S3: Construct the YOLOV8 model as an anchor-free basic model, including the following steps: S31: YOLOv8 adopts an object detection architecture based on the convolutional neural network CNN. The overall structure follows the design of three parts: the backbone network Backbone, the neck structure Neck, and the detection head Head; S32: The Backbone is used to extract the deep features of the input image, the Neck structure fuses multi-scale information, and the Head is used for the final object classification and localization prediction.
5. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, wherein: In step S4, a lightweight vision Transformer model is constructed as the feature extraction network of the model, specifically: The hybrid network architecture MobileViT-v3 that combines CNN and Transformer is used as the Backbone, and MobileViT-v3 combines the convolutional neural network CNN and Transformer.
6. The low-light pedestrian detection method based on the anchor-free lightweight vision Transformer according to claim 1, wherein: Step S5: For the Backbone and Neck parts, a self-attention mechanism that simultaneously focuses on information in both spatial and temporal dimensions is incorporated. The TSA attention mechanism focuses on information in both spatial and temporal dimensions.
7. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, wherein: Step S6: Reparameterized convolutions are introduced to reduce the computational overhead. By adopting the RCS reparameterized convolution method that combines reparameterized convolutions and channel shuffle strategy, the feature expression ability of the model is enhanced during the training phase, and the computational graph is simplified during the inference phase.
8. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, wherein: Step S7: The Wise-IOUv3 loss function is used to replace the CIOU loss. The Wise-IOUv3 loss function handles the deviation of the model when predicting the boundaries.
9. The low-light pedestrian detection method based on an anchor-free lightweight vision Transformer according to claim 1, characterized in that: In step S8, the dataset is input for training iterations until the loss value is stable. The validation set obtains the model with the optimal training weights, and the model parameter weights are saved.