Method and System for Road Condition Monitoring under Complex Lighting Conditions Based on a Lightweight Enhancement Network
By combining the lightweight enhanced network PLED-Net, TSD-Net and Expand-Net combined with the PP-YOLOE-PLUS-Tiny model, the accuracy and real-time problems of the road condition monitoring system in complex light environments are solved, and efficient and low-cost road condition monitoring is achieved.
Patent Information
- Application Number
- CN202510393256.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The accuracy and reliability of existing road condition monitoring systems have significantly decreased in complex light environments (such as low light, astigmatism, and backlight), and relying on multimodal sensors and complex image processing algorithms has increased system complexity and cost, making it difficult to operate efficiently on edge computing platforms.
Lightweight enhancement networks are adopted, including low-light enhancement network PLED-Net, astigmatism enhancement network TSD-Net and backlight enhancement network Expand-Net, combined with the PP-YOLOE-PLUS-Tiny model, and video streams collected by the camera are processed in real time through the USB transmission protocol to achieve image enhancement and object detection.
Improve image quality and detection accuracy under a variety of complex light conditions, reduce dependence on multimodal sensors, maintain efficient real-time processing capabilities, and reduce system complexity and cost.
Smart Images

Figure CN119888647B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent traffic condition detection, and particularly to a traffic condition monitoring method based on multiple lightweight image enhancement networks. Background Art
[0002] With the acceleration of the urbanization process, the traffic condition monitoring system plays a crucial role in ensuring traffic safety and optimizing traffic management. Most of the existing traffic condition monitoring systems rely on cameras to collect video data and use image processing and object detection algorithms to monitor and analyze traffic elements such as vehicles and pedestrians. However, complex light environments (such as strong light, shadows, and night lighting) often pose challenges to the accuracy of the traffic condition monitoring system.
[0003] Although the existing traffic condition monitoring systems can work normally under general lighting conditions, in complex light environments such as low light (decreased image contrast and increased noise), scattered light (blurred image and lost details), and backlight (overexposed image and shadow areas), the object detection accuracy and reliability are significantly reduced. To address these challenges, the existing systems usually rely on multi-modal sensor fusion or complex image processing algorithms, which not only increase the system complexity and cost but also require high hardware performance and are difficult to operate efficiently on edge computing platforms, resulting in insufficient real-time performance. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a traffic condition monitoring method and system under complex light conditions based on lightweight enhancement networks, belonging to the field of intelligent traffic condition detection. The present invention uses the low-light enhancement network PLED-Net to optimize the low contrast and noise problems of images in low-light environments; on the basis of the low-light enhancement network PLED-Net, the scattered light enhancement network TSD-Net is introduced to improve the clarity and detail performance of images in scattered light environments; at the same time, on the basis of the low-light enhancement network PLED-Net, the backlight enhancement network Expand-Net is introduced to effectively solve the overexposure and shadow problems in backlight environments. The present invention significantly improves the image quality under multiple complex light conditions by using lightweight enhancement networks; the traffic condition detection system includes a camera and an edge computing platform, and the traffic condition detection system can work normally under multiple complex light conditions, reducing the dependence on multi-modal sensors, while ensuring efficient real-time processing and reducing the cost of traffic condition monitoring.
[0005] A road condition monitoring system based on a lightweight enhancement network. The road condition monitoring system includes a camera and an edge computing platform. The edge computing platform includes a lightweight enhancement network and a PP-YOLOE-PLUS-Tiny model. The lightweight enhancement network includes a low-light enhancement network PLED-Net, a glare enhancement network TSD-Net, and a backlight enhancement network Expand-Net. The low-light enhancement network PLED-Net, the glare enhancement network TSD-Net, and the backlight enhancement network Expand-Net are three independent enhancement networks. The lightweight enhancement network operates independently according to the real-time road condition environment. When the road condition environment is under various complex lighting conditions, the three independent enhancement networks are combined to enhance the image simultaneously. The camera collects road condition videos and transmits the video stream to the edge computing platform in real time through the USB transmission protocol. The lightweight enhancement network splits and enhances the video stream, inputs the enhanced image into the PP-YOLOE-PLUS-Tiny model for target monitoring, and combines the monitoring results into a video stream for visual display.
[0006] A method for monitoring road conditions under complex lighting based on a lightweight enhancement network. The road condition monitoring method includes the following steps:
[0007] Step 1: Use a USB camera to collect road condition videos and transmit the collected road condition video stream to the edge computing platform in real time through the USB transmission protocol. The edge computing platform includes a lightweight enhancement network and a PP-YOLOE-PLUS-Tiny model. The lightweight enhancement network includes a low-light enhancement network PLED-Net, a glare enhancement network TSD-Net, and a backlight enhancement network Expand-Net.
[0008] Step 2: Initialize the parameters of the lightweight enhancement network and the PP-YOLOE-PLUS-Tiny model. When starting the road condition monitoring system, the parameters are automatically initialized uniformly and loaded into the road condition monitoring system at one time. The lightweight enhancement network and the PP-YOLOE-PLUS-Tiny model remain in an active state during the operation of the road condition monitoring system until all weights are released when the road condition monitoring system exits.
[0009] Step 3: The lightweight enhancement network dynamically selects the low-light enhancement network PLED-Net, the glare enhancement network TSD-Net, and the backlight enhancement network Expand-Net according to the light environment for image enhancement processing to obtain a single-frame enhanced image.
[0010] Step 4: Input the single-frame enhanced image into the PP-YOLOE-PLUS-Tiny model for detection to obtain a road condition monitoring image.
[0011] Step 5, batch-combine the road condition monitoring images into a video stream for visual display; batch-combining means combining 5 images into a batch.
[0012] Furthermore, in Step 3, the single-frame enhanced image is a low-light enhanced image, a defocused light enhanced image, a backlight enhanced image, or a fusion enhanced image; the fusion enhanced image is a low-light defocused light enhanced image, a low-light backlight enhanced image, a defocused light backlight enhanced image, or a low-light defocused light backlight enhanced image.
[0013] Furthermore, in Step 3, the low-light enhancement network PLED-Net is based on the Retinex theory, and the low-light image enhancement process includes five stages: enhancement, denoising, enhancement, denoising, and enhancement in sequence; the low-light enhancement network realizes high-quality image enhancement through progressive enhancement and adaptive denoising; the low-light enhancement network PLED-Net includes a progressive multi-scale residual illumination enhancement module (PMSRIEB) and a detail enhancement and adaptive denoising module (DEANRB);
[0014] In the enhancement stage, the PMSRIEB module is adopted; the PMSRIEB module sequentially includes a multi-scale feature extraction layer, a channel attention module, a color interaction module, a global-local illumination estimation layer, and a spatial attention module; the multi-scale feature extraction layer includes four convolutional layers, and the four convolutional layers are in parallel, and the convolutional kernels of the four convolutional layers are 1x1, 3x3, 5x5, and 7x7 respectively; the color interaction module is a 3x3 grouped convolutional layer; the global-local illumination estimation layer includes two convolutional layers, and the two convolutional layers are in parallel, and the convolutional kernels of the two convolutional layers are 3x3 and 7x7 respectively; the low-light image first passes through the multi-scale feature extraction layer to extract the illumination feature, and then the channel attention module applies attention to the illumination feature; then the color interaction module improves the correlation of the three RGB channels; then the global-local illumination estimation layer estimates the illumination from the global scale and the local scale, and the spatial attention module applies different enhancement degrees according to different regions to obtain a low-light output image; finally, the low-light output image is added to the low-light image residual to obtain a preliminary enhanced image;
[0015] In the denoising stage, the DEANRB module is adopted; the DEANRB module includes differential convolution, a noise estimation module, and a spatial attention module; the low-light image and the preliminary low-light enhanced image are simultaneously used as the inputs of the differential convolution, and the differential convolution extracts detailed features; the preliminary low-light enhanced image passes through the noise estimation module to obtain a noise estimation image; the illumination component feature map obtained by dividing the low-light image by the preliminary low-light enhanced image is used to calculate the SNR map, and the SNR map obtains a noise spatial attention feature map through the spatial attention module, and the noise attention feature map is multiplied by the estimated noise image to achieve adaptive denoising; finally, the noise estimation image is subtracted from the preliminary low-light enhanced image to obtain a preliminary denoised image, and the detailed features are added to the preliminary denoised image to obtain a denoised image;
[0016] The calculation process of the SNR map is as follows: First, median filtering is performed on the illumination component feature map to obtain a filtered illumination component feature map; then, the absolute value is taken after subtracting the filtered illumination component feature map from the illumination component feature map to obtain a noise map; next, the filtered illumination component feature map is divided by the noise map to obtain the SNR map.
[0017] Furthermore, the differential convolution includes central difference, horizontal difference, vertical difference, and alignment difference.
[0018] Furthermore, the noise estimation module sequentially includes 12 convolutional blocks; each convolutional block includes a 3x3 convolutional layer, a normalization layer, and a LeakyReLU layer; the convolutional layer, the normalization layer, and the LeakyReLU layer are connected in series in sequence.
[0019] Furthermore, in step 3, the astigmatism enhancement network TSD-Net includes a multi-scale attention extraction module, a double-branch feature extraction module, and a multi-feature fusion module;
[0020] The multi-scale attention extraction module includes six convolutional units; among the six convolutional units, the first convolutional unit includes maximum pooling, a 3x3 convolutional layer, and a LeakyReLU layer; the second convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the third convolutional unit includes three parallel 3x3 dilated convolutional layers, and the dilation rates of the three dilated convolutional layers are 5, 7, and 9 respectively; the fourth convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the fifth convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the sixth convolutional unit includes a 3x3 convolutional layer and a Sigmoid layer;
[0021] The dual-branch feature extraction module includes two branches; the first branch adopts the UNet architecture, and the second branch adopts the DeseNet network; the first branch sequentially includes a basic feature extraction block, a downsampling block, a basic feature extraction block, a basic feature extraction block, a downsampling block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, an upsampling block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, an upsampling block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, and a basic feature extraction block; the basic feature extraction block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer, the downsampling block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer, and the upsampling block is bilinear interpolation.
[0022] The multi-feature fusion module sequentially includes an input block, 5 feature extraction blocks, and an output block; the input block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer; each feature extraction block includes two convolutional blocks, and each convolutional block includes a 3x3 convolutional layer, a LeakyReLU layer, and a residual block; after the 3x3 convolutional layer and the LeakyReLU layer are connected in series, they are then connected to the residual block; the output block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer.
[0023] Furthermore, the image enhancement process of the astigmatism enhancement network TSD-Net is as follows:
[0024] First, the astigmatic image is subjected to feature extraction through the multi-scale attention extraction module to obtain an astigmatic feature map;
[0025] Then, the astigmatic image and the astigmatic feature image are first multiplied pixel by pixel and then added to obtain an input feature extraction map;
[0026] Next, the input feature extraction map passes through the dual-branch feature extraction module to obtain two preliminary astigmatic enhancement images;
[0027] Finally, the input feature extraction map and the two preliminary astigmatic enhancement images are merged and input into the multi-feature fusion module to obtain an astigmatic enhancement image.
[0028] Furthermore, in step 3, the backlight enhancement network Expand-Net adopts a convolutional network with dual-branch convolution; the convolutional network with dual-branch convolution includes two branches; the first branch includes four 3x3 convolutional layers; the second branch includes seven 3x3 convolutional layers; both branches have 64 channels;
[0029] In step 3.1, the backlight image is processed by the first branch to obtain a detail feature image, and the image size of the detail image remains unchanged;
[0030] In step 3.2, the backlight image is gradually downsampled by the second branch to obtain a global feature image, and the size of the global feature image is reduced to 1x1;
[0031] Step 3.3: Fuse the detailed feature image and the global feature image, and input them into a 3x3 convolutional layer. The number of channels in the convolutional layer is reduced from 128 to 3 to obtain the backlight enhanced image.
[0032] Furthermore, in step 4, the detection process of the PP-YOLOE-PLUS-Tiny model is as follows:
[0033] The PP-YOLOE-PLUS-Tiny model includes an object detection module, an object tracking module, and an extended function module;
[0034] First, the object detection module accurately identifies and classifies a single-frame enhanced image, classifying the single-frame enhanced image as a pedestrian or a vehicle;
[0035] Then, the object tracking module saves the classified image of the object detection module and correlates the information between the upper and lower frames of the classified image of the object detection module to achieve continuous tracking of the same object;
[0036] Next, the extended function module detects the attributes of pedestrians or vehicles; pedestrian attributes include gender, clothing, and age characteristics, and vehicle attributes include vehicle feature information and vehicle violation detection. Vehicle feature information includes vehicle type, color, and brand; the extended function module detects vehicle feature information by segmenting and extracting text information from license plate images, and the extended function module performs vehicle violation detection on vehicle images by combining vehicle feature information and road traffic markings to determine whether there are vehicle violations; finally, a road condition monitoring image is obtained.
[0037] Furthermore, in step 4, the scheduling mode of the PP-YOLOE-PLUS-Tiny model is as follows:
[0038] Reconstruct the call mode of PaddleDetection in the PP-YOLOE-PLUS-Tiny model from command-line driven to modular; the modular scheduling mode enables the road condition monitoring system to dynamically select and call the corresponding function modules according to the input type and target category during operation, thereby enhancing the flexibility and adaptability of the road condition monitoring system; the input type is an image or a video, and the target category is a pedestrian or a vehicle.
[0039] The present invention has the following beneficial effects:
[0040] The image enhancement networks of the present invention are all lightweight networks, which can perform low-light, astigmatic, and backlight processing on the same original image respectively; they can either operate independently and be optimized for specific lighting conditions (such as low illuminance, foggy days, backlight), or be used in combination to handle more complex mixed lighting environments; in the case of the coexistence of low light and backlight, the road condition monitoring system can first perform low-light enhancement through PLED-Net and then perform backlight repair through Expand-Net. The flexible usage method not only improves the visible quality of the image in harsh lighting environments, but also significantly improves the accuracy of the detection algorithm, enhancing the robustness and accuracy of the road condition monitoring system in different lighting conditions; through lightweight design, the road condition monitoring system reduces the computational complexity of the system while maintaining high processing capabilities, enabling it to better adapt to the requirements of real-time road condition monitoring;
[0041] In response to the real-time processing requirements of the road condition monitoring system, the present invention adopts a fully lightweight model framework, which reduces the consumption of computing resources while maintaining the high performance of the road condition monitoring system, realizes low-latency real-time processing, and can quickly respond, monitor and analyze road conditions in a dynamically changing environment to ensure traffic safety; the road condition monitoring system is characterized by being lightweight, easy to deploy, having good robustness, and strong real-time performance, and can improve the image quality and the accuracy of traffic detection under various complex lighting conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the overall architecture diagram of the road condition monitoring method under complex lighting conditions for the lightweight enhancement network;
[0043] Figure 2 It is the schematic diagram of the low-light enhancement network structure;
[0044] Figure 3 It is the schematic diagram of the astigmatic enhancement module structure;
[0045] Figure 4 It is the schematic diagram of the backlight enhancement module structure;
[0046] Figure 5 It is the workflow diagram of the PP-YOLOE-PLUS-Tiny model with improved scheduling architecture. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0049] It should be noted that in the description of the present invention, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0050] The present invention provides a method and system for monitoring road conditions under complex light based on a lightweight enhancement network, belonging to the field of intelligent detection of traffic road conditions. The present invention adopts the weak light enhancement network PLED-Net and optimizes the low contrast and noise problems of images in low light environments; on the basis of the weak light enhancement network PLED-Net, the astigmatism enhancement network TSD-Net is introduced to improve the clarity and detail performance of images in astigmatic environments; at the same time, on the basis of the weak light enhancement network PLED-Net, the backlight enhancement network Expand-Net is introduced to effectively solve the overexposure and shadow problems in backlight environments. The present invention significantly improves the image quality under multiple complex light conditions by using a lightweight enhancement network; the road condition detection system includes a camera and an edge computing platform, and the road condition detection system can work normally under multiple complex light conditions, reducing the dependence on multi-modal sensors, and while ensuring efficient real-time processing, reducing the cost of road condition monitoring.
[0051] A road condition monitoring system under complex light based on a lightweight enhancement network, the road condition monitoring system includes a camera and an edge computing platform; the edge computing platform uses Jetson Orin Nano; the edge computing platform includes a lightweight enhancement network and a PP-YOLOE-PLUS-Tiny model; the lightweight enhancement network includes a weak light enhancement network PLED-Net, an astigmatism enhancement network TSD-Net, and a backlight enhancement network Expand-Net; the weak light enhancement network PLED-Net, the astigmatism enhancement network TSD-Net, and the backlight enhancement network Expand-Net are three independent enhancement networks; the lightweight enhancement network operates independently according to the real-time road condition environment. When the road condition environment is under multiple complex lights, the three independent enhancement networks are combined to enhance the image simultaneously; the camera collects road condition videos and transmits the video stream to the edge computing platform in real time through the USB transmission protocol; the lightweight enhancement network splits and enhances the video stream, inputs the enhanced image into the PP-YOLOE-PLUS-Tiny model for target monitoring, and combines the monitoring results into a video stream for visual display.
[0052] A method for monitoring road conditions under complex light based on a lightweight enhancement network, the road condition monitoring method includes the following steps:
[0053] As Figure 1 shown below:
[0054] Step 1: Use a USB camera to collect road condition videos and transmit the collected road condition video stream to the edge computing platform in real time through the USB transmission protocol; the edge computing platform includes a lightweight enhanced network and a PP-YOLOE-PLUS-Tiny model; the lightweight enhanced network includes a low-light enhancement network PLED-Net, a diffused light enhancement network TSD-Net, and a backlight enhancement network Expand-Net;
[0055] Step 2: Initialize the parameters of the lightweight enhanced network and the PP-YOLOE-PLUS-Tiny model: When starting the road condition monitoring system, the parameters are uniformly and automatically initialized and loaded into the road condition monitoring system at one time; the lightweight enhanced network and the PP-YOLOE-PLUS-Tiny model remain in an active state during the operation of the road condition monitoring system until all weights are released when the road condition monitoring system exits; each module can always maintain a low-latency state to ensure that the road condition monitoring system can operate efficiently and stably when real-time processing image enhancement and object detection tasks in a complex light environment.
[0056] Step 3: The lightweight enhanced network dynamically selects the low-light enhancement network PLED-Net, the diffused light enhancement network TSD-Net, and the backlight enhancement network Expand-Net according to the light environment for image enhancement processing to obtain a single-frame enhanced image; in a single complex light environment, the road condition monitoring system independently calls the corresponding enhancement network: use PLED-Net to improve image contrast and detail performance in low-light environments; use TSD-Net to improve image clarity and detail retention in diffused light environments; reduce overexposure and shadow problems through Expand-Net in backlight environments; for a more complex mixed light environment, the mixed light environment includes the simultaneous presence of low light and backlight, the simultaneous presence of diffused light and backlight, the simultaneous presence of low light and diffused light, and the simultaneous presence of diffused light, low light, and backlight; the road condition monitoring system can combine multiple enhancement networks to perform targeted enhancement processing on the image in sequence to ensure high-quality image output under various complex light conditions and provide better input for subsequent object detection; the input image is sent to different enhancement networks in sequence according to different scenarios, adopting the process of "low-light enhancement → diffused light enhancement → backlight enhancement", and according to actual needs, the input image can be selected to be processed through one or more enhancement networks.
[0057] Step 4: Input the single-frame enhanced image into the PP-YOLOE-PLUS-Tiny model for detection to obtain a road condition monitoring image;
[0058] Step 5: Batch-combine the road condition monitoring images into a video stream for visual display; batch-combining means combining 5 images into a batch.
[0059] In Step 3, the single-frame enhanced image is a low-light enhanced image, a diffused light enhanced image, a backlight enhanced image, or a fusion enhanced image; the fusion enhanced image is a low-light diffused light enhanced image, a low-light backlight enhanced image, a diffused light backlight enhanced image, or a low-light diffused light backlight enhanced image;
[0060] As Figure 2 shown:
[0061] In Step 3, the low-light enhancement network PLED-Net is based on the Retinex theory. The process of the low-light enhanced image is successively five stages of enhancement, denoising, enhancement, denoising, and enhancement; the low-light enhancement network realizes high-quality image enhancement through progressive enhancement and adaptive denoising; the low-light enhancement network PLED-Net includes a progressive multi-scale residual illumination enhancement module (PMSRIEB) and a detail enhancement and adaptive noise reduction module (DEANRB);
[0062] In the enhancement stage, the PMSRIEB module is adopted; the PMSRIEB module successively includes a multi-scale feature extraction layer, a channel attention module, a color interaction module, a global-local illumination estimation layer, and a spatial attention module; the multi-scale feature extraction layer includes four convolutional layers, and the four convolutional layers are in parallel. The convolutional kernels of the four convolutional layers are 1x1, 3x3, 5x5, and 7x7 respectively; the color interaction module is a 3x3 grouped convolutional layer; the global-local illumination estimation layer includes two convolutional layers, and the two convolutional layers are in parallel. The convolutional kernels of the two convolutional layers are 3x3 and 7x7 respectively; the low-light image first passes through the multi-scale feature extraction layer to extract the illumination feature, and then the channel attention module applies attention to the illumination feature; then the color interaction module improves the correlation of the three RGB channels; then the global-local illumination estimation layer estimates the illumination from the global scale and the local scale, and the spatial attention module applies different enhancement degrees according to different regions to obtain a low-light output image; finally, the low-light output image is added to the low-light image residual to obtain a preliminary enhanced image;
[0063] In the denoising stage, the DEANRB module is adopted; the DEANRB module includes differential convolution, a noise estimation module, and a spatial attention module; the low-light image and the preliminary low-light enhanced image are simultaneously used as the inputs of the differential convolution, and the differential convolution extracts detailed features; the preliminary low-light enhanced image passes through the noise estimation module to obtain a noise estimation image; the illumination component feature map obtained by dividing the low-light image by the preliminary low-light enhanced image is used to calculate the SNR map, and the SNR map passes through the spatial attention module to obtain a noise spatial attention feature map, and the noise attention feature map and the estimated noise image are multiplied to achieve adaptive denoising; finally, the noise estimation image is subtracted from the preliminary low-light enhanced image to obtain a preliminary denoised image, and the detailed features are added to the preliminary denoised image to obtain a denoised image;
[0064] The calculation process of the SNR map is as follows: First, median filtering is performed on the illumination component feature map to obtain a filtered illumination component feature map; then, the absolute value is taken after subtracting the filtered illumination component feature map from the illumination component feature map to obtain a noise image; next, the SNR map is obtained by dividing the filtered illumination component feature map by the noise image.
[0065] The low-light enhancement network PLED-Net adopts a phased enhancement and denoising strategy, significantly improving the contrast and detail performance of low-light images, effectively suppressing noise at the same time, and maintaining low computational complexity, making it suitable for efficient operation on edge computing platforms.
[0066] The differential convolution includes central difference, horizontal difference, vertical difference, and alignment difference.
[0067] The noise estimation module sequentially includes 12 convolutional blocks; each convolutional block includes a 3x3 convolutional layer, a normalization layer, and a LeakyReLU layer; the convolutional layer, the normalization layer, and the LeakyReLU layer are connected in series in sequence.
[0068] As Figure 3 shown:
[0069] In step 3, the astigmatism enhancement network TSD-Net includes a multi-scale attention extraction module, a dual-branch feature extraction module, and a multi-feature fusion module;
[0070] The multi-scale attention extraction module includes six convolutional units. Among the six convolutional units, the first convolutional unit includes a maximum pooling layer, a 3x3 convolutional layer, and a LeakyReLU layer; the second convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the third convolutional unit includes three parallel 3x3 dilated convolutional layers with dilation rates of 5, 7, and 9 respectively; the fourth convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the fifth convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the sixth convolutional unit includes a 3x3 convolutional layer and a Sigmoid layer.
[0071] The double-branch feature extraction module includes two branches. The first branch adopts a UNet architecture, and the second branch adopts a DeseNet network. The first branch sequentially includes basic feature extraction blocks, downsampling blocks, basic feature extraction blocks, basic feature extraction blocks, downsampling blocks, basic feature extraction blocks, basic feature extraction blocks, basic feature extraction blocks, basic feature extraction blocks, upsampling blocks, basic feature extraction blocks, basic feature extraction blocks, basic feature extraction blocks, upsampling blocks, basic feature extraction blocks, basic feature extraction blocks, basic feature extraction blocks, and basic feature extraction blocks. The basic feature extraction block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer, the downsampling block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer, and the upsampling block is bilinear interpolation.
[0072] The multi-feature fusion module sequentially includes an input block, 5 feature extraction blocks, and an output block. The input block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer. Each feature extraction block includes two convolutional blocks, and each convolutional block includes a 3x3 convolutional layer, a LeakyReLU layer, and a residual block. The 3x3 convolutional layer and the LeakyReLU layer are connected in series and then connected to the residual block. The output block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer.
[0073] The image enhancement process of the astigmatism enhancement network TSD-Net is as follows:
[0074] First, the astigmatism image is subjected to feature extraction through the multi-scale attention extraction module to obtain an astigmatism feature map.
[0075] Then, the astigmatism image and the astigmatism feature image are first multiplied pixel by pixel and then added to obtain an input feature extraction map.
[0076] Next, the input feature extraction map passes through the double-branch feature extraction module to obtain two preliminary astigmatism enhancement images.
[0077] Finally, the input feature extraction map and the two preliminary astigmatism enhancement images are merged and input into the multi-feature fusion module to obtain an astigmatism enhancement image.
[0078] Such asFigure 4 As shown below:
[0079] In step 3, the backlight enhancement network Expand-Net adopts a convolutional network with dual-branch convolution; the convolutional network with dual-branch convolution contains two branches; the first branch includes four 3x3 convolutional layers; the second branch includes seven 3x3 convolutional layers; both branches have 64 channels.
[0080] In step 3.1, the backlight image is processed by the first branch to obtain a detailed feature image, and the image size of the detailed image remains unchanged.
[0081] In step 3.2, the backlight image is gradually downsampled by the second branch to obtain a global feature image, and the size of the global feature image is reduced to 1x1.
[0082] In step 3.3, the detailed feature image and the global feature image are fused and input into a 3x3 convolutional layer, and the number of channels in the convolutional layer is reduced from 128 to 3 to obtain a backlight enhanced image.
[0083] The design of the convolutional network with dual-branch convolution effectively reduces the ghosting phenomenon that may appear in the restored image by avoiding the use of upsampling operations, and improves the quality of the backlight enhanced image.
[0084] As Figure 5 shown below:
[0085] In step 4, the detection process of the PP-YOLOE-PLUS-Tiny model is as follows:
[0086] The PP-YOLOE-PLUS-Tiny model includes an object detection module, an object tracking module, and an extended function module.
[0087] First, the object detection module accurately identifies and classifies a single-frame enhanced image, and classifies the single-frame enhanced image as a pedestrian or a vehicle.
[0088] Then, the object tracking module saves the classified image of the object detection module and correlates the information between the upper and lower frames of the classified image of the object detection module, so as to realize continuous tracking of the same object.
[0089] Next, the extended function module detects the attributes of pedestrians or vehicles; pedestrian attributes include gender, clothing, and age characteristics, vehicle attributes include vehicle characteristic information and vehicle violation detection, and vehicle characteristic information includes vehicle type, color, and brand; the extended function module realizes the detection of vehicle characteristic information by segmenting and extracting text information from license plate images, and the extended function module detects vehicle violations by combining vehicle characteristic information and road traffic markings to determine whether there are vehicle violation behaviors; finally, a road condition monitoring image is obtained.
[0090] In step 4, the scheduling mode of the PP-YOLOE-PLUS-Tiny model is as follows:
[0091] The calling mode of PaddleDetection in the PP-YOLOE-PLUS-Tiny model is refactored from command-line driven to modular; the modular scheduling mode enables the road condition monitoring system to dynamically select and call the corresponding functional modules according to the input type and target category during operation, thus enhancing the flexibility and adaptability of the road condition monitoring system; the input type is image or video, and the target category is pedestrian or vehicle.
[0092] The PP-YOLOE-PLUS-Tiny model is a lightweight object detection model optimized based on the PP-YOLOE architecture and is suitable for edge computing platforms. It has multi-task adaptation capabilities and can flexibly turn on and off functional modules such as pedestrian attribute recognition, vehicle attribute recognition, and license plate recognition according to different requirements. In the traffic flow monitoring scenario, the pedestrian attribute recognition function can be turned off to focus on vehicle detection and license plate recognition; in the urban road monitoring scenario, the pedestrian and vehicle detection functions can be turned on simultaneously to achieve comprehensive traffic monitoring. The function configuration of the PP-YOLOE-PLUS-Tiny model enables the road condition monitoring system to dynamically adjust the detection tasks according to different application scenarios and output corresponding traffic detection images, improving the practicality and adaptability of the system; in the improved PP-YOLOE-PLUS-Tiny model with the scheduling mode, the target segmentation function uses built-in multi-functional models such as pedestrian attributes, vehicle attributes, and license plate recognition for identification.
[0093] The above specific embodiments are used to explain and illustrate the present invention, rather than to limit the present invention. Any modifications and changes made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A method for monitoring road conditions under complex lighting based on a lightweight enhanced network, characterized in that, The road condition monitoring method includes the following steps: Step 1, Use a USB camera to collect road condition videos and transmit the collected road condition video stream to the edge computing platform in real time through the USB transmission protocol; the edge computing platform includes a lightweight enhancement network and a PP-YOLOE-PLUS-Tiny model; the lightweight enhancement network includes a low-light enhancement network PLED-Net, a astigmatism enhancement network TSD-Net, and a backlight enhancement network ExpandNet; Step 2, Initialize the parameters of the lightweight enhancement network and the PP-YOLOE-PLUS-Tiny model: When the road condition monitoring system is started, the parameters are uniformly and automatically initialized and loaded into the road condition monitoring system at one time; the lightweight enhancement network and the PP-YOLOE-PLUS-Tiny model remain in an active state during the operation of the road condition monitoring system until all weights are released when the road condition monitoring system exits; Step 3, The lightweight enhancement network dynamically selects the low-light enhancement network PLED-Net, the astigmatism enhancement network TSD-Net, and the backlight enhancement network ExpandNet according to the light environment for image enhancement processing to obtain a single-frame enhanced image; The low-light enhancement network PLED-Net is based on the Retinex theory. The process of enhancing the low-light image is successively five stages: enhancement, denoising, enhancement, denoising, and enhancement; the low-light enhancement network achieves high-quality image enhancement through progressive enhancement and adaptive denoising; the low-light enhancement network PLED-Net includes PMSRIEB and DEANRB; The image enhancement process of the astigmatism enhancement network TSD-Net is as follows: First, the astigmatic image is subjected to feature extraction through a multi-scale attention extraction module to obtain an astigmatic feature map; Then, the astigmatic image and the astigmatic feature image are first multiplied pixel by pixel and then added to obtain an input feature extraction map; Next, the input feature extraction map passes through a double-branch feature extraction module to obtain two preliminary astigmatic enhancement images; Finally, the input feature extraction map and the two preliminary astigmatic enhancement images are merged and input into a multi-feature fusion module to obtain an astigmatic enhancement image; The backlight enhancement network ExpandNet uses a convolutional network with double-branch convolution; the convolutional network with double-branch convolution contains two branches; the first branch includes four 3x3 convolutional layers; the second branch includes seven 3x3 convolutional layers; both branches have 64 channels; Step 3.1, The backlight image is processed by the first branch to obtain a detail feature image, and the image size of the detail image remains unchanged; Step 3.2, The backlight image is gradually downsampled by the second branch to obtain a global feature image, and the size of the global feature image is reduced to 1x1; Step 3.3, The detail feature image and the global feature image are fused and input into a 3x3 convolutional layer, and the number of channels of the convolutional layer is reduced from 128 to 3 to obtain a backlight enhancement image; Step 4, Input the single-frame enhanced image into the PP-YOLOE-PLUS-Tiny model for detection to obtain a road condition monitoring image; Step 5, batch-combine the road condition monitoring images into a video stream for visual display; batch-combining means combining 5 images into a batch.
2. The road condition monitoring method according to claim 1, wherein In step 3, the single-frame enhanced image includes a low-light enhanced image, a defocused light enhanced image, a backlight enhanced image, or a fusion enhanced image; the fusion enhanced image includes a low-light defocused light enhanced image, a low-light backlight enhanced image, a defocused light backlight enhanced image, or a low-light defocused light backlight enhanced image.
3. The road condition monitoring method according to claim 1, wherein In step 3, the process of the low-light enhanced image is as follows: In the enhancement stage, the PMSRIEB module is adopted; the PMSRIEB module sequentially includes a multi-scale feature extraction layer, a channel attention module, a color interaction module, a global-local illumination estimation layer, and a spatial attention module; the multi-scale feature extraction layer includes four convolutional layers, and the four convolutional layers are in parallel, and the convolutional kernels of the four convolutional layers are 1x1, 3x3, 5x5, and 7x7 respectively; the color interaction module is a 3x3 grouped convolutional layer; the global-local illumination estimation layer includes two convolutional layers, and the two convolutional layers are in parallel, and the convolutional kernels of the two convolutional layers are 3x3 and 7x7 respectively; the low-light image first passes through the multi-scale feature extraction layer to extract the illumination feature, and then the channel attention module applies attention to the illumination feature; then the color interaction module enhances the correlation of the three RGB channels; Then, the global-local illumination estimation layer performs illumination estimation from the global scale and the local scale, and the spatial attention module applies different enhancement degrees according to different regions to obtain a low-light output image; finally, the low-light output image is added to the low-light image residual to obtain a preliminary enhanced image; In the denoising stage, the DEANRB module is adopted; The DEANRB module includes a differential convolution, a noise estimation module, and a spatial attention module; the low-light image and the preliminary low-light enhanced image are simultaneously used as the inputs of the differential convolution, and the differential convolution extracts the detail features; The preliminary low-light enhanced image passes through the noise estimation module to obtain a noise estimation image; the illumination component feature map obtained by dividing the low-light image by the preliminary low-light enhanced image is used to calculate the SNR map, the SNR map obtains a noise spatial attention feature map through the spatial attention module, and the noise attention feature map is multiplied by the estimated noise image to achieve adaptive denoising; finally, the noise estimation image is subtracted from the preliminary low-light enhanced image to obtain a preliminary denoised image, and the detail features are added to the preliminary denoised image to obtain a denoised image; The calculation process of the SNR map is as follows: First, perform median filtering on the illumination component feature map to obtain a filtered illumination component feature map; then, subtract the illumination component feature map from the filtered illumination component feature map and take the absolute value to obtain a noise image; then, divide the filtered illumination component feature map by the noise image to obtain the SNR map.
4. The road condition monitoring method according to claim 3, characterized in that The differential convolution includes central difference, horizontal difference, vertical difference, and alignment difference.
5. The road condition monitoring method according to claim 3, wherein, The noise estimation module sequentially includes 12 convolutional blocks; each convolutional block includes a 3x3 convolutional layer, a normalization layer, and a LeakyReLU layer; the convolutional layer, the normalization layer, and the LeakyReLU layer are sequentially connected in series.
6. The road condition monitoring method according to claim 1, wherein In step 3, the astigmatism enhancement network TSD-Net includes a multi-scale attention extraction module, a dual-branch feature extraction module, and a multi-feature fusion module; The multi-scale attention extraction module includes six convolutional units; among the six convolutional units, the first convolutional unit includes a maximum pooling, a 3x3 convolutional layer, and a LeakyReLU layer; the second convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the third convolutional unit includes three parallel 3x3 dilated convolutional layers with dilation rates of 5, 7, and 9 respectively; the fourth convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the fifth convolutional unit includes a 3x3 convolutional layer and a LeakyReLU layer; the sixth convolutional unit includes a 3x3 convolutional layer and a Sigmoid layer; The dual-branch feature extraction module includes two branches; the first branch adopts a UNet architecture, and the second branch adopts a DeseNet network; the first branch sequentially includes a basic feature extraction block, a downsampling block, a basic feature extraction block, a basic feature extraction block, a downsampling block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, an upsampling block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, an upsampling block, a basic feature extraction block, a basic feature extraction block, a basic feature extraction block, and a basic feature extraction block; The basic feature extraction block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer, the downsampling block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer, and the upsampling block is bilinear interpolation; The multi-feature fusion module sequentially includes an input block, 5 feature extraction blocks, and an output block; the input block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer; each feature extraction block includes two convolutional blocks, and each convolutional block includes a 3x3 convolutional layer, a LeakyReLU layer, and a residual block; the 3x3 convolutional layer and the LeakyReLU layer are connected in series and then connected to the residual block; the output block sequentially includes a 3x3 convolutional layer and a LeakyReLU layer.
7. The road condition monitoring method according to claim 1, wherein In step 4, the detection process of the PP-YOLOE-PLUS-Tiny model is as follows: The PP-YOLOE-PLUS-Tiny model includes an object detection module, an object tracking module, and an extended function module; First, the object detection module accurately identifies and classifies a single-frame enhanced image, classifying the single-frame enhanced image as a pedestrian or a vehicle; Then, the object tracking module saves the classified image of the object detection module and correlates the information between the upper and lower frames of the classified image of the object detection module, so as to realize continuous tracking of the same object; Next, the extended function module detects the attributes of pedestrians or vehicles; pedestrian attributes include gender, clothing, and age characteristics, and vehicle attributes include vehicle characteristic information and vehicle violation detection. Vehicle characteristic information includes vehicle type, color, and brand. The extended function module detects vehicle characteristic information by segmenting and extracting text information from license plate images, and detects vehicle violations in vehicle images by combining vehicle characteristic information with road traffic markings to determine whether a vehicle has committed a violation; finally, a road condition monitoring image is obtained. The scheduling mode of the PP-YOLOE-PLUS-Tiny model is as follows: The calling mode of PaddleDetection in the PP-YOLOE-PLUS-Tiny model is refactored from command-line driven to modular. The modular scheduling mode dynamically selects and calls the corresponding function modules according to the input type and target category, thereby enhancing the flexibility and adaptability of the road condition monitoring system. The input type is an image or video, and the target category is a pedestrian or vehicle.
8. A road condition monitoring system for performing the method according to any one of claims 1-7, characterized in that, The road condition monitoring system includes a camera and an edge computing platform; the edge computing platform includes a lightweight enhancement network and a PP-YOLOE-PLUS-Tiny model; the lightweight enhancement network includes a low-light enhancement network PLED-Net, a glare enhancement network TSD-Net, and a backlight enhancement network ExpandNet; the low-light enhancement network PLED-Net, the glare enhancement network TSD-Net, and the backlight enhancement network ExpandNet are three independent enhancement networks; the lightweight enhancement network operates independently according to the real-time road condition environment. When the road condition environment is in multiple complex lighting conditions, the three independent enhancement networks are combined to enhance the image simultaneously. The camera collects road condition videos and transmits the video stream to the edge computing platform in real time through the USB transmission protocol; the lightweight enhancement network splits and enhances the video stream, inputs the enhanced image into the PP-YOLOE-PLUS-Tiny model for target monitoring, and combines the monitoring results into a video stream for visual display.
Citation Information
Patent Citations
Low-illumination image classification method based on attention mechanism and capsule network
CN111950649A
Target detection method based on DFLLOD-Net under low-illumination superposition fog weather
CN118918035A