Tunnel electromechanical equipment recognition method and system based on Retinex-DCE-YOLOv5s

By adopting the Retinex-DCE light enhancement algorithm and the improved YOLOv5s network model in tunnel electromechanical equipment inspection, the problems of low patrol efficiency and insufficient detection accuracy in the existing technology are solved, and efficient and accurate tunnel electromechanical equipment identification and operation and maintenance management are achieved.

CN120047776BActive Publication Date: 2025-07-01ZHEJIANG SCI RES INST OF TRANSPORT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510950.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-01
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing tunnel electromechanical equipment inspection methods are inefficient, making it difficult to fully grasp the electromechanical equipment in the tunnel, especially in low-light environments to detect significantly.

Method used

The tunnel electromechanical equipment recognition method based on Retinex-DCE-YOLOv5s is adopted, and the image is brighter and brighter through the Retinex-DCE illumination enhancement algorithm is used to build an improved YOLOv5s network model, replacing the ordinary convolution in the backbone network as CoordConv convolution, and introducing the CBAM attention mechanism and multi-scale detection head.

Benefits of technology

It significantly improves detection efficiency and accuracy, and can accurately identify a variety of tunnel electromechanical equipment in low-light environments, achieve comprehensive inspection, reduce operation and maintenance difficulties, and improve operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047776B_ABST
    Figure CN120047776B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s. The identification method involved includes: S1. Obtaining tunnel electromechanical equipment images collected during vehicle patrol; S2. Using Retinex-DCE to perform brightness enhancement processing on the obtained equipment images; S3. Constructing a YOLOv5s network, replacing ordinary convolutions except the first layer of convolution in the backbone network of YOLOV5s with CoordConv convolutions, introducing a CBAM attention mechanism at the end of the backbone network, and adding detection heads to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; S4. Inputting the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; S5. Inputting the equipment images to be processed into the detection and recognition model for processing, and outputting the equipment images, the positions where the equipment images are located, and the categories of the equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tunnel identification, and particularly to a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s. Background Art

[0002] The tunnel electromechanical system is an important part of the tunnel, which can effectively ensure traffic safety and smoothness and reduce the occurrence of traffic accidents. Therefore, the normal operation of tunnel electromechanical equipment is crucial for the traffic safety of the tunnel. By controlling and maintaining the tunnel electromechanical equipment, the service life of the electromechanical equipment can be extended, the maintenance and replacement costs can be reduced, the operation efficiency of the tunnel can be improved, the normal operation of the equipment can be ensured, and safety accidents caused by equipment failures can be avoided.

[0003] Currently, in the inspection work of tunnel electromechanical facilities in China, the commonly used detection methods are manual inspection, vehicle inspection, and system inspection. Among them, the manual inspection and vehicle inspection methods are not only inefficient but also difficult to comprehensively understand the situation of tunnel electromechanical equipment.

[0004] The vehicle inspection content includes power supply and distribution facilities, tunnel lighting facilities, tunnel ventilation equipment, tunnel fire protection facilities, and monitoring and communication facilities. Vehicle inspection requires 2 inspectors (1 driver and 1 observer to check the tunnel electromechanical facilities and record them in a form). The internal environment of the tunnel is relatively complex and there are many electromechanical facilities. During the vehicle inspection process, inspectors often have problems such as missed inspection and misjudgment.

[0005] Object detection is greatly affected by imaging conditions and environmental lighting. Once in a low-light environment, the change of light will have a great impact on the detection effect of the model. The lightweight model with fewer parameters and weaker expression ability has poor generalization ability, resulting in a sharp decline in detection performance. Due to the complex environment in the tunnel, the images taken during the tunnel vehicle inspection are extremely prone to quality problems such as weak light and blurred edges, which further lose the features of the target and lead to a significant decline in detection performance.

[0006] In view of the above technical problems, the present invention proposes a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s in view of the defects of the prior art.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] The method for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s includes:

[0010] S1. Obtain the tunnel electromechanical equipment images collected during the vehicle inspection process;

[0011] S2. Use Retinex-DCE to perform brightness enhancement processing on the obtained equipment images;

[0012] S3. Construct the YOLOv5s network, replace the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions, introduce the CBAM attention mechanism at the end of the backbone network, and add detection heads to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model;

[0013] S4. Input the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model;

[0014] S5. Input the equipment images to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the equipment images, the positions where the equipment images are located, and the categories of the equipment.

[0015] Further, after the step S5, it further includes:

[0016] S6. Judge whether the equipment is normal according to the equipment images output by the detection and recognition model.

[0017] Further, the step S2 is specifically:

[0018] S21. Convert the RGB color space of the equipment images to the HSV color space;

[0019] S22. Use Gaussian filtering to process the V channel of the HSV color space to obtain the illumination component of the V channel;

[0020] S23. Obtain the reflection component of the V channel according to the principle of light reflection imaging:

[0021] S24. Correct the illumination component based on the ZeroDCE network to obtain the corrected illumination component;

[0022] S25. Combine the corrected illumination component with the reflection component to obtain an enhanced V channel;

[0023] S26. Obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space to an enhanced RGB color space.

[0024] Further, in step S3, replacing the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions specifically includes:

[0025] Adding two coordinate channels to the input feature map, respectively representing the x coordinate and y coordinate of each pixel; then performing ordinary convolution operations on the feature map after adding the coordinate channels.

[0026] Further, the CBAM attention mechanism in step S3 includes a channel attention module and a spatial attention module:

[0027] Processing the feature map through the channel attention module to generate a channel attention map; processing the feature map after being processed by the channel attention module through the spatial attention module to generate a spatial attention map; multiplying the channel attention map and the spatial attention map to generate the final output feature.

[0028] Further, adding a detection head to the Head layer of YOLOV5s in step S3 specifically includes:

[0029] Performing upsampling and downsampling fusion on the feature map to generate feature maps of different scales; performing object detection on the feature maps of different scales.

[0030] Further, the tunnel electromechanical equipment images collected during the vehicle patrol in step S1 are collected by a camera installed on the vehicle.

[0031] Correspondingly, a tunnel electromechanical equipment recognition system based on Retinex-DCE-YOLOv5s is also provided, including:

[0032] An acquisition module, used to acquire tunnel electromechanical equipment images collected during vehicle patrol;

[0033] A processing module, used to perform brightness enhancement processing on the acquired equipment images using Retinex-DCE;

[0034] A construction module, used to construct a YOLOv5s network, replace the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions, introduce the CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model;

[0035] A training module, used to input the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model;

[0036] An output module for inputting a device image to be processed into a detection and recognition model for processing, and the detection and recognition model outputs the device image, the location where the device image is located, and the category of the device.

[0037] Further, it also includes:

[0038] A judgment module for judging whether the device is normal according to the device image output by the detection and recognition model.

[0039] Further, the processing module is specifically:

[0040] A first conversion module for converting the RGB color space of the device image into the HSV color space;

[0041] A first component module for processing the V channel of the HSV color space using Gaussian filtering to obtain the illumination component of the V channel;

[0042] A second component module for obtaining the reflection component of the V channel according to the principle of light reflection imaging:

[0043] A correction module for correcting the illumination component based on the ZeroDCE network to obtain the corrected illumination component;

[0044] A combination module for combining the corrected illumination component with the reflection component to obtain an enhanced V channel;

[0045] A second conversion module for obtaining an enhanced HSV color space according to the enhanced V channel and converting the enhanced HSV color space into an enhanced RGB color space.

[0046] Compared with the prior art, the beneficial effects of the present invention are:

[0047] 1. The detection efficiency is significantly improved: By installing a video camera on the vehicle hood and combining with an improved YOLOv5s model, the rapid detection and recognition of tunnel electromechanical equipment are realized. Compared with the traditional manual inspection and vehicle inspection methods, the detection efficiency is greatly improved, and equipment abnormalities can be discovered more timely.

[0048] 2. The detection accuracy is greatly improved: The Retinex-DCE illumination enhancement algorithm is used to preprocess the acquired images, effectively improving problems such as low illumination and blurred edges in the tunnel and enhancing the saliency of electromechanical equipment. At the same time, the improved YOLOv5s model enhances the detection ability for targets of different scales by introducing technologies such as CoordConv, CBAM attention mechanism, and multi-scale detection heads, significantly improving the detection accuracy.

[0049] 3. Wider detection range: The improved YOLOv5s model can identify a variety of tunnel electromechanical devices, including power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc., covering all types of electromechanical devices in the tunnel and achieving comprehensive detection.

[0050] 4. Reduced operation and maintenance difficulty: The monitoring system can analyze the operating status of electromechanical devices in real time, detect faults in a timely manner and remind operation and maintenance personnel, reducing the operation and maintenance difficulty, improving the operation and maintenance efficiency, and ensuring the stable operation of tunnel electromechanical devices. Description of the Drawings

[0051] Figure 1 is the flow chart of the tunnel electromechanical device recognition method based on Retinex-DCE-YOLOv5s provided in the first embodiment;

[0052] Figure 2 is the schematic diagram of installing a video camera on the vehicle hood provided in the first embodiment;

[0053] Figure 3 is the flow chart of performing brightness enhancement processing on the device image provided in the first embodiment;

[0054] Figure 4 is the ZeroDCE network structure provided in the first embodiment;

[0055] Figure 5 is the structural diagram of the improved YOLOv5s network model provided in the first embodiment;

[0056] Figure 6 is the schematic explanatory diagram of the improved YOLOv5s network model provided in the first embodiment;

[0057] Figure 7 is the CoordConv structure diagram provided in the first embodiment;

[0058] Figure 8 is the CBAM structure diagram provided in the first embodiment;

[0059] Figure 9 is the structure diagram of the channel attention module provided in the first embodiment;

[0060] Figure 10 is the structure diagram of the spatial attention module provided in the first embodiment;

[0061] Figure 11 is the structure diagram of the detection head provided in the first embodiment. Detailed Embodiments

[0062] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0063] The object of the present invention is to provide a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s in view of the defects of the prior art.

[0064] Embodiment 1

[0065] This embodiment provides a method for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s, as Figure 1 shown, including:

[0066] S1. Obtain images of tunnel electromechanical equipment collected during vehicle patrol;

[0067] S2. Use Retinex-DCE to perform brightness enhancement processing on the obtained equipment images;

[0068] S3. Construct a YOLOv5s network, replace the ordinary convolutions except the first layer of convolution in the backbone network of YOLOV5s with CoordConv convolutions, introduce a CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model;

[0069] S4. Input the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model;

[0070] S5. Input the equipment images to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the equipment images, the positions where the equipment images are located, and the categories of the equipment.

[0071] This embodiment conducts a field investigation on a comprehensive management office of a certain highway tunnel and obtains the tunnel electromechanical equipment during vehicle patrol as shown in Table 1 below.

[0072] Table 1 Tunnel electromechanical facilities during vehicle patrol

[0073]

[0074] According to the investigation situation, it is found that the tunnel electromechanical equipment has the following characteristics:

[0075] (1) The illumination in the tunnel is dim, and most of the captured images are low-illumination images.

[0076] (2) The distribution of tunnel electromechanical equipment has certain rules. For example, related equipment such as box transformers and substation rings is distributed on both sides of the tunnel entrance and exit, located on the left and right sides of the acquired images; lighting facilities, ventilation facilities, in-tunnel monitoring equipment, in-tunnel information collection equipment, and in-tunnel publishing equipment are distributed on the tunnel ceiling, located at the top of the acquired images; in-tunnel power distribution cabinets and related equipment of tunnel fire protection facilities are often distributed on the tunnel walls, located on the left and right sides of the acquired images.

[0077] (3) The sizes of tunnel electromechanical facilities that require vehicle patrol inspection vary in the images. For example, related facilities such as box transformers and substation rings are large targets; related facilities such as ventilation facilities are medium targets; related facilities such as tunnel fire protection and tunnel publishing are small targets; related facilities such as lighting and tunnel information collection are extremely small targets.

[0078] In step S1, images of tunnel electromechanical equipment collected during vehicle patrol inspection are acquired.

[0079] Install a video camera on the vehicle hood. The lens height of the camera is 1.5 meters, and the angle between the lens axis and the road surface is 180°, that is, the lens axis is parallel to the road surface. To reduce the impact of vehicle vibration on the image quality of the camera, the camera uses a vehicle-mounted stable bracket, such as Figure 2 shown.

[0080] In this embodiment, relevant electromechanical equipment images in the tunnel are collected by the camera, and the collected electromechanical equipment images are uploaded to the system for subsequent processing.

[0081] In step S2, Retinex-DCE is used to perform brightness enhancement processing on the acquired equipment images.

[0082] According to the illumination-reflection imaging model, an image is formed by the light of the light source reaching the imaging unit after being reflected by the object surface, that is, an image is composed of the product of the illumination component irradiating the scene and the reflection component of the object surface.

[0083] Image processing Retinex algorithm: Based on color constancy, it is assumed that the image observed by the human eye can be decomposed into two parts: illumination and reflectivity. Illumination reflects the distribution of light in the scene, affects the brightness of the object surface, and is separated from the color of the object. Reflectivity represents the inherent property of the object color, that is, the reflection ability of the object to light of different wavelengths, and is independent of the illumination conditions in an ideal situation. Therefore, in this embodiment, it is assumed that the illumination component of the tunnel scene image exists in the low-frequency part of the image and its overall change is gentle; the reflection component exists in the high-frequency part of the image and its local change is strong.

[0084] The ZeroDCE network adjusts the illumination component by means of per-pixel curve adjustment. The illumination component is input into the ZeroDCE network, and the network outputs the adjustment parameters for each pixel of the illumination component. The adjustment parameters are used to adjust each pixel of the illumination component multiple times to obtain the enhanced illumination component.

[0085] According to the above content, the Retinex-DCE illumination enhancement algorithm proposed in this embodiment reasonably equalizes the pixel distribution of low-illumination images in the tunnel, adjusts the image brightness and contrast, improves the saliency of electromechanical equipment, and distinguishes it from the background.

[0086] As Figure 3 shown, step S2 specifically includes:

[0087] S21. Convert the RGB color space of the device image to the HSV color space;

[0088] Convert the obtained device image in the tunnel from the RGB color space to the HSV color space. Compared with the RGB color space, the hue H, saturation S, and brightness V in the HSV color space are independent of each other and are more suitable for illumination improvement. The brightness V is the degree of light and dark that the human eye feels and is related to the reflection of the object; that is, the brightness channel is composed of the illumination component and part of the reflection component. According to the illumination-reflection imaging principle, the mathematical model of the brightness channel image of the input image is:

[0089] V(x,y)=R(x,y)×L(x,y);

[0090] where R(x,y) represents the partial reflection image of the object in the brightness V channel, that is, the reflection component; L(x,y) represents the incident illumination of the input image in the brightness V channel, that is, the illumination component; V(x,y) represents the brightness channel of the input image.

[0091] S22. Process the V channel of the HSV color space using Gaussian filtering to obtain the illumination component of the V channel;

[0092] Perform multi-scale Gaussian filtering (in this embodiment, three-scale Gaussian filtering is used, and the standard deviations σ of its convolution kernels are 15, 80, and 130 respectively) to convolve the V channel image to obtain the incident light image, that is, obtain the illumination component, which is expressed as:

[0093] ;

[0094] where is the Gaussian filtering formula, c represents the size factor, λ represents the normalization constant, and ensures that the Gaussian function G(x,y) satisfies the normalization condition, .

[0095] S23. Obtain the reflection component of the V channel according to the principle of light reflection imaging:

[0096] After extracting the illumination component L(x, y), obtain the reflection component R(x, y) according to the principle of light reflection imaging, which is expressed as:

[0097] ;

[0098] S24. Correct the illumination component based on the ZeroDCE network to obtain the corrected illumination component;

[0099] To enhance the illumination, correct the illumination component L(x, y) using the ZeroDCE algorithm to obtain the corrected illumination component L1(x, y), specifically:

[0100] As Figure 4 shown is the ZeroDCE network structure. The network has 7 layers, each layer contains a number of 3×3 convolutional kernels, the convolutional stride is taken as 1, and the boundary padding size is taken as 1. Since the ReLU function has a high computational efficiency ratio and the stability of gradient propagation, which helps the rapid convergence of the network, the first 6 layers of the ZeroDCE network use the ReLU activation function. The last layer uses the Tanh activation function, which can effectively normalize the output of the network to the desired range. Input the L(x, y) illumination component image into the network. After multiple convolutions, the network outputs the curve adjustment parameters for each pixel point in the image. Each parameter has 8 channels, corresponding to the subsequent 8 iterative adjustment stages. After obtaining the curve adjustment parameters, use this parameter to iterate L(x, y), and the iteration formula is:

[0101] LE n (x,y)=LE n-1 (x,y)+A n (x,y)LE n-1 (x,y)(1-LE n-1 (x,y));

[0102] Among them, (x, y) represents the pixel point coordinates; A n (x,y) represents the curve adjustment parameter of the pixel point coordinates and the illuminance component, LE n (x,y) represents the enhanced image of the given input LE n-1 (x,y), and n represents the number of iterations.

[0103] LE0(x,y)=L(x,y);

[0104] L1(x,y)=LE8(x,y);

[0105] Among them, L1(x, y) represents the corrected illumination component; LE8(x, y) == LE7(x, y) + A8(x, y)LE7(x, y)(1 - LE7(x, y)).

[0106] S25. Combine the corrected illumination component with the reflection component to obtain an enhanced V channel;

[0107] Combine the corrected illumination component L1(x, y) with the reflection component R(x, y) to obtain an enhanced luminance channel V1(x, y), expressed as:

[0108] V1(x, y) = R(x, y) × L1(x, y)

[0109] Among them, V1(x, y) represents the enhanced luminance V channel.

[0110] S26. Obtain an enhanced HSV color space based on the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space to achieve brightness enhancement of low-light images in the tunnel.

[0111] In step S3, construct a YOLOv5s network, replace the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions, introduce a CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model.

[0112] As Figures 5 - 6 shown, and Figure 6 includes three modules: CBS module, C3 module, and SPPF module. In this embodiment, by replacing the ordinary convolutions in the YOLOV5s backbone except the first layer of convolution with CoordConv structures, the model can more easily learn and utilize the spatial layout and structural information in the image, better understand the position information of tunnel electromechanical equipment, and further improve the positioning accuracy of the model for tunnel electromechanical facilities; by introducing a CBAM mechanism at the end of the backbone network to update the feature map before fusion to ignore some irrelevant information and enhance the model's attention to the target area; by adding a detection head to the Head layer of YOLOV5s to change the three-scale detection of the model to four-scale detection, the detection performance of the model for tunnel electromechanical facilities of different scales is improved.

[0113] Due to the translation invariance of traditional convolutional operations, that is, they are not sensitive to the position of the target. However, in many tasks (such as object detection), the position information of the target is very important. In this embodiment, CoordConv enables the model to better learn the spatial layout and position information of the target by explicitly introducing position information, thereby improving the positioning accuracy of the target.

[0114] As Figure 7 shown, the CoordConv structure is an improved convolutional operation designed to enhance the model's perception of the target position by introducing position information. The CoordConv structure includes an addition part for coordinate channels and an ordinary convolutional operation part.

[0115] Addition part for coordinate channels (Concat operation): Based on the input feature map, two additional channels are added, representing the x - coordinate and y - coordinate of each pixel respectively. These two coordinate channels are fixed and represent the position information of the pixels in the feature map. After adding the coordinate channels, the number of channels of the feature map increases from the original C to C + 2.

[0116] Ordinary convolutional operation part: After adding the coordinate channels, the feature map is processed through traditional convolutional operations. When calculating the convolutional kernel, it not only learns the local information of the feature map but also utilizes the position information in the coordinate channels to enhance the perception ability of the target position.

[0117] The processing flow after the feature map is input into CoordConv is as follows:

[0118] The shape of the input feature map is H×W×C, where H is the height of the feature map; W is the width of the feature map; and C is the number of channels of the feature map.

[0119] Add two coordinate channels to the input feature map: The x - coordinate channel represents the horizontal position (lateral coordinate) of each pixel in the feature map; the y - coordinate channel represents the vertical position (longitudinal coordinate) of each pixel in the feature map; after adding the coordinate channels, the shape of the feature map becomes H×W×(C + 2).

[0120] Process the feature map with added coordinate channels using traditional convolutional operations; where the size and stride of the convolutional kernel are the same as those of ordinary convolution, but the number of input channels is C + 2, and the number of output channels is C′ (determined according to specific design). When calculating the convolutional kernel, it will learn both the local information of the feature map and the position information in the coordinate channels.

[0121] After the convolutional operation is completed, the shape of the output feature map is H′×W′×C′, where H′ and W′ represent the size of the convolutional feature map; C′ represents the number of channels after convolution.

[0122] The position added by CoordConv in this embodiment enables it to perform local operations on the image, increasing the correlation between local and global image information, thereby more accurately determining the target position. CoordConv also has the translational invariance of traditional convolution, allowing pixels at different positions of the image to share the same convolution kernel parameters during convolution operations, and thus more effectively learning the essential features of the image.

[0123] As Figure 8 shown, CBAM (Convolutional Block Attention Module) is a simple, lightweight, and efficient attention module for feed-forward convolutional neural networks. It is a lightweight attention mechanism that can improve the model's ability to focus on target features in both spatial and channel dimensions. CBAM can perform adaptive feature optimization, enabling the network to focus more on useful information, which is beneficial for improving the detection accuracy of small targets and can be seamlessly integrated into any CNN network. It achieves double refinement of the feature map by combining the channel attention module CAM (Channel Attention) and the spatial attention module SAM (Spatial Attention).

[0124] As Figure 9 shown, the channel attention module compresses the spatial dimension while keeping the channel dimension unchanged, enabling the model to pay more attention to the information in the image. Specifically:

[0125] Max pooling and average pooling: Perform max pooling (MaxPool) and average pooling (AvgPool) operations on the input feature map F to generate two 1×1 feature maps with different dimensions. These two operations can capture the global information of the feature map from different perspectives.

[0126] Shared multi-layer perceptron (MLP): Input the two 1×1 feature maps with different dimensions obtained from max pooling and average pooling into a shared multi-layer perceptron respectively to compress the parameters. The shared multi-layer perceptron outputs two feature maps. This MLP usually contains two fully connected layers for modeling channel features.

[0127] Summation and activation: Add the two feature maps output by the MLP and generate the channel attention weight M c (F) through the sigmoid function. The shared layer here is a multi-layer perceptron plus a hidden layer, and the activation size is R C / r×1×1 , where r represents the reduction ratio.

[0128] ;

[0129] where σ represents the sigmoid activation function; W0 ∈ R C / r×CWith W1 ∈ R C×C / r represents the weights of the multi-layer perceptron, which are shared by two inputs, and the weight W0 is obtained using the RELU function.

[0130] As Figure 10 shown, the spatial attention module compresses the channel dimension while keeping the spatial dimension unchanged, enabling the model to focus more on the location information of the target. Specifically:

[0131] Max pooling and average pooling: The feature map F’ processed by the channel attention module is further subjected to max pooling and average pooling operations, which compress it along the channels to obtain two feature maps of H×W×1.

[0132] Convolution operation: The two generated two-dimensional feature maps are subjected to a convolution operation to generate the spatial attention map M s . This convolution operation usually uses a 7×7 convolution kernel, which can capture the spatial context information in the feature map.

[0133] Activation and fusion: The feature map F’ processed by the sigmoid function in the channel attention module is activated, and F’ is multiplied pixel by pixel with the generated spatial attention map Ms to obtain a more refined feature map M s (F’).

[0134] ;

[0135] where σ represents the sigmoid activation function; f 7×7 represents a convolution operation with a convolution kernel size of 7×7; AvgPool() and MaxPool() represent average pooling and max pooling operations respectively.

[0136] Output feature: The feature map processed by the CBAM attention mechanism is output as the input of the SimSPPF module for subsequent multi-scale feature extraction.

[0137] As Figure 11 shown, the network structure of adding a detection head to the Head layer of YOLOv5s is as follows:

[0138] When using a convolutional neural network to extract image features, as the number of convolutional layers deepens, the resolution of the feature map gradually decreases, and the range of the original image that the local receptive field can perceive continuously expands. Moreover, the feature maps closer to the top layer tend to focus more on the global information of the image, so the deep features extracted by the deep neural network are very unfavorable for small target detection. The local receptive field of the shallow feature map is smaller, containing more location and detail information, which can make up for the deficiency of deep features to a certain extent. However, if only shallow features are used to detect targets, due to the lack of guidance from high-level semantic information, it is very easy to cause serious misdetection and missed detection phenomena.

[0139] In YOLOV5s, if the scale of the input image is 640×640×3, feature maps of 80×80, 40×40, and 20×20 are obtained through downsampling by 8 times, 16 times, and 32 times respectively. The network finally performs object detection on the feature maps at these three different scales. Among the feature maps at these three scales, the one with the smallest local receptive field is the feature map after 8 - fold downsampling. Mapping this feature map to the original input image, each grid corresponds to an original Figure 8 ×8 region. For targets with relatively small resolutions, the receptive field of the feature map obtained by 8 - fold downsampling is still relatively large, and it is easy to lose the positions and detailed information of some small targets.

[0140] Since YOLOV5s misses a large number of small targets with resolutions lower than 64 pixels, its performance in detecting extremely small targets is poor. To improve the situation of missed object detection, in this embodiment, the Head structure of YOLOV5s is optimized. Based on the original three - scale detection head, a detection head for detecting extremely small targets is added. By adding a small - scale detection head, the network can detect small targets with resolutions not lower than 16 pixels, effectively improving the ability of the network to detect small targets.

[0141] In this embodiment, a detection head with 4 - fold downsampling is added to the YOLOV5s model. The minimum target resolution that can be detected is 4×4, which is more powerful than the original network. A total of 12 different - scale prior boxes are generated by clustering, and targets of different scales can be predicted on the feature maps after 4 - fold, 8 - fold, 16 - fold, and 32 - fold downsampling, greatly improving the multi - scale object detection performance of the algorithm.

[0142] Since a small - scale detection head is added to the improved model, the feature fusion method at the Neck end has changed accordingly, but the overall still follows the FPN + PAN structure. Figure 11 In it, C2, C3, C4, and C5 respectively correspond to the feature maps after 4 - fold, 8 - fold, 16 - fold, and 32 - fold downsampling extracted by the backbone network. F3 and F4 are feature fusion layers, and P2, P3, P4, and P5 are detection layers.

[0143] The Neck network first performs upsampling fusion on the input feature maps, then downsampling fusion, and then predicts the fused feature maps separately. Specifically:

[0144] Upsampling fusion: Perform an upsampling operation on the C5 feature map input by the backbone network and fuse it with the C4 feature map to obtain the F4 feature map. After performing operations such as convolution on F4 and then upsampling, fuse it with C3 to obtain a finer - grained feature map of the F3 layer. Continue to upsample F3 to obtain a larger - scale feature map, then fuse it with the C2 feature map, perform extremely small target detection on the fused large - scale feature map, and finally output the detection result by the P2 detection layer.

[0145] Downsampling fusion: First, perform downsampling on the feature map of the P2 detection layer and fuse it with the F3 feature map at the same scale. Then, the detection results of small targets are output by the P3 detection layer. Continue with downsampling and fuse with F4 and C5 respectively to detect medium and large targets on the 16-fold and 32-fold downsampled feature maps. Finally, the prediction results are output by the P4 and P5 detection layers.

[0146] In step S4, the device image after brightness enhancement processing is input into the improved YOLOv5s network model for training to obtain a trained detection and recognition model.

[0147] According to the image dataset of tunnel electromechanical equipment obtained in step S1, ensure that the dataset contains images of various tunnel electromechanical equipment, such as power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc. And divide the obtained dataset into a training set and a validation set, usually divided according to the ratio of 80% training set and 20% validation set.

[0148] Set training parameters, including input image size (such as 640x640), batch size (such as 16 or 32), number of training epochs (such as 100 to 300 epochs), dataset path, and weight file (you can choose to start training from a pre-trained model or start from scratch).

[0149] Use the training set to train the improved YOLOv5s network model, and adjust the model parameters by optimizing the loss function (such as GIoULoss or Focal Loss) to improve the detection accuracy and robustness of the model.

[0150] During the training process, use data augmentation methods such as rotation, translation, scaling, and brightness adjustment to increase the sample size and improve the robustness and generalization ability of the model.

[0151] Use the validation set to validate the trained model and evaluate the performance metrics of the model, such as mAP (mean average precision), to ensure that the detection performance of the model meets the requirements.

[0152] According to the validation results, further optimize the model, including adjusting hyperparameters (such as learning rate, weight decay, etc.), to further improve the detection performance of the model.

[0153] Through the above training method, the performance and accuracy of the improved YOLOv5s network model in tunnel electromechanical equipment detection can be effectively improved, and a trained detection and recognition model can be obtained.

[0154] In step S5, the device image to be processed is input into the detection and recognition model for processing, and the detection and recognition model outputs the device image, the location where the device image is located, and the category of the device.

[0155] The position of the device image output by the detection and recognition model includes the coordinates of the device (the upper left and lower right coordinates of the bounding box) and the confidence of the device (the certainty of the model about the existence of the device).

[0156] The device categories output by the detection and recognition model include power supply and distribution facilities (such as box transformers, ring main units, distribution cabinets, etc.), tunnel lighting facilities (such as lighting devices), tunnel ventilation facilities (such as ventilation devices), tunnel fire protection facilities (such as fire cabinets, signs, etc.), and monitoring and communication facilities (such as monitoring devices, information collection devices, message boards, etc.).

[0157] Through the detection and recognition method of this embodiment, the inspection personnel can quickly understand the distribution and status of the electromechanical devices in the tunnel, thereby improving the inspection efficiency and accuracy.

[0158] After step S5, it further includes:

[0159] S6. Judge whether the device is normal according to the device image output by the detection and recognition model.

[0160] The inspection personnel judge whether the tunnel electromechanical device image output by the detection and recognition model is normal, which specifically includes:

[0161] Detect the appearance status of the device: Check whether there are obvious damages, deformations or abnormal conditions on the appearance of the device. For example, whether the lighting fixtures are intact and whether the ventilation devices are damaged.

[0162] Observe the operation status of the device: Manually observe the operation of the device. For example, whether the brightness of the lighting fixtures is normal and whether the ventilation devices are operating normally.

[0163] Check whether the functions of the device are normal: For information release devices, check whether the displayed content is timely and correct. For power supply devices, check whether the power supply is stable and whether transformers, switchgear, etc. are normal.

[0164] Technical index detection: Use professional instruments to measure the technical indexes of the device. For example, whether the grounding resistance is ≤4Ω and whether the insulation resistance is ≥50Ω.

[0165] Data transmission and communication function: Check whether the communication between the device and the control center is normal and whether there is no packet loss and out-of-step in data transmission.

[0166] Through the above method, the inspection personnel can comprehensively check the appearance, operation status, functions and technical indexes of the device by combining the device position and category detected by YOLOv5s, so as to judge whether the device is operating normally.

[0167] Compared with the prior art, the beneficial effects of the present invention are:

[0168] 1. Significantly improved detection efficiency: By installing a video camera on the vehicle hood and combining it with the improved YOLOv5s model, rapid detection and identification of tunnel electromechanical equipment are achieved. Compared with traditional manual inspection and vehicle inspection methods, the detection efficiency is greatly improved, and equipment abnormalities can be detected more promptly.

[0169] 2. Greatly enhanced detection accuracy: The Retinex-DCE light enhancement algorithm is used to preprocess the acquired images, effectively improving problems such as low light and blurred edges in the tunnel and enhancing the saliency of electromechanical equipment. At the same time, the improved YOLOv5s model enhances the detection ability for targets of different scales and significantly improves the detection accuracy by introducing technologies such as CoordConv, CBAM attention mechanism, and multi-scale detection heads.

[0170] 3. Wider detection range: The improved YOLOv5s model can identify a variety of tunnel electromechanical equipment, including power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc., covering all types of electromechanical equipment in the tunnel and achieving comprehensive detection.

[0171] 4. Reduced operation and maintenance difficulty: The monitoring system can perform real-time analysis on the operating status of electromechanical equipment, promptly detect faults and alert operation and maintenance personnel, reducing the operation and maintenance difficulty, improving the operation and maintenance efficiency, and ensuring the stable operation of tunnel electromechanical equipment.

[0172] Example 2

[0173] This example provides a tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s, including:

[0174] An acquisition module for acquiring images of tunnel electromechanical equipment collected during vehicle inspections;

[0175] A processing module for performing brightness enhancement processing on the acquired equipment images using Retinex-DCE;

[0176] A construction module for constructing a YOLOv5s network, replacing the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions, introducing a CBAM attention mechanism at the end of the backbone network, and adding detection heads to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model;

[0177] A training module for inputting the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and identification model;

[0178] An output module for inputting the device image to be processed into a detection and recognition model for processing, and the detection and recognition model outputs the device image, the position where the device image is located, and the category of the device.

[0179] Furthermore, it further includes:

[0180] A judgment module for judging whether the device is normal according to the device image output by the detection and recognition model.

[0181] Furthermore, the processing module is specifically:

[0182] A first conversion module for converting the RGB color space of the device image into the HSV color space;

[0183] A first component module for processing the V channel of the HSV color space using Gaussian filtering to obtain the illumination component of the V channel;

[0184] A second component module for obtaining the reflection component of the V channel according to the principle of light reflection imaging:

[0185] A correction module for correcting the illumination component based on the ZeroDCE network to obtain the corrected illumination component;

[0186] A combination module for combining the corrected illumination component and the reflection component to obtain the enhanced V channel;

[0187] A second conversion module for obtaining the enhanced HSV color space according to the enhanced V channel and converting the enhanced HSV color space into the enhanced RGB color space.

[0188] It should be noted that the tunnel electromechanical device recognition system based on Retinex-DCE-YOLOv5s provided in this embodiment is similar to that in Embodiment 1, and will not be elaborated here.

[0189] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s, characterized in that: include: S1. Obtain images of tunnel electromechanical equipment collected during vehicle inspection; S2. Use Retinex-DCE to perform brightness enhancement processing on the acquired device image; S3. Build a YOLOv5s network, replace the ordinary convolution except the first layer of convolution in the YOLOV5s backbone network with CoordConv convolution, introduce the CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; Replace the ordinary convolution except the first layer of convolution in the YOLOV5s backbone network with CoordConv convolution as follows: Add two coordinate channels to the input feature map, representing the x-coordinate and y-coordinate of each pixel respectively; then perform a normal convolution operation on the feature map after adding the coordinate channels; The CBAM attention mechanism includes a channel attention module and a spatial attention module: Process the feature map through the channel attention module to generate a channel attention map; process the feature map processed by the channel attention module through the spatial attention module to generate a spatial attention map; multiply the channel attention map and the spatial attention map to generate the final output feature; The specific steps for adding a detection head to the Head layer of YOLOV5s are as follows: Upsample and downsample the feature maps to generate feature maps of different scales; perform target detection on feature maps of different scales; S4. Input the device image after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; S5. Input the device image to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the device image and the location of the device image and the category of the device.

2. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: After step S5, the following steps are also included: S6. Determine whether the device is normal based on the device image output by the detection and recognition model.

3. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: The step S2 is specifically as follows: S21. Convert the RGB color space of the device image into the HSV color space; S22. Processing the V channel of the HSV color space using Gaussian filtering to obtain a light component of the V channel; S23. Obtain the reflection component of the V channel according to the principle of light reflection imaging: S24. Correcting the illumination component based on the ZeroDCE network to obtain a corrected illumination component; S25. Combining the corrected illumination component with the reflection component to obtain an enhanced V channel; S26. Obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space.

4. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: The tunnel electromechanical equipment images collected during the vehicle inspection in step S1 are collected by a camera installed on the vehicle.

5. Tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s, characterized by: include: An acquisition module is used to acquire images of tunnel electromechanical equipment collected during vehicle inspection; A processing module, used for performing brightness enhancement processing on the acquired device image by using Retinex-DCE; The construction module is used to build the YOLOv5s network. The ordinary convolution except the first layer of convolution in the YOLOV5s backbone network is replaced with CoordConv convolution. At the same time, the CBAM attention mechanism is introduced at the end of the backbone network, and a detection head is added to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model. Replace the ordinary convolution except the first layer of convolution in the YOLOV5s backbone network with CoordConv convolution as follows: Add two coordinate channels to the input feature map, representing the x-coordinate and y-coordinate of each pixel respectively; then perform a normal convolution operation on the feature map after adding the coordinate channels; The CBAM attention mechanism includes a channel attention module and a spatial attention module: Process the feature map through the channel attention module to generate a channel attention map; process the feature map processed by the channel attention module through the spatial attention module to generate a spatial attention map; multiply the channel attention map and the spatial attention map to generate the final output feature; The specific steps for adding a detection head to the Head layer of YOLOV5s are as follows: Upsample and downsample the feature maps to generate feature maps of different scales; perform target detection on feature maps of different scales; The training module is used to input the device image after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; The output module is used to input the device image to be processed into the detection and recognition model for processing. The detection and recognition model outputs the device image and the location of the device image and the category of the device.

6. The tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s according to claim 5 is characterized in that: Also includes: The judgment module is used to judge whether the device is normal according to the device image output by the detection and recognition model.

7. The tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s according to claim 5 is characterized in that: The processing module is specifically: A first conversion module, used for converting the RGB color space of the device image into the HSV color space; The first component module is used to process the V channel of the HSV color space by using Gaussian filtering to obtain the illumination component of the V channel; The second component module is used to obtain the reflection component of the V channel according to the principle of light reflection imaging: A correction module, used for correcting the illumination component based on the ZeroDCE network to obtain the corrected illumination component; A combining module, used for combining the corrected illumination component with the reflection component to obtain an enhanced V channel; The second conversion module is used to obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space.

Citation Information

Patent Citations

  • Coal flow foreign matter identification method for coal mine belt conveyor based on machine vision

    CN116665011A

  • Converter assembly automatic quality inspection method and system based on improved YOLOv8 detection algorithm

    CN118711038A