Tunnel electromechanical equipment identification method and system based on Retinex-DCE-YOLOv5s

Through the tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s, the problems of low inspection efficiency of tunnel electromechanical equipment and poor detection performance in low light environments in the prior art are solved, and efficient and accurate equipment identification and operation and maintenance management are achieved.

CN120047776AActive Publication Date: 2025-05-27ZHEJIANG SCI RES INST OF TRANSPORT

Patent Information

Application Number
CN202510510950.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing tunnel electromechanical equipment inspection methods are inefficient, making it difficult to fully grasp the electromechanical equipment in the tunnel, especially in low-light environments to detect significantly.

Method used

The tunnel electromechanical equipment recognition method based on Retinex-DCE-YOLOv5s is adopted, and the brightness enhancement processing is performed through the Retinex-DCE algorithm, and an improved YOLOv5s network model is constructed, the ordinary convolution in the backbone network is replaced as CoordConv convolution, the CBAM attention mechanism is introduced, and the detection head is added to the head layer.

Benefits of technology

It significantly improves detection efficiency and accuracy, and can accurately identify a variety of electromechanical equipment in the tunnel in a low-light environment, covering power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc., reducing the difficulty of operation and maintenance and improving operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047776A_ABST
    Figure CN120047776A_ABST
Patent Text Reader

Abstract

The invention discloses a tunnel electromechanical equipment identification method and system based on Retinex-DCE-YOLOv5s, and the method comprises the steps: S1, obtaining a tunnel electromechanical equipment image collected in a vehicle inspection process; s2, carrying out brightness enhancement processing on the obtained equipment image by adopting Retinex-DCE; s3, constructing a YOLOv5s network, replacing common convolution except the first layer of convolution in the backbone network of the YOLOV5s with CoordConv convolution, introducing a CBAM attention mechanism at the tail end of the backbone network, and adding a detection head at a Head layer of the YOLOV5s to obtain an improved YOLOv5s network model; s4, inputting the equipment image subjected to brightness enhancement processing into an improved YOLOv5s network model for training to obtain a trained detection and recognition model; and S5, inputting a to-be-processed equipment image into the detection and identification model for processing, and outputting the equipment image, the position of the equipment image and the type of the equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tunnel identification, and particularly to a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s. Background Art

[0002] The tunnel electromechanical system is an important part of the tunnel, which can effectively ensure traffic safety and smoothness and reduce the occurrence of traffic accidents. Therefore, the normal operation of tunnel electromechanical equipment is crucial for the traffic safety of the tunnel. By controlling and maintaining tunnel electromechanical equipment, the service life of electromechanical equipment can be extended, the maintenance and replacement costs can be reduced, the operation efficiency of the tunnel can be improved, the normal operation of the equipment can be ensured, and safety accidents caused by equipment failures can be avoided.

[0003] Currently, in the inspection work of tunnel electromechanical facilities in China, the commonly used detection methods are manual inspection, vehicle inspection, and system inspection. Among them, the manual inspection and vehicle inspection methods are not only inefficient but also difficult to comprehensively understand the situation of tunnel electromechanical equipment.

[0004] The content of vehicle inspection includes power supply and distribution facilities, tunnel lighting facilities, tunnel ventilation equipment, tunnel fire protection facilities, and monitoring and communication facilities. Vehicle inspection requires 2 inspectors (1 drives the vehicle, and 1 observes the tunnel electromechanical facilities and records them in a form). The internal environment of the tunnel is relatively complex, and there are many electromechanical facilities. During vehicle inspection, inspectors often have problems such as missed inspection and misjudgment.

[0005] Object detection is greatly affected by imaging conditions and environmental lighting. Once in a low-light environment, the change of light will have a great impact on the detection effect of the model. Lightweight models with fewer parameters and weaker expression ability have poor generalization ability, resulting in a sharp decline in detection performance. Due to the complex environment in the tunnel, the images taken during tunnel vehicle inspection are extremely prone to quality problems such as weak light and blurred edges, further losing the features of the target and resulting in a significant decline in detection performance.

[0006] In view of the above technical problems, the present invention proposes a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s in view of the defects of the prior art.

[0008] To achieve the above purpose, the present invention adopts the following technical solutions: A method for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s includes: S1. Obtain the images of tunnel electromechanical equipment collected during vehicle inspection tours; S2. Use Retinex-DCE to perform brightness enhancement processing on the obtained equipment images; S3. Construct a YOLOv5s network, replace the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions, introduce a CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; S4. Input the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; S5. Input the equipment images to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the equipment images, the positions where the equipment images are located, and the categories of the equipment.

[0009] Further, after the step S5, it further includes: S6. Judge whether the equipment is normal according to the equipment images output by the detection and recognition model.

[0010] Further, the step S2 is specifically: S21. Convert the RGB color space of the equipment images to the HSV color space; S22. Use Gaussian filtering to process the V channel of the HSV color space to obtain the illumination component of the V channel; S23. Obtain the reflection component of the V channel according to the principle of light reflection imaging: S24. Correct the illumination component based on the ZeroDCE network to obtain the corrected illumination component; S25. Combine the corrected illumination component with the reflection component to obtain an enhanced V channel; S26. Obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space to an enhanced RGB color space.

[0011] Further, the replacement of the ordinary convolutions in the backbone network of YOLOV5s except the first layer of convolution with CoordConv convolutions in the step S3 is specifically: Add two coordinate channels to the input feature map, representing the x coordinate and y coordinate of each pixel respectively; then perform ordinary convolution operations on the feature map after adding the coordinate channels.

[0012] Further, the CBAM attention mechanism in the step S3 includes a channel attention module and a spatial attention module: Process the feature map through a channel attention module to generate a channel attention map; process the feature map processed by the channel attention module through a spatial attention module to generate a spatial attention map; multiply the channel attention map and the spatial attention map to generate the final output feature.

[0013] Further, specifically adding a detection head in the Head layer of YOLOV5s in step S3 is as follows: Perform upsampling and downsampling fusion on the feature map to generate feature maps of different scales; perform object detection on the feature maps of different scales.

[0014] Further, in step S1, the tunnel electromechanical equipment images collected during vehicle inspection are collected by a camera installed on the vehicle.

[0015] Correspondingly, a tunnel electromechanical equipment recognition system based on Retinex-DCE-YOLOv5s is also provided, including: An acquisition module, used to acquire tunnel electromechanical equipment images collected during vehicle inspection; A processing module, used to perform brightness enhancement processing on the acquired equipment images using Retinex-DCE; A construction module, used to construct a YOLOv5s network, replace the ordinary convolutions in the backbone network of YOLOV5s except for the first convolutional layer with CoordConv convolutions, introduce a CBAM attention mechanism at the end of the backbone network, and add a detection head in the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; A training module, used to input the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; An output module, used to input the equipment images to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the equipment images, the positions where the equipment images are located, and the categories of the equipment.

[0016] Further, it also includes: A judgment module, used to judge whether the equipment is normal according to the equipment images output by the detection and recognition model.

[0017] Further, the processing module is specifically: A first conversion module, used to convert the RGB color space of the equipment image into the HSV color space; A first component module, used to process the V channel of the HSV color space using Gaussian filtering to obtain the illumination component of the V channel; A second component module, used to obtain the reflection component of the V channel according to the principle of light reflection imaging: Correction module, which is used to correct the illumination component based on the ZeroDCE network to obtain the corrected illumination component; Combination module, which is used to combine the corrected illumination component with the reflection component to obtain the enhanced V channel; Second conversion module, which is used to obtain the enhanced HSV color space according to the enhanced V channel and convert the enhanced HSV color space into the enhanced RGB color space.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The detection efficiency is significantly improved: By installing a video camera on the vehicle hood and combining with the improved YOLOv5s model, the rapid detection and identification of tunnel electromechanical equipment are realized. Compared with the traditional manual inspection and vehicle inspection methods, the detection efficiency is greatly improved, and equipment abnormalities can be found more timely.

[0019] 2. The detection accuracy is greatly improved: The Retinex-DCE illumination enhancement algorithm is used to preprocess the acquired images, effectively improving problems such as low illumination and blurred edges in the tunnel and enhancing the saliency of electromechanical equipment. At the same time, the improved YOLOv5s model enhances the detection ability for targets of different scales by introducing technologies such as CoordConv, CBAM attention mechanism, and multi-scale detection heads, significantly improving the detection accuracy.

[0020] 3. The detection range is wider: The improved YOLOv5s model can identify a variety of tunnel electromechanical equipment, including power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc., covering all types of electromechanical equipment in the tunnel and realizing comprehensive detection.

[0021] 4. The operation and maintenance difficulty is reduced: The monitoring system can analyze the operating status of electromechanical equipment in real time, discover faults in a timely manner and remind the operation and maintenance personnel, reducing the operation and maintenance difficulty, improving the operation and maintenance efficiency, and ensuring the stable operation of tunnel electromechanical equipment. Description of the Drawings

[0022] Figure 1 It is a flowchart of the method for identifying tunnel electromechanical equipment based on Retinex-DCE-YOLOv5s provided in the first embodiment; Figure 2 It is a schematic diagram of installing a video camera on the vehicle hood provided in the first embodiment; Figure 3 It is a flowchart of the brightness enhancement processing of the equipment image provided in the first embodiment; Figure 4 It is the ZeroDCE network structure provided in the first embodiment; Figure 5It is the structural diagram of the improved YOLOv5s network model provided in the first embodiment; Figure 6 It is the schematic diagram for the description of the improved YOLOv5s network model provided in the first embodiment; Figure 7 It is the structural diagram of CoordConv provided in the first embodiment; Figure 8 It is the structural diagram of CBAM provided in the first embodiment; Figure 9 It is the structural diagram of the channel attention module provided in the first embodiment; Figure 10 It is the structural diagram of the spatial attention module provided in the first embodiment; Figure 11 It is the structural diagram of the detection head provided in the first embodiment. Specific implementation mode

[0023] The following uses specific specific examples to illustrate the implementation mode of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation modes. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0024] The purpose of the present invention is to provide a tunnel electromechanical equipment recognition method and system based on Retinex-DCE-YOLOv5s in view of the defects of the prior art.

[0025] The first embodiment

[0026] This embodiment provides a tunnel electromechanical equipment recognition method based on Retinex-DCE-YOLOv5s, as Figure 1 shown, including: S1. Obtain tunnel electromechanical equipment images collected during vehicle patrol; S2. Use Retinex-DCE to perform brightness enhancement processing on the obtained equipment images; S3. Construct a YOLOv5s network, replace the ordinary convolution except the first layer convolution in the backbone network of YOLOV5s with CoordConv convolution, introduce the CBAM attention mechanism at the end of the backbone network at the same time, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; S4. Input the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; S5. Input the device image to be processed into the detection and recognition model for processing. The detection and recognition model outputs the device image, the location where the device image is located, and the category of the device.

[0027] In this embodiment, a field investigation was conducted on a comprehensive management office of a highway tunnel, and the vehicle inspection of tunnel electromechanical equipment is shown in Table 1 below.

[0028] Table 1 Vehicle inspection of tunnel electromechanical facilities

[0029] According to the investigation, the following characteristics of tunnel electromechanical equipment were found: (1) The lighting in the tunnel is dim, and most of the captured images are low-illuminance images.

[0030] (2) The distribution of tunnel electromechanical equipment has a certain pattern. For example, related equipment such as box transformers and ring network stations is distributed on both sides of the tunnel entrance and exit, located on the left and right sides of the acquired image; lighting facilities, ventilation facilities, in-tunnel monitoring equipment, in-tunnel information collection equipment, and in-tunnel publishing equipment are distributed on the tunnel ceiling, located at the top of the acquired image; in-tunnel distribution cabinets and related equipment of tunnel fire protection facilities are often distributed on the tunnel wall, located on the left and right sides of the acquired image.

[0031] (3) The sizes of tunnel electromechanical facilities that need to be inspected by vehicle vary in the image. For example, related facilities such as box transformers and ring network stations are large targets; related facilities such as ventilation facilities are medium targets; related facilities such as tunnel fire protection and tunnel publishing are small targets; related facilities such as lighting and tunnel information collection are extremely small targets.

[0032] In step S1, acquire the tunnel electromechanical equipment images collected during the vehicle inspection process.

[0033] Install a video camera on the vehicle hood. The height of the camera lens is 1.5 meters, and the angle between the lens axis and the road surface is 180°, that is, the lens axis is parallel to the road surface. To reduce the impact of vehicle vibration on the image quality of the camera, the camera uses a vehicle-mounted stable bracket, such as Figure 2 shown.

[0034] In this embodiment, relevant electromechanical equipment images in the tunnel are collected by the camera, and the collected electromechanical equipment images are uploaded to the system for subsequent processing.

[0035] In step S2, use Retinex-DCE to perform brightness enhancement processing on the acquired device images.

[0036] According to the light-reflection imaging model, an image is formed by the light of the light source reaching the imaging unit after being reflected by the object surface, that is, an image is composed of the product of the illumination component of the light irradiating the scene and the reflection component of the object surface.

[0037] Image processing Retinex algorithm: Based on color constancy, it is assumed that the image observed by the human eye can be decomposed into two parts: illumination and reflectance. The illumination reflects the distribution of light in the scene, affects the brightness of the object surface, and is separated from the color of the object. The reflectance represents the inherent property of the object color, that is, the reflection ability of the object to light of different wavelengths, and is independent of the lighting conditions in ideal cases. Therefore, this embodiment assumes that the illumination component of the tunnel scene image exists in the low-frequency part of the image and its overall change is gentle; the reflection component exists in the high-frequency part of the image and its local change is strong.

[0038] The ZeroDCE network adjusts the illumination component by means of pixel-by-pixel curve adjustment. The illumination component is input into the ZeroDCE network, and the network outputs the adjustment parameters for each illumination component pixel. The adjustment parameters are used to adjust each illumination component pixel multiple times to obtain the enhanced illumination component.

[0039] According to the above content, the Retinex-DCE illumination enhancement algorithm proposed in this embodiment reasonably equalizes the pixel distribution of low-illumination images in the tunnel, adjusts the image brightness and contrast, improves the saliency of electromechanical equipment, and distinguishes it from the background.

[0040] As Figure 3 shown, step S2 specifically includes: S21. Convert the RGB color space of the device image to the HSV color space; Convert the obtained device image in the tunnel from the RGB color space to the HSV color space. Compared with the RGB color space, the hue H, saturation S, and brightness V in the HSV color space are independent of each other and are more suitable for lighting improvement. The brightness V is the degree of light and dark that the human eye feels and is related to the reflection of the object; that is, the brightness channel is composed of the illumination component and part of the reflection component. According to the illumination-reflection imaging principle, the mathematical model of the brightness channel image of the input image is: V(x,y)=R(x,y)×L(x,y); where, R(x,y) represents the partial reflection image of the object in the brightness V channel, that is, the reflection component; L(x,y) represents the incident illumination of the input image in the brightness V channel, that is, the illumination component; V(x,y) represents the brightness channel of the input image.

[0041] S22. Process the V channel of the HSV color space using Gaussian filtering to obtain the illumination component of the V channel; Perform multiple-degree Gaussian filtering (in this embodiment, three-scale Gaussian filtering is used, and the standard deviations σ of its convolution kernels are 15, 80, and 130 respectively) to convolve the V channel image to obtain the incident light image, that is, obtain the illumination component, which is expressed as: ; Among them, is the Gaussian filtering formula, c represents the size factor, and λ represents the normalization constant to ensure that the Gaussian function G(x, y) satisfies the normalization condition. .

[0042] S23. Obtain the reflection component of the V channel according to the principle of light reflection imaging: After extracting the light component L(x, y), obtain the reflection component R(x, y) according to the principle of light reflection imaging, which is expressed as: ; S24. Correct the light component based on the ZeroDCE network to obtain the corrected light component; To enhance the light, correct the light component L(x, y) using the ZeroDCE algorithm to obtain the corrected light component L 1 (x, y), specifically: As Figure 4 shown is the ZeroDCE network structure. The network has 7 layers, each layer contains a number of 3×3 convolutional kernels, the convolution stride is 1, and the boundary padding size is 1. Since the ReLU function has a high computational efficiency ratio and stability of gradient propagation, which helps the network to converge quickly, the first 6 layers of the ZeroDCE network use the ReLU activation function. The last layer uses the Tanh activation function, which can effectively normalize the output of the network to the desired range. Input the L(x, y) light component image into the network. After multiple convolutions, the network outputs the curve adjustment parameters for each pixel point in the image. Each parameter has 8 channels, corresponding to the subsequent 8 iterative adjustment stages. After obtaining the curve adjustment parameters, use this parameter to iterate L(x, y), and the iterative formula is: LE n (x, y) = LE n-1 (x, y) + A n (x, y)LE n-1 (x, y)(1 - LE n-1 (x, y)); Among them, (x, y) represents the pixel point coordinates; A n (x, y) represents the curve adjustment parameter of the pixel point coordinates and the illuminance component, and LE n (x, y) represents the enhanced image of the given input LE n-1 (x, y), and n represents the number of iterations.

[0043] LE 0 (x, y) = L(x, y); L 1 (x, y) = LE 8 (x, y); Among them, L 1 (x, y) represents the corrected illumination component; LE 8 (x, y) == LE 7 (x, y) + A 8 (x, y)LE 7 (x, y)(1 - LE 7 (x, y)).

[0044] S25. Combine the corrected illumination component with the reflection component to obtain an enhanced V channel; Combine the corrected illumination component L 1 (x, y) with the reflection component R(x, y) to obtain an enhanced luminance channel V 1 (x, y), expressed as: V 1 (x, y) = R(x, y) × L 1 (x, y) Among them, V 1 (x, y) represents the enhanced luminance V channel.

[0045] S26. Obtain an enhanced HSV color space based on the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space to achieve brightness enhancement of low-light images in the tunnel.

[0046] In step S3, construct a YOLOv5s network, replace the ordinary convolutions except the first-layer convolution in the backbone network of YOLOV5s with CoordConv convolutions, introduce a CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model.

[0047] As Figure 5 - Figure 6 shown, and Figure 6 includes three modules: CBS module, C3 module, and SPPF module. In this embodiment, by replacing the ordinary convolutions except the first-layer convolution in the YOLOV5s backbone with CoordConv structures, the model can more easily learn and utilize the spatial layout and structural information in the image, better understand the position information of tunnel electromechanical equipment, and further improve the positioning accuracy of the model for tunnel electromechanical facilities; by introducing a CBAM mechanism at the end of the backbone network, the feature map before fusion is updated to ignore some irrelevant information and enhance the model's attention to the target area; by adding a detection head to the Head layer of YOLOV5s, the three-scale detection of the model is changed to four-scale detection, improving the detection performance of the model for tunnel electromechanical facilities of different scales.

[0048] Due to the translational invariance of traditional convolutional operations, that is, they are insensitive to the position of the target. However, in many tasks (such as object detection), the position information of the target is very important. In this embodiment, CoordConv enables the model to better learn the spatial layout and position information of the target by explicitly introducing position information, thereby improving the positioning accuracy of the target.

[0049] As Figure 7 shown, the CoordConv structure is an improved convolutional operation designed to enhance the model's perception of the target position by introducing position information. The CoordConv structure includes an addition part of coordinate channels and a part of ordinary convolutional operations.

[0050] Addition part of coordinate channels (Concat operation): Based on the input feature map, two additional channels are added, representing the x - coordinate and y - coordinate of each pixel respectively. These two coordinate channels are fixed and represent the position information of the pixels in the feature map. After adding the coordinate channels, the number of channels of the feature map increases from the original C to C + 2.

[0051] Ordinary convolutional operation part: After adding the coordinate channels, the feature map is processed through traditional convolutional operations. When calculating the convolutional kernel, it not only learns the local information of the feature map but also utilizes the position information in the coordinate channels to enhance the perception ability of the target position.

[0052] The processing flow after the feature map is input into CoordConv is as follows: The shape of the input feature map is H×W×C, where H is the height of the feature map; W is the width of the feature map; C is the number of channels of the feature map.

[0053] Add two coordinate channels to the input feature map: The x - coordinate channel represents the horizontal position (lateral coordinate) of each pixel in the feature map; the y - coordinate channel represents the vertical position (longitudinal coordinate) of each pixel in the feature map; after adding the coordinate channels, the shape of the feature map becomes H×W×(C + 2).

[0054] Process the feature map with added coordinate channels using traditional convolutional operations; where the size and stride of the convolutional kernel are the same as those of ordinary convolution, but the number of input channels is C + 2, and the number of output channels is C′ (determined according to specific design). When calculating the convolutional kernel, it will learn the local information of the feature map and the position information in the coordinate channels simultaneously.

[0055] After the convolutional operation, the shape of the output feature map is H′×W′×C′, where H′ and W′ represent the size of the feature map after convolution; C′ represents the number of channels after convolution.

[0056] The position added by CoordConv in this embodiment can enable it to perform local operations on the image, increasing the relevance between local information and global information of the image, thereby more accurately determining the target position. CoordConv also has the translational invariance of traditional convolution, enabling pixels at different positions of the image to share the same convolution kernel parameters during convolution operations, and thus more effectively learning the essential features of the image.

[0057] As Figure 8 shown, CBAM (Convolutional Block Attention Module) is a simple, lightweight, and efficient attention module for feed-forward convolutional neural networks. It is a lightweight attention mechanism that can improve the model's ability to focus on target features in both spatial and channel dimensions. CBAM can perform adaptive feature optimization, enabling the network to focus more on useful information, thereby facilitating the improvement of small target detection accuracy and can be seamlessly integrated into any CNN network. It achieves double refinement of the feature map by combining the channel attention module CAM (Channel Attention) and the spatial attention module SAM (Spatial Attention).

[0058] As Figure 9 shown, the channel attention module compresses the spatial dimension while keeping the channel dimension unchanged, allowing the model to pay more attention to the information in the image. Specifically: Max pooling and average pooling: Perform max pooling (MaxPool) and average pooling (AvgPool) operations on the input feature map F to generate two 1×1 feature maps with different dimensions. These two operations can capture the global information of the feature map from different perspectives.

[0059] Shared multi-layer perceptron (MLP): Input the two 1×1 feature maps with different dimensions obtained by max pooling and average pooling into a shared multi-layer perceptron respectively, compress the parameters, and the shared multi-layer perceptron outputs two feature maps. This MLP usually contains two fully connected layers for modeling channel features.

[0060] Summation and activation: Add the two feature maps output by the MLP and generate the channel attention weight M c (F) through the sigmoid function. The shared layer here is a multi-layer perceptron plus a hidden layer, and the activation size is R C / r×1×1 , where r represents the reduction ratio.

[0061] ; where σ represents the sigmoid activation function; W 0 ∈R C / r×C and W 1 ∈RC×C / r Represents the multilayer perceptron weights, which are shared by two inputs, and the weight W is obtained using the RELU function 0 .

[0062] like Figure 10 As shown in the figure, the spatial attention module keeps the spatial dimension unchanged and compresses the channel dimension, allowing the model to pay more attention to the location information of the target. Specifically: Max pooling and average pooling: The feature map F' processed by the channel attention module is again subjected to max pooling and average pooling operations, which are compressed on the channel to obtain two H×W×1 feature maps.

[0063] Convolution operation: Perform convolution operation on the two generated two-dimensional feature maps to generate a spatial attention map M s This convolution operation usually uses a 7×7 convolution kernel, which can capture the spatial context information in the feature map.

[0064] Activation and fusion: The feature map F' processed by the sigmoid function channel attention module is activated, and F' is multiplied pixel by pixel with the generated spatial attention map Ms to obtain a more refined feature map M s (F').

[0065] ; Among them, σ represents the sigmoid activation function; f 7×7 Represents a convolution operation with a convolution kernel size of 7×7; AvgPool() and MaxPool() represent average pooling and maximum pooling operations respectively.

[0066] Output features: The feature map processed by the CBAM attention mechanism is output as the input of the SimSPPF module for subsequent multi-scale feature extraction.

[0067] like Figure 11 As shown in the figure, the network structure of adding a detection head to the Head layer of YOLOv5s is as follows: When using convolutional neural networks to extract image features, as the number of convolutional layers increases, the resolution of the feature map gradually decreases, and the range of the original image that the local receptive field can perceive continues to expand. Moreover, the feature map closer to the top layer tends to focus on the global information of the image, so the deep features extracted by deep neural networks are very unfavorable for small target detection. The local receptive field of shallow feature maps is smaller and contains more position and detail information, which can make up for the lack of deep features to a certain extent. However, if only shallow features are used to detect targets, due to the lack of guidance from high-level semantic information, it is very easy to cause serious false detection and missed detection.

[0068] In YOLOV5s, if the scale of the input image is 640×640×3, feature maps of 80×80, 40×40, and 20×20 are obtained through 8x, 16x, and 32x downsampling respectively. The network finally performs object detection on the feature maps at these three different scales. Among the feature maps at these three scales, the one with the smallest local receptive field is the 8x downsampled feature map. That is, when mapping this feature map to the original input image, each grid corresponds to the original Figure 8 ×8 area. For targets with relatively small resolutions, the receptive field of the feature map obtained by 8x downsampling is still relatively large, and it is easy to lose the position and detailed information of some small targets.

[0069] Since YOLOV5s misses a large number of small targets with resolutions lower than 64 pixels, its performance for extremely small targets is poor. To improve the situation of missed object detection, in this embodiment, the Head structure of YOLOV5s is optimized. Based on the original three-scale detection head, a detection head for extremely small target detection is added. By adding a small-scale detection head, the network can detect small targets with resolutions not lower than 16 pixels, effectively improving the ability of the network to detect small targets.

[0070] In this embodiment, a detection head with 4x downsampling is added to the YOLOV5s model. The smallest target resolution that can be detected is 4×4, which is more powerful than the original network. A total of 12 different-scale prior boxes are generated by clustering, and targets of different scales can be predicted on the feature maps with 4x, 8x, 16x, and 32x downsampling, greatly improving the multi-scale object detection performance of the algorithm.

[0071] Since a small-scale detection head is added to the improved model, the feature fusion method at the Neck end has changed accordingly, but the overall still follows the FPN+PAN structure. Figure 11 In it, C2, C3, C4, and C5 respectively correspond to the feature maps with 4x, 8x, 16x, and 32x downsampling extracted by the backbone network. F3 and F4 are feature fusion layers, and P2, P3, P4, and P5 are detection layers.

[0072] The Neck network first performs upsampling fusion on the input feature maps, then downsampling fusion, and then predicts the fused feature maps separately. Specifically: Upsampling fusion: Perform an upsampling operation on the C5 feature map input by the backbone network and fuse it with the C4 feature map to obtain the F4 feature map. After performing convolution and other operations on F4 and then upsampling it, fuse it with C3 to obtain a finer-grained feature map of the F3 layer. Continue to upsample F3 to obtain a larger-scale feature map, then fuse it with the C2 feature map, perform extremely small target detection on the fused large-scale feature map, and finally output the detection result by the P2 detection layer.

[0073] Downsampling fusion: First, perform downsampling on the feature map of the P2 detection layer, fuse it with the F3 feature map at the same scale, and then the P3 detection layer outputs the detection results of small targets. Continue with downsampling, fuse with F4 and C5 respectively, detect medium and large targets on the 16x and 32x downsampled feature maps, and finally the P4 and P5 detection layers output the prediction results.

[0074] In step S4, input the device image after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model.

[0075] According to the image dataset of tunnel electromechanical equipment obtained in step S1, ensure that the dataset contains images of various tunnel electromechanical equipment, such as power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc. And divide the obtained dataset into a training set and a validation set, usually divided according to the ratio of 80% training set and 20% validation set.

[0076] Set training parameters, including input image size (such as 640x640), batch size (such as 16 or 32), number of training epochs (such as 100 to 300 epochs), dataset path, and weight file (you can choose to start training from a pre-trained model or start from scratch).

[0077] Use the training set to train the improved YOLOv5s network model, and adjust the model parameters by optimizing the loss function (such as GIoULoss or Focal Loss) to improve the detection accuracy and robustness of the model.

[0078] During training, use data augmentation methods such as rotation, translation, scaling, and brightness adjustment to increase the sample size and improve the robustness and generalization ability of the model.

[0079] Use the validation set to validate the trained model, evaluate the performance metrics of the model, such as mAP (mean average precision), to ensure that the detection performance of the model meets the requirements.

[0080] According to the validation results, further optimize the model, including adjusting hyperparameters (such as learning rate, weight decay, etc.) to further improve the detection performance of the model.

[0081] Through the above training method, the performance and accuracy of the improved YOLOv5s network model in tunnel electromechanical equipment detection can be effectively improved, and a trained detection and recognition model can be obtained.

[0082] In step S5, input the device image to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the device image, the location where the device image is located, and the category of the device.

[0083] The position of the device image output by the detection and recognition model includes the coordinates of the device (the upper left and lower right coordinates of the bounding box) and the confidence of the device (the certainty of the model about the existence of the device).

[0084] The device categories output by the detection and recognition model include power supply and distribution facilities (such as box transformers, ring main units, distribution cabinets, etc.), tunnel lighting facilities (such as lighting devices), tunnel ventilation facilities (such as ventilation devices), tunnel fire protection facilities (such as fire cabinets, signs, etc.), and monitoring and communication facilities (such as monitoring devices, information collection devices, message boards, etc.).

[0085] Through the detection and recognition method of this embodiment, the inspection personnel can quickly understand the distribution and status of the electromechanical devices in the tunnel, thereby improving the inspection efficiency and accuracy.

[0086] After step S5, it further includes: S6. Judge whether the device is normal according to the device image output by the detection and recognition model.

[0087] The inspection personnel judge whether the tunnel electromechanical device image output by the detection and recognition model is normal, which specifically includes: Detect the appearance status of the device: Check whether there are obvious damages, deformations or abnormal conditions on the appearance of the device. For example, whether the lighting fixtures are intact and whether the ventilation devices are damaged.

[0088] Observe the operating status of the device: Manually observe the operating conditions of the device. For example, whether the brightness of the lighting fixtures is normal and whether the ventilation devices are operating normally.

[0089] Check whether the functions of the device are normal: For information release devices, check whether the displayed content is timely and correct. For power supply devices, check whether the power supply is stable and whether transformers, switchgear, etc. are normal.

[0090] Technical index detection: Use professional instruments to measure the technical indexes of the device. For example, whether the grounding resistance is ≤4Ω and whether the insulation resistance is ≥50Ω.

[0091] Data transmission and communication function: Check whether the communication between the device and the control center is normal and whether there is no packet loss and out-of-step in data transmission.

[0092] Through the above method, the inspection personnel can comprehensively check the appearance, operating status, functions and technical indexes of the device by combining the device position and category detected by YOLOv5s, so as to judge whether the device is operating normally.

[0093] Compared with the prior art, the beneficial effects of the present invention are: 1. Significantly improved detection efficiency: By installing a video camera on the vehicle hood and combining it with an improved YOLOv5s model, rapid detection and identification of tunnel electromechanical equipment are achieved. Compared with traditional manual inspection and vehicle inspection methods, the detection efficiency is greatly improved, and equipment abnormalities can be detected more promptly.

[0094] 2. Greatly enhanced detection accuracy: The Retinex-DCE illumination enhancement algorithm is used to preprocess the acquired images, effectively improving problems such as low illumination and blurred edges in the tunnel and enhancing the saliency of electromechanical equipment. At the same time, the improved YOLOv5s model enhances the detection ability for targets of different scales and significantly improves the detection accuracy by introducing technologies such as CoordConv, CBAM attention mechanism, and multi-scale detection heads.

[0095] 3. Wider detection range: The improved YOLOv5s model can identify a variety of tunnel electromechanical equipment, including power supply and distribution facilities, lighting facilities, ventilation equipment, fire protection facilities, etc., covering all types of electromechanical equipment in the tunnel and achieving comprehensive detection.

[0096] 4. Reduced operation and maintenance difficulty: The monitoring system can perform real-time analysis on the operating status of electromechanical equipment, promptly detect faults and alert operation and maintenance personnel, reducing the operation and maintenance difficulty, improving the operation and maintenance efficiency, and ensuring the stable operation of tunnel electromechanical equipment.

[0097] Embodiment 2 This embodiment provides a tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s, including: An acquisition module for acquiring images of tunnel electromechanical equipment collected during vehicle inspections; A processing module for performing brightness enhancement processing on the acquired equipment images using Retinex-DCE; A construction module for constructing a YOLOv5s network, replacing the ordinary convolutions except the first layer convolution in the backbone network of YOLOV5s with CoordConv convolutions, introducing a CBAM attention mechanism at the end of the backbone network, and adding detection heads to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; A training module for inputting the equipment images after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and identification model; An output module for inputting the equipment images to be processed into the detection and identification model for processing, and the detection and identification model outputs the equipment images, the positions where the equipment images are located, and the categories of the equipment.

[0098] Furthermore, it further includes: A judgment module, configured to judge whether the device is normal according to the device image output by the detection and recognition model.

[0099] Furthermore, the processing module is specifically: A first conversion module, configured to convert the RGB color space of the device image into the HSV color space; A first component module, configured to process the V channel of the HSV color space by using Gaussian filtering to obtain the illumination component of the V channel; A second component module, configured to obtain the reflection component of the V channel according to the principle of light reflection imaging: A correction module, configured to correct the illumination component based on the ZeroDCE network to obtain the corrected illumination component; A combination module, configured to combine the corrected illumination component with the reflection component to obtain an enhanced V channel; A second conversion module, configured to obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space.

[0100] It should be noted that the tunnel electromechanical device recognition system based on Retinex-DCE-YOLOv5s provided in this embodiment is similar to that in Embodiment 1, and will not be elaborated here.

[0101] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described here, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s, characterized in that: include: S1. Obtain images of tunnel electromechanical equipment collected during vehicle inspection; S2. Use Retinex-DCE to perform brightness enhancement processing on the acquired device image; S3. Build a YOLOv5s network, replace the ordinary convolution except the first layer of convolution in the YOLOV5s backbone network with CoordConv convolution, introduce the CBAM attention mechanism at the end of the backbone network, and add a detection head to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model; S4. Input the device image after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; S5. Input the device image to be processed into the detection and recognition model for processing, and the detection and recognition model outputs the device image and the location of the device image and the category of the device.

2. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: After step S5, the following steps are also included: S6. Determine whether the device is normal based on the device image output by the detection and recognition model.

3. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: The step S2 is specifically as follows: S21. Convert the RGB color space of the device image into the HSV color space; S22. Processing the V channel of the HSV color space using Gaussian filtering to obtain a light component of the V channel; S23. Obtain the reflection component of the V channel according to the principle of light reflection imaging: S24. Correcting the illumination component based on the ZeroDCE network to obtain a corrected illumination component; S25. Combining the corrected illumination component with the reflection component to obtain an enhanced V channel; S26. Obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space.

4. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: In step S3, the ordinary convolution except the first convolution layer in the backbone network of YOLOV5s is replaced with CoordConv convolution as follows: Add two coordinate channels to the input feature map, representing the x-coordinate and y-coordinate of each pixel respectively; then perform a normal convolution operation on the feature map after adding the coordinate channels.

5. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: The CBAM attention mechanism in step S3 includes a channel attention module and a spatial attention module: The feature map is processed by the channel attention module to generate a channel attention map; the feature map processed by the channel attention module is processed by the spatial attention module to generate a spatial attention map; the channel attention map and the spatial attention map are multiplied to generate the final output feature.

6. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: In step S3, adding a detection head to the Head layer of YOLOV5s is specifically as follows: The feature maps are upsampled and downsampled to generate feature maps of different scales; and target detection is performed on feature maps of different scales.

7. The tunnel electromechanical equipment identification method based on Retinex-DCE-YOLOv5s according to claim 1 is characterized in that: The tunnel electromechanical equipment images collected during the vehicle inspection in step S1 are collected by a camera installed on the vehicle.

8. Tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s, characterized by: include: An acquisition module is used to acquire images of tunnel electromechanical equipment collected during vehicle inspection; A processing module, used for performing brightness enhancement processing on the acquired device image by using Retinex-DCE; The construction module is used to build the YOLOv5s network. The ordinary convolution except the first layer of convolution in the YOLOV5s backbone network is replaced with CoordConv convolution. At the same time, the CBAM attention mechanism is introduced at the end of the backbone network, and a detection head is added to the Head layer of YOLOV5s to obtain an improved YOLOv5s network model. The training module is used to input the device image after brightness enhancement processing into the improved YOLOv5s network model for training to obtain a trained detection and recognition model; The output module is used to input the device image to be processed into the detection and recognition model for processing. The detection and recognition model outputs the device image and the location of the device image and the category of the device.

9. The tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s according to claim 8 is characterized in that: Also includes: The judgment module is used to judge whether the device is normal according to the device image output by the detection and recognition model.

10. The tunnel electromechanical equipment identification system based on Retinex-DCE-YOLOv5s according to claim 8, characterized in that: The processing module is specifically: A first conversion module, used for converting the RGB color space of the device image into the HSV color space; The first component module is used to process the V channel of the HSV color space by using Gaussian filtering to obtain the illumination component of the V channel; The second component module is used to obtain the reflection component of the V channel according to the principle of light reflection imaging: A correction module, used for correcting the illumination component based on the ZeroDCE network to obtain the corrected illumination component; A combining module, used for combining the corrected illumination component with the reflection component to obtain an enhanced V channel; The second conversion module is used to obtain an enhanced HSV color space according to the enhanced V channel, and convert the enhanced HSV color space into an enhanced RGB color space.

Citation Information

Patent Citations

  • Coal flow foreign matter identification method for coal mine belt conveyor based on machine vision

    CN116665011A

  • Converter assembly automatic quality inspection method and system based on improved YOLOv8 detection algorithm

    CN118711038A

  • Aluminum profile defect detection and identification method based on YOLOv7-ESC

    CN119205614A

  • Improved YOLOv8 coal flow visual detection method

    CN119741594A

  • High-altitude electric power operation violation identification method

    CN119785431A

Cited By

  • Image recognition method and device, storage medium and electronic equipment

    CN120656000A

  • Intelligent invoice identification method and system based on AI technology

    CN121545161A

  • An invoice intelligent identification method and system based on AI technology

    CN121545161B