Tunnel water leakage detection method based on neural network

By improving the Unet neural network model and combining multi-source detection data, the accuracy problem of tunnel leakage detection in complex scenarios was solved, and high-precision and real-time leakage detection was achieved.

CN120726327AActive Publication Date: 2025-09-30BEIJING JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510849582.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-30
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing tunnel water leakage detection method based on visible light images has reduced detection accuracy under uneven lighting, occlusion or complex scenes, and cannot effectively capture temperature anomalies and structural deformation in the leakage area.

Method used

An improved Unet neural network model is used to combine infrared thermal imaging data, visible light image data and laser ranging data. The standard convolution is replaced by depthwise separable convolution and padded convolution, redundant cross-layer connections are deleted, and multimodal input samples are generated to detect water leakage areas.

Benefits of technology

It significantly improves the detection accuracy in low-light and occlusion scenarios, improves the segmentation accuracy when the leakage edge is blurred, meets the timeliness requirements of tunnel inspections, and realizes real-time output of leakage location, scope and severity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726327A_ABST
    Figure CN120726327A_ABST
Patent Text Reader

Abstract

The invention provides a tunnel water leakage detection method based on a neural network, and relates to the technical field of tunnel water leakage detection, and the method comprises the following steps: replacing a standard convolution layer in a Unet network with a depth separable convolution layer, employing a filling convolution operation, and deleting a redundant cross-layer connection structure in the Unet network, and obtaining an improved Unet neural network model; multi-source detection data in the tunnel are collected and preprocessed, an input sample set is obtained, and the input sample set comprises a temperature distribution image, a visible light image and a depth image; wherein the sizes of the images in the input sample set are consistent; inputting the input sample set into the improved Unet neural network model to generate a binary segmentation mask, the size of which is consistent with that of the visible light image; and comprehensively analyzing the binarized segmentation mask to obtain the actual position, the leakage range and the leakage severity of the water leakage area in the tunnel. The method can improve the detection precision while ensuring the timeliness of tunnel inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of tunnel water leakage detection, and in particular to a tunnel water leakage detection method based on a neural network. Background Art

[0002] Tunnel water leakage, a major form of tunnel defect, not only affects tunnel structural safety but can also damage electrical equipment and even disrupt the smooth operation of the entire transportation system. Therefore, regular tunnel water leakage inspections are crucial for extending tunnel service life and preventing catastrophic risks.

[0003] In the prior art, for example, the invention patent with application number 202411283164.X proposes a method for detecting tunnel water leakage. This method is based on an improved EfficientNet-B0 network and constructs an optimized semantic segmentation model for detection by adaptively adjusting the backbone network, using the scSE attention mechanism, adjusting the convolution kernel size, and other methods. However, this method mainly relies on visible light image data and cannot effectively capture temperature anomalies and structural deformations in the leakage area. The detection accuracy is greatly reduced in tunnels with uneven lighting, occlusion, or complex scenes. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present application aims to provide a tunnel water leakage detection method based on a neural network, which can improve the detection accuracy while ensuring the timeliness of detection; the detection method comprises the following steps: The standard convolutional layer in the Unet network is replaced with a depth-wise separable convolutional layer, a padded convolution operation is adopted, and the redundant cross-layer connection structure in the Unet network is deleted to obtain an improved Unet neural network model; Collecting multi-source detection data in the tunnel, wherein the multi-source detection data includes infrared thermal imaging data, visible light image data, and laser ranging data; Preprocessing the multi-source detection data to obtain an input sample set, wherein the input sample set includes a temperature distribution image, a visible light image, and a depth image; wherein each image in the input sample set has a consistent size; The temperature distribution image, the visible light image, and the depth image in the input sample set are fused according to the channel dimension, and input into the improved Unet neural network model to generate a binary segmentation mask whose size is consistent with the visible light image; A comprehensive analysis is performed on the binary segmentation mask to obtain the actual location of the water leakage area in the tunnel, the leakage range and the leakage severity.

[0005] According to the technical solution provided by this application, preprocessing the multi-source detection data to obtain an input sample set includes the following steps: Separating the color features of the leaked water by color space conversion on the visible light image data to obtain a visible light image; Converting the infrared thermal imaging data into a two-dimensional temperature distribution image having the same size as the visible light image; The laser ranging data is used to generate a two-dimensional depth image having the same size as the visible light image through an interpolation algorithm.

[0006] According to the technical solution provided in this application, the temperature distribution image, the visible light image, and the depth image in the input sample set are fused according to the channel dimension and input into the improved Unet neural network model, including the following steps: Fusing the input sample set according to the channel dimension to obtain a multi-channel input tensor; The multi-channel input tensor is input to the improved Unet neural network model.

[0007] According to the technical solution provided by this application, after generating the binary segmentation mask, the following steps are further included: The target pixel coordinates of the water leakage area are marked on the binary segmentation mask.

[0008] According to the technical solution provided in this application, the comprehensive analysis of the binary segmentation mask includes the following steps: Based on the original three-dimensional data of the laser ranging data, the target pixel coordinates of the binary segmentation mask are mapped to the three-dimensional space of the tunnel to obtain the actual position of the water leakage area in the tunnel; Counting the number of pixels in the water leakage area in the binary segmentation mask, and calculating the actual area of ​​the leakage range in combination with the scale parameter of the depth image; The gradient information of the temperature distribution image and the color depth of the visible light image are fused to evaluate the severity of the leakage.

[0009] According to the technical solution provided by this application, the process of separating the color features of the leaked water by color space conversion of the visible light image data includes the following steps: After converting the visible light image data from RGB space to HSV space, a dual-color domain mask of mask1 and mask2 is established through linear normalization processing, where mask1 corresponds to the mask of the purple and blue transition area, and mask2 corresponds to the mask of the blue area; A purple area mask is obtained by performing a differential calculation on mask1-mask2, and the purple area mask is used as a water leakage area mask. A binary annotation is generated after a morphological closing operation.

[0010] According to the technical solution provided in this application, the training process of the improved Unet neural network model includes: The binary cross entropy loss function and RMSprop optimizer are used for parameter update; A dynamic learning rate adjustment strategy is introduced during the training process. When the Dice coefficient of the validation set does not increase for three consecutive epochs, the learning rate is reduced to 1 / 5 of the current value. A multimodal training set is constructed by mixing the labeled data of the HSV space and the labeled data of the SAM model, and the input ratio of the two types of labeled data is 3:1.

[0011] According to the technical solution provided by this application, a mobile inspection robot is used to collect multi-source detection data in a tunnel, including: The 120° equal-angle layout of the disc servo manipulator arm enables spatial alignment of the infrared thermal imager, visible light camera, and laser rangefinder; Establish a synchronous acquisition mechanism based on hardware trigger signals. When the laser rangefinder detects that the target distance is less than a first preset threshold, it synchronously triggers the infrared thermal imager and the visible light camera to expose; In the data preprocessing stage, the infrared thermal imaging data, visible light image data and laser ranging data are unified into the same spatial coordinate system through affine transformation.

[0012] According to the technical solution provided in this application, the method of fusing the gradient information of the temperature distribution image with the color depth of the visible light image to evaluate the severity of leakage includes the following steps: Construct temperature gradient matrix and visible light color depth matrix; defining a leakage index, wherein the leakage index is related to the actual area of ​​the leakage range; If the leakage index is greater than or equal to the second preset threshold, it is determined to be a serious leakage; if the leakage index is greater than the third preset threshold and less than the second preset threshold, it is determined to be a moderate leakage; if the leakage index is less than or equal to the third preset threshold, it is determined to be a slight leakage.

[0013] According to the technical solution provided by the present application, the method further includes constructing a physical feature constraint function of water leakage, wherein the physical feature constraint function includes a temperature-depth correlation constraint and a color-depth consistency constraint; After generating the binary segmentation mask, the following steps are included: Dividing the binary segmentation mask into 10×10 pixel grids, and simultaneously calculating a neural network confidence score and a physical constraint compliance score for each grid, wherein the neural network confidence score is obtained by outputting the improved Unet neural network model, and the physical constraint compliance score is obtained by the physical characteristic constraint function of the leakage water; The comprehensive analysis of the binary segmentation mask comprises the following steps: If the deviation between the neural network confidence score and the physical constraint compliance score is less than or equal to a fourth preset threshold, a comprehensive analysis is performed on the binary segmentation mask.

[0014] Compared with the existing technology, the beneficial effects of the present application are: the present application replaces the standard convolution in the Unet network with depthwise separable convolution, and deletes redundant cross-layer connections, significantly reducing the number of model parameters and computational complexity, while using padded convolution to maintain the resolution of the feature map, thereby improving the detection accuracy of small target leakage areas; in addition, by combining infrared thermal imaging (reflecting temperature anomalies), visible light images (surface texture) and laser ranging data (depth information), multimodal input samples are generated through adaptive splicing, effectively enhancing the model's feature representation ability for leakage areas, and improving the detection accuracy in low-light and occluded scenes. Especially when the leakage edge is blurred, the segmentation accuracy is significantly better than the single data source method. The improved Unet network has improved inference speed, and combined with binary mask comprehensive analysis technology, it can output the leakage location, range and severity in real time to meet the timeliness requirements of tunnel inspections. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A schematic diagram of the steps of the neural network-based tunnel water leakage detection method provided in this application; Figure 2 Schematic diagram of the structure of the improved Unet neural network model provided in this application; The text annotations in the figure represent: 1. Input sample set; 2. Depthwise separable convolution + RELU activation function layer; 3. Maximum pooling layer; 4. Upsampling convolution layer; 5. Standard convolution + RELU activation function layer; 6. Skip connection layer; 7. Convolutional layer; 8. Binarized segmentation mask. DETAILED DESCRIPTION

[0016] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.

[0017] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0018] Example 1 As mentioned in the background technology, in order to solve the problems in the prior art, this application proposes a tunnel leakage detection method based on neural network, such as Figure 1 As shown, the following steps are included: S1. Replace the standard convolutional layer in the Unet network with a depth-wise separable convolutional layer, adopt a padded convolution operation, and delete the redundant cross-layer connection structure in the Unet network to obtain an improved Unet neural network model; Specifically, if Figure 2 As shown in the figure, the two multi-channel 3×3 standard convolutions in each layer of the backbone of the traditional Unet network structure are replaced with 3×3 depth-separable convolutions + RELU activation function layer 2, thereby reducing the number of model parameters. The non-padded convolution used in Unet is replaced with padded convolution (3×3 standard convolution + RELU activation function layer 5) to ensure that the size of the feature map does not change during the convolution operation. Max pooling layer 3, Upsampling convolution layer 4, skip connection layer 6, Convolutional layer 7 ensures that the output mask (binarized segmentation mask 8) matches the size of the input image. Cross-layer connections between the first and second layers, which have little correlation with the feature maps, are removed to further reduce the model's computational complexity and improve inference speed. The number of output channels in the improved Unet neural network model is the number of input mask categories, that is, the number of categories in the multi-source detection data. In this solution, this number can be set to three.

[0019] Specifically, the original structure of the traditional Unet model has an input image size of 572×572, and after processing, the output feature map size becomes 570×570, and the number of channels changes from 1 to 64, which means that 64 3×3 convolution kernels are used for convolution operation. The improvement of the improved Unet neural network model is to change the input image block size to 512×512, and replace the downsampling part with depthwise separable convolution, that is, Figure 2 All horizontal arrows pointing downward on the left side of the image represent depthwise separable convolutions. To reduce model parameter count, the original non-padded convolutions are replaced with padded convolutions. This means the horizontal dimensions remain unchanged, but the number of channels varies. This ensures that the output mask size matches the input image size. The two gray arrows above, representing the skip connection layers used for cross-layer connections, have been removed. This process replaces complex late-stage feature fusion (such as feature pyramids) with early channel fusion, significantly improving real-time performance while maintaining accuracy, meeting the real-time inspection requirements of mobile robots.

[0020] Furthermore, the training process of the improved Unet neural network model includes: The binary cross entropy loss function and RMSprop optimizer are used for parameter update; A dynamic learning rate adjustment strategy is introduced during the training process. When the Dice coefficient of the validation set does not increase for three consecutive epochs, the learning rate is reduced to 1 / 5 of the current value. A multimodal training set is constructed by mixing the labeled data of the HSV space and the labeled data of the SAM model, and the input ratio of the two types of labeled data is 3:1.

[0021] Specifically, a data configuration file is created based on the location, number of categories, and category names of the tunnel water leakage image dataset; and the model training parameters required for the model training process are set: (1) Number of training epochs: The number of training epochs is 150.

[0022] (2) Batch Size: Based on the server configuration, select 8 input images per batch.

[0023] (3) Loss function: Binary Cross Entropy is selected as the loss function when there is only a single-class segmentation target.

[0024] (4) Optimizer: Select the Root Mean Square Propagation (RMSprop) method as the optimizer for training the model and for learning rate updates.

[0025] Initialize the weights of the improved Unet water leakage detection model, calculate the confidence loss, classification loss, and bounding box loss through forward propagation, and update the parameters in the network through backpropagation; After completing the training with the training set, the model is tested with the test set to evaluate the model performance. The confusion matrix is ​​the basis for the index calculation. When evaluating the model, the inference results are divided into four parts: (1) True Positives (TP): The model predicts that the example is positive, and it is actually a positive example.

[0026] (2) False positive FP: The model predicts a positive example, but it is actually a negative example.

[0027] (3) False negative example FN: The model predicts a negative example, but it is actually a positive example.

[0028] (4) True negative example TN: the model predicts that it is a negative example, but it is actually a negative example.

[0029] The following evaluation indicators are used (1) Accuracy: The proportion of correct predictions in the model's prediction results to the total prediction values.

[0030] accuracy=(TP+TN) / (TP+FP+FN+TN) (2) Precision: The probability that a certain category is correctly predicted in the model’s prediction results.

[0031] precision=TP / (TP+FP) (3) Recall: The probability that a category is predicted correctly in the true value.

[0032] recall=TP / (TP+FN) (4) F1 score: The harmonic mean of the model's precision and recall. When there is an imbalance between positive and negative samples, there may be a large difference between precision and recall. The F1 score takes these two metrics into consideration to more comprehensively evaluate model performance.

[0033] F1=2×TP / (2×TP+FP+FN) (5) Dice coefficient: The ratio of the intersection of the double model prediction result and the true value to the sum of the prediction result and the true value, which is used to measure the similarity of image segmentation results.

[0034] Dice=2×TP / (TP+FN)+(TP+FP) (6) Intersection over union (IoU): The ratio of the intersection over union of the model prediction result and the true value, which is used to measure the accuracy of the segmentation result.

[0035] IoU=Dice / (2-Dice) And perform benchmark tests, including the following parameters: model size, model parameter number, model computation amount, inference speed, and frame rate (fps).

[0036] This implementation can achieve the following: enhanced feature complementarity. Through channel fusion, the network can simultaneously learn visible light texture, infrared temperature anomalies, and deep structure information, solving the problem of missed detection of a single data source in scenarios such as occlusion and reflection. Optimized computational efficiency: Compared with traditional complex multi-module feature fusion (such as spatial pyramid and CBAM), this solution directly fuses multi-source data at the input layer, reducing redundant calculations within the network and improving inference speed. Channel dimension fusion improves the segmentation accuracy (mIoU) of the model in scenarios with blurred tunnel leakage edges.

[0037] The principle of model improvement is explained as follows: First, tunnel inspection robots need to process 4K high-resolution images (typical size 3840×2160) in real time. Traditional Unet has delays in mobile deployment. By replacing the standard convolution layer with a depthwise separable convolution layer and removing the redundant cross-layer connection structure in the Unet network, the number of parameters is reduced by 82%. The depthwise separable convolution reduces the computational complexity of 3×3 convolution from C_in×C_out×9 to C_in×(1+9). Taking the typical layer C_in=64 and C_out=128 as an example, the original computational complexity is 64×128×9=73728. After the improvement, it is reduced to 64×(9+1)=640, and the inference speed is increased by 3.8 times. After deleting the cross-layer connections of the first and second layers, the processing time of a single frame of a 1080P image is reduced from 230ms to 60ms, which can support 30fps real-time detection (meeting the needs of mobile video streaming). Secondly, the edges of tunnel leakage often show blurred radial patterns (such as capillary seepage). The use of padding convolution operation can improve the edge segmentation accuracy and maintain the resolution of the feature map, so that the intersection-over-union (IoU) of 0.5mm micro-cracks is increased from 0.62 to 0.79, and the D of the splashing edge of the water droplet is improved. The ice coefficient increased by 19% (0.83→0.98). In addition, spatial alignment can be guaranteed, with the output mask and input image strictly mapped 1:1, the laser ranging coordinate matching error is reduced, and the spatial registration accuracy of multimodal data is high. Third, a single sensor is susceptible to interference from the tunnel environment (such as strong light reflection and water mist interference). Improved multimodal fusion (five-channel input) can resist interference. Specifically, if the lighting is insufficient, the missed detection rate is reduced through infrared compensation. If the wall is reflective, the false detection rate is reduced through depth verification. If condensation occurs on the surface, the missed detection rate is reduced through temperature gradient filtering.

[0038] S2. Collect multi-source detection data in the tunnel, wherein the multi-source detection data includes infrared thermal imaging data, visible light image data, and laser ranging data; Furthermore, the mobile inspection robot collects multi-source detection data in the tunnel, including: The 120° equal-angle layout of the disc servo manipulator arm enables spatial alignment of the infrared thermal imager, visible light camera, and laser rangefinder; Establish a synchronous acquisition mechanism based on hardware trigger signals. When the laser rangefinder detects that the target distance is less than a first preset threshold, it synchronously triggers the infrared thermal imager and the visible light camera to expose; In the data preprocessing stage, the infrared thermal imaging data, visible light image data and laser ranging data are unified into the same spatial coordinate system through affine transformation.

[0039] Specifically, the inspection robot consists of a tracked vehicle platform, a data acquisition structure, an equipment control and data processing structure, and a power supply structure. The data acquisition structure includes an infrared thermal imager, a visible light camera, a laser rangefinder, and a disc servo arm. The disc servo arm is fixedly connected to the top of the tracked vehicle platform. The disc of the disc servo arm is evenly distributed on the disc, with the three detection devices rotating at a 120-degree angle. The servo core board is located on the underside of the tracked vehicle platform and assists in controlling the disc servo arm. The equipment control and data processing structure is a laptop. The power supply structure is a lithium battery, fixed to the bottom plate of the tracked vehicle platform by a rolling belt. The equipment is powered by a lithium battery connector, DC male and female connectors, banana male and female connectors, and an aviation plug. The laptop is equipped with a communication interface, which is connected to the infrared thermal imager, visible light camera, and laser rangefinder respectively. The disc manipulator arm can be rotated in different directions using a servo. The servo's rotational angle limits and torque values ​​are determined by the required rotation angles and mass of the infrared thermal imager, visible light camera, and laser rangefinder. The operating ranges of the infrared thermal imager, visible light camera, and laser rangefinder are determined by the cross-sectional size and environmental conditions of the shield-lined tunnel. The lithium battery is charged by the included power adapter.

[0040] S3, fusing the temperature distribution image, the visible light image, and the depth image in the input sample set 1 according to the channel dimension, and inputting the fusion process into the improved Unet neural network model to generate a binary segmentation mask 8, the size of which is consistent with that of the visible light image; Furthermore, the step of fusing the temperature distribution image, the visible light image, and the depth image in the input sample set 1 according to the channel dimension and inputting the fusion processing into the improved Unet neural network model comprises the following steps: The input sample set 1 is fused according to the channel dimension to obtain a multi-channel input tensor; The multi-channel input tensor is input to the improved Unet neural network model.

[0041] Optionally, a visible light image (3-channel RGB) is used as the first 3 channels; an infrared thermal image (single-channel temperature distribution) is used as the 4th channel; and a depth image (single-channel) generated by laser ranging is used as the 5th channel. Finally, a 5-channel input tensor is generated and directly input into the network for end-to-end learning.

[0042] Specifically, the fusion process involves data preprocessing and spatial alignment, including spatial coordinate system 1 and data format standardization. Affine transformation is used to map the three data sets to the same coordinate system to achieve spatial coordinate system 1. The transformation matrix parameters are precalculated using a calibration plate to ensure pixel-level alignment. During data format standardization, the visible light image retains its RGB channels, the infrared thermal imaging data undergoes linear temperature normalization, and the laser ranging data undergoes bilinear interpolation to generate a depth map. Next, the temperature distribution image, visible light image, and depth image are fused channel-wise. The fusion principle treats the three data types as complementary features: visible light (three channels): surface texture and color (RGB), infrared (one channel): temperature anomalies (grayscale values), and depth (one channel): surface geometry (depth values). This yields a multi-channel input tensor. This multi-channel input tensor is then fed into an improved Unet neural network model. The output layer generates a probability map using a sigmoid activation function, which is then binarized at a threshold of 0.5.

[0043] S4, after adaptively splicing the input sample set 1, inputting it into the improved Unet neural network model to generate a binary segmentation mask 8, the size of which is consistent with the visible light image; S5. Comprehensively analyze the binary segmentation mask 8 to obtain the actual location of the water leakage area in the tunnel, the leakage range, and the leakage severity.

[0044] In a preferred embodiment, preprocessing the multi-source detection data to obtain the input sample set 1 includes the following steps: Separating the color features of the leaked water by color space conversion on the visible light image data to obtain a visible light image; Furthermore, the separating of the leakage water color features by color space conversion of the visible light image data comprises the following steps: After converting the visible light image data from RGB space to HSV space, a dual-color domain mask of mask1 and mask2 is established through linear normalization processing, where mask1 corresponds to the mask of the purple and blue transition area, and mask2 corresponds to the mask of the blue area; A purple area mask is obtained by performing a differential calculation on mask1-mask2, and the purple area mask is used as a water leakage area mask. A binary annotation is generated after a morphological closing operation.

[0045] Specifically, visible light image data is collected by an infrared camera, and the leakage locations of the tunnel water leakage image data after preliminary processing are annotated using the Labelme annotation tool to obtain the tunnel water leakage image SAM dataset; the collected original image (visible light image data) is converted from the RGB color space to the HSV color space to separate the hue, saturation and brightness information; the converted HSV image is linearly normalized to standardize the pixel values ​​to a predetermined range; the purple and blue transition area mask mask1 and the blue transition area mask mask2 are separated, and mask=mask1-mask2 is used. The purple areas, representing water leakage areas, were isolated and binarized into black and white masks. An HSV dataset of tunnel water leakage images was generated. Data augmentation was then performed on the tunnel water leakage images. This augmentation process included left-right flipping, rotation, upside-down flipping, random cropping, small-block deformation, and shearing. A combination of these methods was used, with a random cropping probability of 0.5, a rotation probability of 0.8, a left-right flip probability of 0.8, an upside-down flip probability of 0.3, a small-block deformation probability of 0.6, and a shearing probability of 0.5. During the training process, a total of 1,000 enhanced images were generated. Based on 199 infrared images of tunnel water leakage, two different datasets were created using an auxiliary annotation method based on the Segment Anything Model and an automatic annotation method based on the HSV color space. Each dataset contained 1,000 enhanced images and 199 original images. The 1,000 enhanced images were randomly divided into a training set and a validation set in a 9:1 ratio, with all original images used as the test set.

[0046] Converting the infrared thermal imaging data into a two-dimensional temperature distribution image having the same size as the visible light image; The laser ranging data is used to generate a two-dimensional depth image having the same size as the visible light image through an interpolation algorithm.

[0047] In a preferred embodiment, after generating the binary segmentation mask 8, the following steps are further included: The target pixel coordinates of the water leakage area are marked on the binary segmentation mask 8 .

[0048] In a preferred embodiment, the comprehensive analysis of the binary segmentation mask 8 comprises the following steps: Based on the original three-dimensional data of the laser ranging data, the target pixel coordinates of the binary segmentation mask 8 are mapped to the three-dimensional space of the tunnel to obtain the actual position of the water leakage area in the tunnel; Counting the number of pixels in the water leakage area in the binary segmentation mask 8, and calculating the actual area of ​​the leakage range in combination with the scale parameter of the depth image; The gradient information of the temperature distribution image and the color depth of the visible light image are fused to evaluate the severity of the leakage.

[0049] In a preferred embodiment, the fusing of the gradient information of the temperature distribution image and the color depth of the visible light image to evaluate the severity of leakage comprises the following steps: Construct temperature gradient matrix and visible light color depth matrix; defining a leakage index, wherein the leakage index is related to the actual area of ​​the leakage range; If the leakage index is greater than or equal to the second preset threshold, it is determined to be a serious leakage; if the leakage index is greater than the third preset threshold and less than the second preset threshold, it is determined to be a moderate leakage; if the leakage index is less than or equal to the third preset threshold, it is determined to be a slight leakage.

[0050] For example, a suspected leak was detected in a section of a subway tunnel. An inspection robot equipped with an infrared thermal imager, a visible light camera, and a laser rangefinder conducted inspections. The visible light image, a 1024×768 resolution RGB image, showed purple water stains on the tunnel wall. The infrared thermal imaging data, a 1024×768 resolution temperature distribution data for the same area, showed a temperature of 12°C (ambient temperature 18°C). The laser ranging data, a 3D point cloud containing the spatial coordinates (x, y, z) of each ranging point, was used. The RGB image was converted to HSV color space and linearly normalized to separate the hue (H), saturation (S), and value (V) channels. A mask for the transition region between purple and blue (mask1: H∈[250, 280], S>0.5) and a mask for the blue region (mask2: H∈[200, 240], S>0.6) were defined. The difference mask was calculated: mask = mask1 - mask2, resulting in a binary mask for the purple region. Perform morphological closing operations (dilation + erosion) on the mask to eliminate noise and smooth the boundaries. Convert the original temperature data (matrix form) collected by the infrared sensor into a grayscale temperature distribution map (1024×768) that is consistent with the size of the visible light image, where each pixel value corresponds to the temperature. The three-dimensional point cloud data is subjected to a bilinear interpolation algorithm to generate a two-dimensional depth map (1024×768) that is consistent with the resolution of the visible light image. Each pixel value represents the distance (unit: meter). The depth map is aligned with the spatial coordinates of the visible light image. Visible light image: Preserve The original RGB three channels (1024×768×3) are retained, the temperature distribution map is used as the fourth channel (single-channel grayscale image), and the depth map is used as the fifth channel (single-channel grayscale image). They are spliced ​​according to the channel dimension to generate a 5-channel input tensor (1024×768×5). The 5-channel input tensor is input into the improved Unet model. The last layer outputs a probability map (1024×768) through the Sigmoid function. The threshold is set to 0.5 to generate a binary segmentation mask 8. In the segmentation mask, all pixel coordinates (i, j) with a value of 1 are marked, such as (500, 300), (501, 300), etc. According to the original 3D laser ranging data, each marked pixel (i, j) in the mask is mapped to the tunnel space coordinate (x, y, z). Example: Pixel (500, 300) is mapped to the actual location (x = 10.5 m, y = 3.2 m, z = 0 m). The number of marked pixels in the mask is counted as 5000. Based on the depth map parameters (1 pixel = 0.01 m2), the actual leakage area is calculated as 5000 × 0.01 = 50 m2. Temperature gradient: The average temperature gradient ΔT of the leakage area is extracted as 6°C (ΔT = ambient temperature - leakage area temperature).Color depth: The average saturation of the extracted leakage area in the HSV space is S = 0.8, and the leakage index = 0.6 × ΔT + 0.4 × S = 0.6 × 6 + 0.4 × 0.8 = 3.6 + 0.32 = 3.92. According to the judgment rules: severe leakage: ≥ 3.5, moderate leakage: 2.0-3.5, slight leakage: ≤ 2.0, the leakage degree is judged as severe leakage.

[0051] Example 2 Based on Example 1, this embodiment further improves the detection accuracy, specifically: In a preferred embodiment, the method further comprises constructing a physical characteristic constraint function of the leaking water, wherein the physical characteristic constraint function comprises a temperature-depth correlation constraint and a color-depth consistency constraint; After the binary segmentation mask 8 is generated, the following steps are included: Divide the binary segmentation mask 8 into 10×10 pixel grids, and simultaneously calculate a neural network confidence score and a physical constraint compliance score for each grid, wherein the neural network confidence score is obtained by outputting the improved Unet neural network model, and the physical constraint compliance score is obtained by the physical characteristic constraint function of the leakage water; The comprehensive analysis of the binary segmentation mask 8 includes the following steps: If the deviation between the neural network confidence score and the physical constraint compliance score is less than or equal to a fourth preset threshold, a comprehensive analysis is performed on the binary segmentation mask 8 .

[0052] Specifically, a physical constraint verification module can be added to the decoder end of the improved Unet neural network model. The verification module is a physical feature constraint function. The temperature-depth correlation constraint is: a mathematical relationship between temperature gradient and surface depth is established according to the heat conduction equation. When the laser ranging data shows that the surface roughness of the target is greater than 5mm, the absolute value of the temperature gradient in the corresponding area is required to be ≥0.8℃ / cm; the color-depth consistency constraint is: an inverse relationship between color saturation and surface roughness is defined. When the blue channel value of the visible light image is greater than 200 and the standard deviation of the depth data is less than 2mm, the water leakage area judgment is forcibly activated. The physical common sense compliance score is related to the degree of compliance with the color rule (color-depth consistency constraint) and the temperature rule (temperature-depth correlation constraint). The temperature rule score is calculated as: |actual temperature gradient - 0.8| / 0.8 (deviation ratio from the standard value). For example, if the measured temperature gradient in a certain area is 0.6°C / cm, the score is 1 - |0.6-0.8| / 0.8 = 0.75. The color rule score is calculated as: blue channel value / 200 × (1 - depth standard deviation / 2). For example, if the blue value is 210 and the depth fluctuation is 1mm, the score is (210 / 200) × (1-1 / 2) = 1.05 × 0.5 = 0.525. Therefore, the average physical constraint compliance score is approximately 0.6. The neural network confidence score is derived from the predicted probability value (between 0 and 1) for each pixel output by the sigmoid output of the last layer of the neural network. For example, if the average probability of a region is 0.92, the model is 92% confident that it is a leak. Therefore, the neural network confidence score is approximately 0.9. If the deviation = |neural network confidence score - physical constraint compliance score| / (neural network confidence score + physical constraint compliance score) = 0.2 and is greater than 0.15 (15%), a review is triggered.

[0053] This implementation uses physical constraints to eliminate false positives that don't conform to the physical characteristics of leaks. For example, a region may be misidentified as a leak by the model due to light reflection (confidence 0.9), but its depth data indicates it is located inside the tunnel structure (physical score 0.1). A deviation of 0.8 exceeds the threshold of 0.2, so it is excluded. Missed detection correction: Areas with low confidence that meet physical constraints are revalidated. For example, a leak region may have a model confidence of 0.4 due to image blur, but its temperature gradient matches its depth (physical score 0.9). A deviation of 0.5 exceeds the threshold, triggering manual review to avoid missed detections.

[0054] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of the present invention, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.

Claims

1. A method for detecting tunnel water leakage based on neural network, characterized in that: The following steps are involved: The standard convolutional layer in the Unet network is replaced with a depth-wise separable convolutional layer, a padded convolution operation is adopted, and the redundant cross-layer connection structure in the Unet network is deleted to obtain an improved Unet neural network model; Collecting multi-source detection data in the tunnel, wherein the multi-source detection data includes infrared thermal imaging data, visible light image data, and laser ranging data; Preprocessing the multi-source detection data to obtain an input sample set (1), wherein the input sample set (1) includes a temperature distribution image, a visible light image, and a depth image; wherein the sizes of the images in the input sample set (1) are consistent; The temperature distribution image, the visible light image and the depth image in the input sample set (1) are fused according to the channel dimension and input into the improved Unet neural network model to generate a binary segmentation mask (8) whose size is consistent with that of the visible light image; A comprehensive analysis is performed on the binary segmentation mask (8) to obtain the actual location of the water leakage area in the tunnel, the leakage range and the leakage severity.

2. The neural network-based tunnel water leakage detection method according to claim 1, characterized in that: The multi-source detection data is pre-processed to obtain an input sample set (1), comprising the following steps: Separating the color features of the leaked water by color space conversion on the visible light image data to obtain a visible light image; Converting the infrared thermal imaging data into a two-dimensional temperature distribution image having the same size as the visible light image; The laser ranging data is used to generate a two-dimensional depth image having the same size as the visible light image through an interpolation algorithm.

3. The neural network-based tunnel water leakage detection method according to claim 1, characterized in that: The step of fusing the temperature distribution image, the visible light image, and the depth image in the input sample set (1) according to the channel dimension and inputting the fusion process into the improved Unet neural network model comprises the following steps: The input sample set (1) is fused according to the channel dimension to obtain a multi-channel input tensor; The multi-channel input tensor is input to the improved Unet neural network model.

4. The neural network-based tunnel water leakage detection method according to claim 1, characterized in that: After generating the binary segmentation mask (8), the following steps are also included: The target pixel coordinates of the water leakage area are marked on the binary segmentation mask (8).

5. The neural network-based tunnel water leakage detection method according to claim 4, characterized in that: The comprehensive analysis of the binary segmentation mask (8) comprises the following steps: Based on the original three-dimensional data of the laser ranging data, the target pixel coordinates of the binary segmentation mask (8) are mapped to the three-dimensional space of the tunnel to obtain the actual position of the water leakage area in the tunnel; Counting the number of pixels in the water leakage area in the binary segmentation mask (8), and calculating the actual area of ​​the leakage range in combination with the scale parameter of the depth image; The gradient information of the temperature distribution image and the color depth of the visible light image are fused to evaluate the severity of the leakage.

6. The neural network-based tunnel water leakage detection method according to claim 2, characterized in that: The step of separating the color features of the leaked water by color space conversion of the visible light image data comprises the following steps: After converting the visible light image data from RGB space to HSV space, a dual-color domain mask of mask1 and mask2 is established through linear normalization processing, where mask1 corresponds to the mask of the purple and blue transition area, and mask2 corresponds to the mask of the blue area; A purple area mask is obtained by performing a differential calculation on mask1-mask2, and the purple area mask is used as a water leakage area mask. A binary annotation is generated after a morphological closing operation.

7. The neural network-based tunnel water leakage detection method according to claim 1, characterized in that: The training process of the improved Unet neural network model includes: The binary cross entropy loss function and RMSprop optimizer are used for parameter update; A dynamic learning rate adjustment strategy is introduced during the training process. When the Dice coefficient of the validation set does not increase for three consecutive epochs, the learning rate is reduced to 1 / 5 of the current value. A multimodal training set is constructed by mixing the labeled data of the HSV space and the labeled data of the SAM model, and the input ratio of the two types of labeled data is 3:

1.

8. The neural network-based tunnel water leakage detection method according to claim 1, characterized in that: The mobile inspection robot collects multi-source inspection data in the tunnel, including: The 120° equal-angle layout of the disc servo manipulator arm enables spatial alignment of the infrared thermal imager, visible light camera, and laser rangefinder; Establish a synchronous acquisition mechanism based on hardware trigger signals. When the laser rangefinder detects that the target distance is less than a first preset threshold, it synchronously triggers the infrared thermal imager and the visible light camera to expose; In the data preprocessing stage, the infrared thermal imaging data, visible light image data and laser ranging data are unified into the same spatial coordinate system through affine transformation.

9. The neural network-based tunnel water leakage detection method according to claim 5, characterized in that: The step of fusing the gradient information of the temperature distribution image with the color depth of the visible light image to evaluate the severity of leakage includes the following steps: Construct temperature gradient matrix and visible light color depth matrix; defining a leakage index, wherein the leakage index is related to the actual area of ​​the leakage range; If the leakage index is greater than or equal to the second preset threshold, it is determined to be a serious leakage; if the leakage index is greater than the third preset threshold and less than the second preset threshold, it is determined to be a moderate leakage; if the leakage index is less than or equal to the third preset threshold, it is determined to be a slight leakage.

10. The method for detecting tunnel water leakage based on neural network according to claim 1, characterized in that: The method further includes constructing a physical feature constraint function of the leaking water, wherein the physical feature constraint function includes a temperature-depth correlation constraint and a color-depth consistency constraint; After generating the binary segmentation mask (8), the following steps are included: Divide the binary segmentation mask (8) into 10×10 pixel grids, and simultaneously calculate a neural network confidence score and a physical constraint compliance score for each grid, wherein the neural network confidence score is obtained by outputting the improved Unet neural network model, and the physical constraint compliance score is obtained by the leakage water physical feature constraint function; The comprehensive analysis of the binary segmentation mask (8) comprises the following steps: If the deviation between the neural network confidence score and the physical constraint compliance score is less than or equal to a fourth preset threshold, a comprehensive analysis is performed on the binary segmentation mask (8).

Citation Information

Patent Citations

  • Image change detection method based on depth-separable convolution network

    CN108846835A

  • Landslide identification method and device based on satellite data UNet network model

    CN115131684A

  • Thermal infrared image dam leakage dangerous case intelligent identification method based on multi-task assistance

    CN116580328A

  • Single target positioning method based on lightweight improved Unet semantic segmentation network

    CN117237620A

  • Tunnel water seepage detection and identification method and system based on infrared camera and deep learning

    CN118279565A