Concrete seepage detection system based on multispectral video stream
By using a multispectral video stream detection system, combined with visible light and infrared cameras, multimodal and temporal background information is extracted and fused, solving the problem of identifying tunnel leakage in low light and humid environments, and achieving efficient and accurate leakage detection.
Patent Information
- Application Number
- CN202510944271.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing deep learning-based tunnel leakage detection methods struggle to effectively identify leakage areas in low-light and damp environments. Furthermore, existing dual-modal detection methods neglect the flow phenomenon in the seepage area, limiting their detection performance.
A detection system based on multispectral video streams is adopted, which combines images acquired by visible light and infrared cameras. Features are extracted through a shared-weight backbone network. By utilizing the spatial and temporal features of the visible light sequence and the motion information of the infrared sequence, multimodal and temporal background information are fused, and a motion information fusion module and a time-space information fusion module are designed to optimize the leakage detection model.
It significantly improves the accuracy of seepage area identification in low light and humid environments, reduces false detection and false negative rates, and achieves real-time processing and efficient seepage detection.
Smart Images

Figure CN120807457A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of tunnel seepage detection, and in particular to a concrete seepage detection system based on multi-spectral video stream. BACKGROUND
[0002] The statements in this section merely provide background information related to the present disclosure and can not constitute the prior art.
[0003] Tunnels are an important part of modern transportation and infrastructure construction. However, due to the combined effects of factors such as groundwater erosion, unreasonable construction technology, and aging and deterioration of lining materials, tunnels may experience seepage. If not detected and repaired in a timely manner, water leakage may affect the stability of the tunnel structure and endanger the safety of the building structure. Therefore, in order to ensure the structural safety of the tunnel, it is necessary to regularly inspect the water leakage phenomenon of the tunnel lining.
[0004] Currently, the commonly used tunnel seepage detection mainly relies on visual recognition and recording by manual on-site detection. The detection process is time-consuming, and the accuracy of the recognition is highly dependent on the experience and subjective judgment of the personnel. In addition, it is extremely dangerous to expose personnel to the roadway working face where working face instability and water inrush may occur. With the development of computer technology, scholars have adopted various image processing techniques and deep learning (DL) models based on images, gradually tending towards automated detection to replace manual detection, improving the efficiency and accuracy of seepage detection. At the same time, compared to traditional digital image processing techniques for seepage detection, DL-based seepage detection can automatically learn feature representations, has stronger expression ability and adaptability. In addition, this method has better applicability to changes in lighting conditions and image noise, thereby providing better recognition stability. Therefore, tunnel water leakage intelligent detection methods based on deep learning algorithms have gradually become mainstream.
[0005] However, the existing water leakage recognition methods based on deep learning still have certain limitations, mainly including the following two aspects: (1) Currently, seepage data acquisition is mainly through visible light cameras to capture images, relying on the unique color, texture, and edge information of the water leakage area to distinguish water leakage from other background interference. This method is highly dependent on the image quality captured by the visible light camera. However, visible light cameras are very sensitive to light intensity, and in some scenes with weak light intensity, it is often difficult to effectively detect seepage relying solely on single-mode visible light images captured by optical cameras. (2) Existing research mainly focuses on the recognition of obvious water leakage phenomena, while ignoring seepage detection in the presence of wall wetness. Due to the wall wetness and the presence of accumulated water, the interference is large, and it is difficult to effectively identify whether the area is wall water or a new seepage area. Therefore, how to detect seepage in scenes with weak light and wall wetness is a problem that remains to be solved. SUMMARY
[0006] The purpose of the present application is to provide a concrete seepage detection system based on multi-spectral video stream, which considers that the temperature of seepage water is different from that of wall area water, introduces a thermal infrared (IR) camera, combines a visible light (VIS) color camera to shoot a dual-mode multi-spectral image, provides rich and complementary information, and thus improves the detection performance in a scene with insufficient light intensity. The existing dual-mode detection method is based on a single image, which ignores the flow phenomenon of the seepage area and limits its performance in processing such dynamic scenes. By introducing a video stream, the difference between consecutive video frames can effectively perceive the flow part, thus capturing dynamic motion information. Therefore, in order to simultaneously utilize the complementary information and motion information of dual-mode data, a method of jointly learning semantic representation from multi-spectral and temporal context is also proposed.
[0007] The technical scheme of the present application is as follows: A concrete seepage detection system based on multi-spectral video stream, comprising: a visible light camera, an infrared camera, and a seepage detection model. Wherein the visible light camera and the infrared camera are used to synchronously collect visible light video stream and infrared video stream. The seepage detection model comprises: a group of main stems for extracting initial features of dual-mode image sequences, branches for spatial domain and temporal domain feature extraction of visible light sequences, branches for motion information extraction of infrared sequences, two fusion modules for fusing motion information and final features, respectively, and a final detection head.
[0008] Further, the main stem adopts a backbone network sharing weight parameters, which can extract visible light sequence initial feature maps and infrared sequence initial feature maps of seepage areas from visible light video stream and infrared video stream.
[0009] Further, the backbone network is constructed by DarkNet53 and PAFPN; wherein the deepest two convolution layers of PAFPN are simplified.
[0010] Further, the branch for spatial domain and temporal domain feature extraction of visible light sequences comprises: a visible light spatial information extraction branch and a visible light motion information extraction branch.
[0011] Further, the visible light spatial information extraction branch comprises: (a) the reference frame is summarized by a concatenation operation, and then the summarized reference features are learned by two basic 3x3 convolution to extract potential information, and the features are converted into calibration factors by using a sigmoid function ; (b) using a calibration factor The current key frame is calibrated, and the potential features are fully utilized in a multiplexing manner; the first multiplexed feature fd is obtained using Hadamard product operation, and the initial key frame information fc is further multiplexed through residual connection to form a multi-feature dual multiplexing enhanced key frame spatial information; (c) The dual multiplexing features fc and fd are spliced at the channel level and input again to two layers of 3x3 basic convolution components to obtain the final visible light spatial information, i.e., spatial feature .
[0012] Further, the visible light motion information extraction branch comprises: The difference information of all adjacent frames in the time window is superimposed, the motion features are further amplified by adding residual blocks to enhance the difference information; at the same time, in order to reduce the calculation cost of time dynamic coding, the pooling method is used to reduce the resolution while highlighting the difference information to cope with the dynamic changes of the target in the time dimension; finally, two residual blocks are used to enhance local motion, and motion dependent relationship and target features are deeply fused to finally obtain visible light motion information.
[0013] Further, the branch for extracting motion information of the infrared sequence comprises: The continuous 5 frames of images of the infrared video stream are spliced in the channel dimension, and the spliced features are enhanced through the frequency domain feature enhancement module; The pixel average value of the 5 frames of images is calculated to obtain a multi-channel residual map, and the multi-channel residual map is used to calibrate the frequency domain enhanced image, reduce the interference of static background, and capture the motion information of the infrared video sequence, and finally obtain the infrared motion information.
[0014] Further, the two fusion modules are respectively: a motion information fusion module and a time-space information fusion module; The motion information fusion module fuses the visible light motion information and the infrared motion information by channel splicing; at the same time, a depth separable convolution with a convolution kernel size of 3x3 is used to further mine the motion information of the two modalities; The time-space information fusion module is used to fuse the time feature and the spatial feature after the motion information fusion module, to obtain a feature .
[0015] Further, the time-space information fusion module is configured to perform the following processing: First, a global average pooling operation is used to respectively process the time feature and spatial features The pooling dimension reduction is performed to obtain global information modeling features and ; Then, a 1*1 convolution is used to adjust the channel, and a ReLu activation function is used to improve the nonlinear learning ability, and the obtained features are input into a Sigmoid activation function to obtain a spatial calibration factor and a time calibration factor ; meanwhile, considering the interaction between the two features, the Hadamard product is used to model the interactive calibration of the two features, and the interaction between the two features is enhanced; Then, the features after interactive calibration are preliminarily calibrated and fused by element addition, and a convolution operation and a ReLu activation function are used to obtain more fully fused features ; Finally, a sigmoid activation function is used to further improve the nonlinear ability while further reusing the time features and the spatial features to further calibrate, so as to obtain features fully fused with time and space information .
[0016] Further, in the detection head, a special hyperparameter is introduced to control the classification loss term for the single-class detection of leakage.
[0017] Compared with the prior art, the present application has the following beneficial effects: 1. The present application effectively overcomes the detection blind area of single visible light mode in weak light and humid environment through synchronous acquisition and joint analysis of visible light and infrared video stream, and significantly improves the seepage area recognition accuracy by using the sensitivity of infrared spectrum to temperature difference.
[0018] 2. The present application adopts double-path multiplexing spatial enhancement and time difference motion amplification branch for visible light sequence, fully extracts the spatial texture and dynamic flow characteristics of seepage; and introduces frequency domain enhancement and residual calibration for infrared sequence to suppress static background interference and enhance the sensitivity of small seepage motion.
[0019] 3. The present application adopts a simplified PAFPN backbone network with shared weights to reduce redundant calculation and realize real-time processing; at the same time, through interactive spatio-temporal calibration fusion, the contribution weight of different modal features is adaptively balanced to improve the robustness in complex scenes.
[0020] 4. The present application introduces a single-class classification loss hyperparameter in the detection head to optimize the seepage detection task, so that the false detection rate and the missed detection rate are greatly reduced. DETAILED DESCRIPTION
[0021] Figure 1A block diagram of a concrete seepage detection system based on a multi-spectral video stream. DETAILED DESCRIPTION
[0022] It should be noted that the terms "first", "second", and so on, and the like relational terms only serve to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0023] The features and performances of the present application are further described in detail below in conjunction with embodiments.
[0024] Embodiment One Please refer to Figure 1 A concrete seepage detection system based on a multi-spectral video stream, comprising: a visible light camera, an infrared camera, and a seepage detection model; wherein the visible light camera and the infrared camera are used to synchronously collect visible light video streams and infrared video streams; the seepage detection model comprises: a group of backbones for extracting initial features from a double-mode image sequence, branches for spatial domain and time domain feature extraction of the visible light sequence, a branch for motion information extraction of the infrared sequence, two fusion modules for fusing motion information and final features, respectively, and a final detection head; The running algorithm of the system is roughly as follows: First, a group of backbones sharing weight parameters are used to extract initial features from a double-mode seepage frame sequence, to obtain infrared sequence features and visible light sequence features. Since the visible light information contains rich spatial texture information, two parallel branches are used to extract spatial information and time information, respectively, for the visible light sequence features. At the same time, since the texture information in the infrared sequence is less, but the change of the seepage area is more sensitive, a separate branch is used to extract motion information for the infrared sequence features. Then, since the infrared sequence and the visible light sequence are shooting the same scene seepage phenomenon, the motion information contained should have high similarity, so a motion information fusion module is used to fuse the motion of the two modalities, to obtain complementary motion features. Finally, the time features and the spatial features are fused to enhance the feature representation, which is used for model training and detection of the seepage area by the training head.
[0025] In this embodiment, specifically, the backbone adopts a backbone network sharing weight parameters, which can extract visible light sequence initial feature maps and infrared sequence initial feature maps of the leakage area from the visible light video stream and the infrared video stream; it should be noted that YOLOX is adopted in this embodiment; In this embodiment, specifically, the backbone network is constructed by DarkNet53 and PAFPN; wherein, the two deepest convolution layers of the PAFPN are simplified; that is, the two deepest convolution layers commonly used for extracting and fusing common target features in the PAFPN are simplified in full consideration of the particularity of the leakage; in addition, in order to simplify the parameter size, all backbone networks share the same set of parameters; after the sequence frame passes through the backbone, the initial feature map of the leakage area is obtained.
[0026] In this embodiment, specifically, the branch for spatial domain and time domain feature extraction of the visible light sequence includes a visible light spatial information extraction branch and a visible light motion information extraction branch.
[0027] In this embodiment, specifically, the visible light spatial information extraction branch includes: (a) the reference frame is summarized through a splicing operation, and then the summarized reference feature is learned through two basic 3 × 3 convolution operations to extract potential information, and the feature is converted into a calibration factor by using a sigmoid function ; (b) the calibration factor is used to calibrate the current key frame, and a multiplexing manner is adopted to fully utilize the potential features; the first multiplexed feature fd is obtained by using Hadamard product operation, and the initial key frame information fc is further multiplexed by using a residual connection manner, so as to form a multi-feature double-multiplexing enhanced key frame spatial information; (c) the double-multiplexed features fc and fd are spliced at the channel level, and are input into two layers of 3 × 3 basic convolution components again to obtain the final visible light spatial information, i.e. spatial feature .
[0028] In the embodiment, it is to be noted that more color, texture and other detail information in the visible light sequence are crucial for leakage detection. Therefore, a spatial information extraction module is designed in the embodiment, which fully considers the multiple frames of visible light images and further extracts more complete spatial information. When considering leakage detection, the current frame is mainly relied on for detection, while the information of multiple frames of images in the video sequence is also referred to. Therefore, the main idea of extracting visible light spatial information is to use the features of multiple reference frames to perform multi-feature dual multiplexing operation on the current leakage image frame, which is specifically as follows: First, the reference frames are summarized through splicing operation, and then the summarized reference features are learned through two basic 3x3 convolution to extract potential information, and the features are converted into calibration factors using sigmoid function . This process can be represented by the following formula:
[0029] Then, the calibration factor is used to calibrate the current key frame, and the multi-multiplexing method is used to fully utilize the potential features. The first multi-multiplexing feature fd is obtained by using Hadamard product operation, and the initial key frame information fc is further multiplexed through residual connection, so as to form a multi-feature dual multiplexing enhanced key frame spatial information. Finally, the dual multiplexing features fc and fd are spliced in the channel layer, and input into two layers of 3x3 basic convolution components again to obtain the final visible light spatial information fspatial, which is represented by the following formula:
[0030] In the embodiment, specifically, the visible light motion information extraction branch includes: The difference information of all adjacent frames in the time window is superimposed, the difference information is enhanced, and the residual block is added to further amplify the motion features. At the same time, in order to reduce the calculation cost of time dynamic coding, the pooling method is used to reduce the resolution while highlighting the difference information to cope with the dynamic change of the target in the time dimension. Finally, two residual blocks are used to enhance local motion, and the motion dependent relationship and target features are deeply fused to finally obtain the visible light motion information.
[0031] In this embodiment, it should be noted that in order to more effectively extract the leaked motion pattern, the differential information of all adjacent frames is superimposed within the time window, while solving the sparsity problem observed in a single differential. In addition, the motion features are further amplified by enhancing the differential information and adding residual blocks. In addition, in order to reduce the computational cost of temporal dynamic coding, pooling is used to reduce the resolution while highlighting the differential information to cope with the dynamic changes of the target in the temporal dimension. Finally, two residual blocks are used to enhance local motion and deeply fuse motion dependencies and target features.
[0032]
[0033]
[0034]
[0035]
[0036] In this embodiment, specifically, the branch for extracting motion information of infrared sequences includes: The five consecutive frames of infrared video stream are spliced in the channel dimension, and the spliced features are enhanced by the frequency domain feature enhancement module; The multi-channel residual map is obtained by calculating the pixel average of 5 frames of images, and the multi-channel residual map is used to calibrate the image after frequency domain enhancement, thereby reducing the interference of static background while capturing the motion information of the infrared video sequence, and finally obtaining the infrared motion information.
[0037] In this embodiment, it should be noted that during the identification of water leakage in tunnels, water flow typically exhibits dynamic changes due to factors such as lighting conditions, texture, and color. However, static backgrounds such as stains and damp water can seriously interfere with the identification of water leakage. Unlike general visible light images, infrared images typically lack obvious spatial textures but contain rich leakage contours and shapes. Therefore, the focus is on detecting the temporal characteristics of moving targets from video sequence frames. From a frequency domain perspective, low-frequency information contains positional relationships and spatial information, while high-frequency information primarily contains specific boundary details. To this end, a multi-channel motion perception branch is designed to capture temporal features.
[0038] First, five consecutive frames of the video stream are concatenated in the channel dimension and features are enhanced in the frequency domain to avoid underutilization of spatial information and better extract temporal motion features. Next, a multi-channel residual image is obtained by calculating the pixel average of the five frames. This image is used to calibrate the frequency-domain enhanced image, reducing static background interference while capturing motion information in the infrared video sequence.
[0039] In the embodiment, specifically, the two fusion modules are: a motion information fusion module and a time-space information fusion module.
[0040] In the embodiment, specifically, the motion information fusion module fuses the visible light motion information and the infrared motion information in a channel splicing manner; and a depth separable convolution with a convolution kernel size of 3*3 is used to further mine the motion information of the two modalities. The time-space information fusion module is configured to fuse the time feature and the space feature fused by the motion information fusion module to obtain a feature fully fused with time and space information. .
[0041] It should be noted that, since the video streams of the visible light modality and the infrared modality collect data of the same leakage phenomenon, the data of the two modalities have high similarity in motion information. The embodiment directly fuses the motion information features of the two modalities in a channel splicing manner to sufficiently utilize their potential for leakage detection. In addition, in order to reduce the amount of calculation and improve the operation efficiency, a depth separable convolution (DWConv) with a convolution kernel size of 3*3 is used to further mine the motion information of the two modalities. The whole process can be expressed by a formula as follows:
[0042] In the embodiment, specifically, the time-space information fusion module is configured to perform the following processing: First, a global average pooling operation is used to respectively pool and reduce the dimension of the time feature and the space feature to obtain global information modeling features and . Then, a 1*1 convolution is used to adjust the channel, and a ReLu activation function is used to improve the nonlinear learning ability, and the obtained features are respectively input into a Sigmoid activation function to obtain a space calibration factor and a time calibration factor . At the same time, considering the interaction between the two features, a Hadamard product is used to interactively calibrate the two features to strengthen the interaction between the two features. Next, the features after interactive calibration are preliminarily calibrated and fused in an element addition manner, and a convolution operation and a ReLu activation function are used to obtain more fully fused features . Finally, a sigmoid activation function is used to further improve the nonlinear ability while preliminarily calibrating and fusing the time feature and the space feature .
[0043] In this embodiment, it should be noted that the time feature and the space feature are obtained by the time motion perception branch and the space branch respectively. However, due to the great semantic difference between the two features, there is an essential difference. Simple feature addition or splicing often cannot fully utilize their potential for target detection. In view of this, we specially design a new complementary correction fusion module to realize the cross-domain fusion of time and space features. Specifically, first, the global average pooling operation is used to respectively pool and reduce the dimension of the two branches to obtain the global information modeling features Ftem and Fspa. Then, the 1x1 convolution is used to adjust the channel, and the ReLu activation function is used to improve the nonlinear learning ability, and the obtained features are respectively input into the Sigmoid activation function to obtain the spatial calibration factor and the time calibration factor . At the same time, considering the interaction between the two branch features, the Hadamard product is used to model the interactive calibration of the two features, and the interaction ability between the two branches is strengthened. Then, the element addition method is used to preliminarily calibrate and fuse the features after interactive calibration, and the convolution operation and the ReLu activation function are used to obtain more fully fused features . The formula representation of the whole process is as follows:
[0044]
[0045]
[0046] Finally, the sigmoid activation function is used to further improve the nonlinear ability while further reusing and calibrating the initial time feature and the space feature to obtain the final time and space information fully fused feature .
[0047]
[0048] In this embodiment, specifically, the original YOLOX detection head is designed for multi-class object detection, so the classification loss in the loss function is designed for multi-class classification. However, the leakage detection task only involves single-class target detection (i.e. flow leakage), so we introduce a special hyperparameter to control the classification loss term based on the original YOLOX detection head for single-class detection of leakage, instead of using the original multi-class loss term.
[0049]
[0050] The above-described embodiments are merely illustrative for the present application and are described in more detail and specifically, but should not be understood as a limitation to the protection scope of the present application. It should be noted that, for those skilled in the art, some modifications and improvements can be made without departing from the technical concept of the present application, and these all belong to the protection scope of the present application.
[0051] This Background section is intended to provide a general overview of the context of the application, the work of the current named inventors, the work described in the Background section to the extent it is not considered prior art by the applicant, and is not admitted to be prior art to the application by this Background section description, either explicitly or implicitly, at the time of filing.
Claims
1. A concrete seepage detection system based on multispectral video stream, characterized in that: include: Visible light cameras, infrared cameras, and leak detection models; Among them, the visible light camera and the infrared camera are used to synchronously capture the visible light video stream and the infrared video stream; The leakage detection model includes: a backbone for extracting initial features from a bimodal image sequence, branches for extracting spatial and temporal features from a visible light sequence, a branch for extracting motion information from an infrared sequence, two fusion modules for fusing motion information and final features respectively, and a final detection head.
2. The concrete seepage detection system based on multispectral video stream according to claim 1 is characterized in that: The backbone adopts a backbone network with shared weight parameters, and can extract the initial feature map of the visible light sequence and the initial feature map of the infrared sequence of the leakage area from the visible light video stream and the infrared video stream.
3. The concrete seepage detection system based on multispectral video stream according to claim 2 is characterized in that: The backbone network is constructed by DarkNet53 and PAFPN; wherein, the deepest two convolutional layers of PAFPN are simplified.
4. The concrete seepage detection system based on multispectral video stream according to claim 3 is characterized in that: The branches used for extracting spatial and temporal features of visible light sequences include: a visible light spatial information extraction branch and a visible light motion information extraction branch.
5. The concrete seepage detection system based on multispectral video stream according to claim 4 is characterized in that: The visible light spatial information extraction branch includes: (a) The reference frame The concatenation operation is used to aggregate the reference features, and then two basic 3×3 convolutions are used to learn the aggregated reference features to extract potential information, and the sigmoid function is used to convert the features into calibration factors. ; (b) Using calibration factors The current keyframe is calibrated and the potential features are fully utilized by multiplexing. The first multiplexed feature fd is obtained by Hadamard product operation. At the same time, the initial keyframe information fc is further reused through residual connection to form multi-feature dual-multiplexing to enhance the spatial information of the keyframe. (c) The dual-channel multiplexed features fc and fd are concatenated at the channel level and input into the two-layer 3 × 3 basic convolution component again to obtain the final visible light spatial information, i.e., the spatial feature .
6. The concrete seepage detection system based on multispectral video stream according to claim 5, characterized in that: The visible light motion information extraction branch includes: The differential information of all adjacent frames is superimposed within the time window. By enhancing the differential information, residual blocks are added to further amplify the motion features. At the same time, to reduce the computational cost of temporal dynamic encoding, pooling is used to reduce the resolution while highlighting the differential information to cope with the dynamic changes of the target in the temporal dimension. Finally, two residual blocks are used to enhance local motion, and motion dependencies and target features are deeply fused to finally obtain visible light motion information.
7. The concrete seepage detection system based on multispectral video stream according to claim 6, characterized in that: The branch for motion information extraction of infrared sequences includes: The five consecutive frames of infrared video stream are spliced in the channel dimension, and the spliced features are enhanced by the frequency domain feature enhancement module; The multi-channel residual map is obtained by calculating the pixel average of 5 frames of images, and the multi-channel residual map is used to calibrate the image after frequency domain enhancement, thereby reducing the interference of static background while capturing the motion information of the infrared video sequence, and finally obtaining the infrared motion information.
8. The concrete seepage detection system based on multispectral video stream according to claim 7, characterized in that: The two fusion modules are: motion information fusion module and time-space information fusion module; The motion information fusion module uses channel splicing to fuse visible light motion information with infrared motion information. At the same time, it uses depthwise separable convolution with a convolution kernel size of 3×3 to further mine the motion information of the two modalities. The time-space information fusion module is used to combine the time features fused by the motion information fusion module and spatial characteristics The fusion is performed here to obtain the final feature that fully integrates the time and space information .
9. The concrete seepage detection system based on multispectral video stream according to claim 8, characterized in that: The temporal-spatial information fusion module is configured to perform the following processing: First, use the global average pooling operation to pool the temporal features and spatial characteristics Perform pooling dimensionality reduction to obtain global information modeling features and ; Then use 1×1 convolution to adjust the channel and ReLu activation function to improve nonlinear learning ability, and input the obtained feature input into the Sigmoid activation function to obtain the spatial calibration factor and time calibration factor At the same time, considering the interaction between the two features, the Hadamard product is used to perform interactive calibration modeling on the two features to enhance the interaction ability between the two features; Then, the interactively calibrated features are preliminarily calibrated and fused by element-wise addition, and convolution operation and ReLu activation function are used to obtain more fully fused features. ; Finally, the sigmoid activation function is used to further enhance the nonlinear ability while the temporal characteristics and spatial characteristics Further multiplexing calibration is performed to obtain the final feature that fully integrates time and space information .
10. The concrete seepage detection system based on multispectral video stream according to claim 1, characterized in that: In the detection head, a dedicated hyperparameter is introduced to control the classification loss term for single-category detection such as leakage.