An underwater structure damage identification method and system based on optical polarization imaging

By combining optical polarization imaging with a deep learning detection model and utilizing Stokes vector and polarization difference image generation physical attention mechanism, the problem of difficult identification of tiny cracks in high-turbidity underwater environments has been solved, achieving high-precision and robust underwater structural damage identification.

CN121482045BActive Publication Date: 2026-03-27ANHUI WATER CONSERVANCY TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In high-turbidity underwater environments, traditional detection methods struggle to effectively identify minute cracks in underwater structures, and existing deep learning models lack physical interpretability, resulting in high rates of missed detections and false alarms.

Method used

An optical polarization imaging-based method is adopted to acquire light intensity images through multi-angle polarization imaging, calculate Stokes vectors and polarization difference images, construct multi-channel input data, and introduce a deep learning detection model with physical attention mechanism and global attention mechanism to output damage information and risk level.

Benefits of technology

It significantly improves the detection accuracy and robustness of micro-cracks in high-turbidity underwater environments, achieving highly reliable intelligent monitoring, and is suitable for the safe operation and maintenance of critical infrastructure such as dams and bridges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482045B_ABST
    Figure CN121482045B_ABST
Patent Text Reader

Abstract

The application provides an underwater structure damage identification method and system based on optical polarization imaging, relates to the field of deep learning, and solves the technical problems of low crack detection precision caused by serious scattering and lack of physical prior guidance in the prior art in a high-turbidity underwater environment. The method comprises the following steps: acquiring light intensity images of an underwater target area at a plurality of preset polarization angles; calculating a plurality of physical quantities representing polarization states based on the light intensity images, and constructing multi-channel input data containing the plurality of physical quantities; inputting the multi-channel input data into a deep learning detection model to output damage information; and evaluating a risk level based on the damage information. The application is used in the process of underwater structure damage identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of deep learning, in particular to a method and system for identifying underwater structure damage based on optical polarization imaging. BACKGROUND

[0002] In water conservancy projects, underwater structures such as dams are prone to sub-millimeter cracks due to long-term erosion, corrosion, and stress from water flow. If not detected in time, it may cause leakage or even dam collapse. Traditional detection relies on manual diving or sonar, which has low efficiency, strong subjectivity, high safety risk, and low imaging signal-to-noise ratio in turbid water with turbidity >100 NTU. In recent years, although some studies have used deep learning such as YOLO series or polarization imaging for underwater detection, it is difficult to distinguish between real cracks and interference such as algae and sediments. Therefore, the present application provides a method and system for identifying underwater structure damage based on optical polarization imaging. SUMMARY

[0003] The present application provides a method for identifying underwater structure damage based on optical polarization imaging, which solves the technical problem of low crack detection accuracy in high turbidity underwater environment due to severe scattering and lack of physical prior guidance in the prior art. The present application also provides a system for identifying underwater structure damage based on optical polarization imaging.

[0004] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0005] In a first aspect, a method for identifying underwater structure damage based on optical polarization imaging is provided, comprising:

[0006] Obtaining light intensity images of a target underwater area at a plurality of preset polarization angles;

[0007] Calculating a multi-dimensional physical quantity representing the polarization state based on the light intensity images, and constructing multi-channel input data containing the multi-dimensional physical quantity; the multi-dimensional physical quantity includes Stokes vector and polarization difference image;

[0008] Inputting the multi-channel input data into a deep learning detection model, which is obtained based on the fusion of physical attention mechanism in YOLOv7-AC architecture; wherein the physical attention mechanism is an attention generation mechanism driven by external polarization imaging physical quantity, which is used to guide the deep learning detection model to pay attention to the structure damage sensitive area;

[0009] Spatial modulation of the intermediate feature map based on the physical attention mechanism and the global attention mechanism; the intermediate feature map is a multi-channel feature tensor generated by the deep learning detection model in the backbone network or the feature pyramid network;

[0010] The deep learning detection model outputs damage information according to the intermediate feature map of spatial modulation, and evaluates a risk level based on the damage information; wherein the damage information comprises a damage position and a damage parameter.

[0011] Based on the above technical solution, in the underwater structure damage identification method based on optical polarization imaging provided in the present application, in a high-turbidity underwater environment, traditional visual methods cannot effectively detect small cracks on the surface of the structure, mainly limited by strong scattering, low contrast and background interference, resulting in damage features being submerged, AI models being easily misled by noise, and high false negative and false positive rates. To solve this problem, the present application proposes an intelligent identification scheme that deeply integrates optical polarization physical characteristics and improves the deep learning architecture. The scheme first acquires light intensity images through multi-angle polarization imaging, calculates Stokes vectors and polarization difference images (PDI), and converts polarization information with clear physical meaning into multi-channel input; then a detection model based on YOLOv7-AC is constructed, and a physical attention mechanism is introduced, that is, a spatial attention mask is generated using PDI as an "optical prior" to guide the model to focus on the real damage area, realizing the detection of "cracks that can be seen but not seen by scattering"; at the same time, a global attention mechanism is combined to modulate the intermediate feature map in a double-path space, taking into account data adaptability and physical interpretability; finally, damage information including position and parameters is outputted and used to quantify the risk level. This technical solution breaks through the performance bottleneck of pure data-driven methods in harsh underwater environments, significantly improves detection accuracy, robustness and engineering practicability, and provides a highly reliable and deployable intelligent monitoring method for the safe operation and maintenance of key infrastructure such as dams and bridges.

[0012] In combination with the above first aspect, in a possible implementation manner, the physical attention mechanism generates a first spatial attention mask based on the polarization difference image, and modulates the intermediate feature map in the YOLOv7-AC architecture using the first spatial attention mask.

[0013] In combination with the above first aspect, in a possible implementation manner, the deep learning detection model comprises a feature fusion module, a polarization feature enhancement module and a multi-scale crack detection module.

[0014] The feature fusion module is configured to receive multi-channel input data and generate an initial feature map through a convolution layer, batch normalization and a nonlinear activation function.

[0015] The polarization feature enhancement module is configured to introduce a global attention mechanism; the global attention mechanism comprises a channel attention submodule and a spatial attention submodule.

[0016] The multi-scale crack detection module is configured to expand the structure of a backbone network or a feature pyramid network, and introduce a physical guided attention mechanism.

[0017] With reference to the first aspect above, in a possible implementation manner, the spatial modulation on the intermediate feature map based on the physical attention mechanism and the global attention mechanism comprises:

[0018] The physical attention mechanism generates a first spatial attention mask based on a polarization difference image calculated from a multi-angle polarization light intensity image;

[0019] The global attention mechanism generates a second spatial attention mask based on the intermediate feature map itself by modeling dependence of channel and spatial dimensions;

[0020] The deep learning detection model further comprises a learnable fusion module configured to weight and fuse the first spatial attention mask and the second spatial attention mask to obtain a comprehensive attention mask, and to multiply the comprehensive attention mask with the intermediate feature map element by element to perform spatial modulation on the intermediate feature map.

[0021] With reference to the first aspect above, in a possible implementation manner, the risk level is evaluated based on the damage information, which comprises:

[0022] An equivalent length L and a spatial position of a crack are extracted from the damage information (L, x, y), and a comprehensive risk index IRI is calculated in combination with a material degradation law, a stress concentration effect and a structure sensitivity prior, the IRI being defined as:

[0023]

[0024] wherein, L0 is a reference length, m is a material degradation coefficient, m>1, β is a proportional coefficient, β [0,1], is a spatial direction angle of the damage, is a stress concentration coefficient of the damage area under the damage and a material brittleness parameter k, , is a half crack length, , is a crack tip curvature radius, is a principal stress direction, is a structure sensitivity weight field;

[0025] The IRI is mapped to a discrete risk level as follows:

[0026]

[0027] wherein, , is a preset threshold, and . ​​​​

[0028] With reference to the first aspect, in a possible implementation manner, the structure-sensitive weight field comprises:

[0029] The structure-sensitive weight field is composed of prior knowledge, and an expression thereof is:

[0030] ;

[0031] wherein, is a distance from a spatial point to the nearest key component, the key component including a load-bearing beam, an anchoring area or a stress concentration area; and λ is an attenuation coefficient, used to represent a rate of attenuation of structure sensitivity with distance.

[0032] With reference to the first aspect, in a possible implementation manner, the method for generating the first spatial attention mask comprises

[0033] The polarized difference image is processed through at least one lightweight convolution layer, and a first spatial attention mask belonging to the interval [0, 1] is generated using a Sigmoid activation function.

[0034] With reference to the first aspect, in a possible implementation manner, the method for generating the second spatial attention mask comprises:

[0035] The channel attention sub-module performs global pooling on the intermediate feature map and processes it through a multi-layer perceptron to generate channel weights, and performs channel dimension weighting on the intermediate feature map based on the channel weights;

[0036] The spatial attention sub-module performs channel dimension aggregation on the intermediate feature map and processes it through spatial convolution with a large convolution kernel to generate spatial weights, and performs spatial dimension weighting on the channel-weighted feature map based on the spatial weights;

[0037] The second attention mask is generated through the channel dimension weighting and the spatial dimension weighting of the feature map.

[0038] With reference to the first aspect, in a possible implementation manner, the method for obtaining the multi-dimensional physical quantity comprises:

[0039] Based on the light intensity image, a Stokes vector is calculated, including total light intensity , horizontal-vertical polarization difference , polarization difference , right-left circular polarization difference ;

[0040] Based on the Stokes vector, a polarized difference image PDI is calculated through the formula PDI= .

[0041] In a second aspect, the application provides an underwater structure damage identification system based on optical polarization imaging, comprising: an acquisition module, a detection module and an evaluation module; wherein the acquisition module is configured to acquire light intensity images of a target underwater area at a plurality of preset polarization angles; based on the light intensity images, a plurality of physical quantities representing polarization states are calculated, and a multi-channel input data containing the plurality of physical quantities is constructed; the detection module is configured to input the multi-channel input data into a deep learning detection model, which is obtained based on a YOLOv7-AC architecture with a fusion of physical attention mechanism; the intermediate feature map is spatially modulated based on the physical attention mechanism and the global attention mechanism; the evaluation module is configured to output damage information based on the spatially modulated intermediate feature map and the deep learning detection model, and evaluate the risk level based on the damage information.

[0042] The application provides an underwater structure damage identification method and system based on optical polarization imaging, which can efficiently and accurately identify small cracks and other damages on the surface of concrete or steel structures in complex underwater environments such as high turbidity and low light, and realize automatic quantitative evaluation of damage location, geometric parameters and risk level. Specifically, the method acquires light intensity images of underwater targets through multi-angle polarization imaging, extracts polarization features with physical meaning by combining Stokes vectors and polarization difference images (PDI), effectively suppresses water scattering noise and enhances structural surface details; further, PDI is used as a physical prior to drive the physical attention mechanism in the deep learning model, which cooperates with the global attention mechanism to modulate feature representation, significantly improving the perception ability of real damage areas; finally, based on the detection results, the material degradation law, stress concentration effect and structure sensitivity prior are fused to calculate the comprehensive risk index (IRI), realizing the whole process of intelligent monitoring from "seeing" to "judging accurately and evaluating scientifically". The system can be deployed on a remotely operated vehicle (ROV) or an autonomous underwater vehicle (AUV), and is suitable for regular inspection, hidden trouble investigation and emergency response of key water conservancy and marine infrastructure such as dams, bridge piers, submarine tunnels and nuclear power cooling towers, greatly reducing the safety risks and costs of manual underwater detection, improving the intelligence, standardization and resilience level of infrastructure operation and maintenance, realizing accurate detection of water body, whether turbidity occurs, and measuring the process according to the turbidity level.

[0043] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in this application does not imply that all features and advantages can be achieved in any single embodiment. On the contrary, it can be understood that the description of a feature or a beneficial effect means that the specific technical feature, technical solution or beneficial effect is included in at least one embodiment. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Further, the technical features, technical solutions and beneficial effects described in this embodiment can be combined in any appropriate manner. Those skilled in the art will understand that the embodiments can be implemented without one or more specific technical features, technical solutions or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects can be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 A system architecture diagram of an underwater structure damage identification system based on optical polarization imaging is provided for embodiments of the present application.

[0045] Figure 2 A flowchart of an underwater structure damage identification method based on optical polarization imaging is provided for embodiments of the present application.

[0046] Figure 3 A detection process flowchart of a deep learning detection model is provided for embodiments of the present application. DETAILED DESCRIPTION

[0047] The technical solutions of the present application will be described in detail below with reference to embodiments. Obviously, the described embodiments are only a part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0048] An underwater structure damage identification method based on optical polarization imaging is provided by embodiments of the present application, which can be applied to an underwater structure damage identification system based on optical polarization imaging, as shown in Figure 1 The system includes an acquisition module, a detection module and an evaluation module.

[0049] The acquisition module is configured to acquire light intensity images of a target underwater area at a plurality of preset polarization angles; calculate a multi-dimensional physical quantity representing a polarization state based on the light intensity images; and construct multi-channel input data containing the multi-dimensional physical quantity.

[0050] The detection module is configured to input the multi-channel input data into a deep learning detection model, and the deep learning detection model is obtained based on a physical attention mechanism in a YOLOv7-AC architecture; and the deep learning detection model is configured to perform spatial modulation on an intermediate feature map based on the physical attention mechanism and a global attention mechanism.

[0051] The evaluation module is configured to output damage information by the deep learning detection model according to the spatially modulated intermediate feature map, and evaluate a risk level based on the damage information.

[0052] To solve the technical problem of low crack detection accuracy in a high-turbidity underwater environment due to severe scattering and lack of physical prior guidance in the prior art, an embodiment of the present application provides an underwater structure damage identification method based on optical polarization imaging, which comprises: acquiring light intensity images of a target underwater region at a plurality of preset polarization angles; calculating a multi-dimensional physical quantity representing a polarization state based on the light intensity images, and constructing multi-channel input data containing the multi-dimensional physical quantity; inputting the multi-channel input data into a deep learning detection model, and obtaining the deep learning detection model based on a physical attention mechanism in a YOLOv7-AC architecture; performing spatial modulation on an intermediate feature map based on the physical attention mechanism and a global attention mechanism; outputting damage information by the deep learning detection model according to the spatially modulated intermediate feature map, and evaluating a risk level based on the damage information. Therefore, by deeply coupling the Stokes vector physical characteristics of polarization light imaging and the global attention mechanism of the improved YOLOv7-AC model, the present application realizes high-precision, self-adaptive extraction and multi-scale identification of underwater dam cracks in turbid water environments, and is suitable for dam safety maintenance, hidden danger investigation, emergency response and structure health monitoring scenarios.

[0053] As shown in Figure 2 An embodiment of the present application provides an underwater structure damage identification method based on optical polarization imaging, which comprises:

[0054] S201, acquiring light intensity images of a target underwater region at a plurality of preset polarization angles.

[0055] In some implementations, a remotely operated vehicle (ROV) inspects the underwater part of the dam along a grid path (spacing 0.3 m) at a speed of 2 m / s, and a polarization camera collects multi-angle polarization images at 40 FPS. The present application collects 4 angles of 0°, 45°, 90° and 135°, and there are usually two implementation methods as follows:

[0056] Single sensor + rotating polaroid: every 4 frames form a complete polarization observation group (actual effective sampling rate ≈ 10 groups / s);

[0057] Multi-sensor / focal plane polarization camera: a single frame can simultaneously obtain 4 angles (at this time 40 FPS = 40 groups / s).

[0058] In this application, polarized image acquisition can adopt two ways: one is single sensor combined with rotating polarizer, low cost, high resolution, but need 4 frames to synthesize a set of polarization data, effective sampling rate is about 10 groups per second; Two is to use focal plane polarization camera (such as multi-sensor integration scheme), which can synchronously acquire 0°, 45°, 90°, 135° four polarization angle images in a single frame, realize 40 groups per second real-time sampling. The latter is especially suitable for the scene of ROV with 2 m / s high-speed inspection, which can effectively avoid the multi-frame misalignment caused by motion, improve the polarization information calculation accuracy and system robustness, and is more suitable for high reliability detection demand in underwater dynamic environment.

[0059] S202, calculate a multi-dimensional physical quantity representing a polarization state based on the light intensity image, and construct a multi-channel input data containing the multi-dimensional physical quantity.

[0060] Among them, the multi-dimensional physical quantity includes Stokes vector and polarization difference image.

[0061] It should be pointed out that constructing a multi-channel input data containing the multi-dimensional physical quantity means converting the light intensity images at different polarization angles obtained from the underwater target area by optical polarization imaging technology into multi-dimensional physical quantities that can reflect the polarization characteristics of the target object surface, and integrating these physical quantities into a multi-channel data format as the input of the deep learning model. The specific steps include:

[0062] A series of light intensity images are obtained by shooting the target area at multiple preset polarization angles (such as 0°, 45°, 90° and 135°). Then, based on these light intensity images, the Stokes vector is calculated, which contains total light intensity, horizontal-vertical polarization difference, diagonal polarization difference and right-left circular polarization difference information. In addition, polarization difference image (PDI) is also calculated according to the Stokes vector, which is used to quantify the degree of linear polarization at different positions.

[0063] The multi-dimensional physical quantities calculated above, including the components of Stokes vector and polarization difference image, are combined with the original light intensity image to form a multi-channel input data structure. For example, if RGB image plus polarization feature is used, a 7-channel data structure may be formed (RGB image occupies 3 channels, Stokes vector (S0, S1, S2, S3) occupies 4 channels, and PDI occupies 1 channel), so that the deep learning model can utilize color information, intensity information and polarization characteristics simultaneously for more accurate structure damage identification and analysis. This multi-channel input data is then sent to the deep learning detection model for training or prediction. 、 、

[0064] ​S203, inputting the multi-channel input data into a deep learning detection model, the deep learning detection model being obtained based on a physical attention mechanism fused in a YOLOv7-AC architecture.

[0065] The physical attention mechanism is an attention generation mechanism driven by an external polarization imaging physical quantity and used for guiding the deep learning detection model to focus on a structure damage sensitive area.

[0066] The physical attention mechanism generates a first spatial attention mask based on a polarization difference image and modulates an intermediate feature map in the YOLOv7-AC architecture by using the first spatial attention mask.

[0067] It should be noted that the essence of the physical attention mechanism is to convert optical physical laws (PDI) revealed by polarization imaging into learnable attention signals as "physical priori" to inject into the deep learning model to realize "seeing the scattering and not seeing the cracks".

[0068] S204, the deep learning detection model spatially modulates the intermediate feature map based on the physical attention mechanism and a global attention mechanism, the intermediate feature map being a multi-channel feature tensor generated by the deep learning detection model in a backbone network or a feature pyramid network.

[0069] S205, the deep learning detection model outputs damage information according to the spatially modulated intermediate feature map, and evaluates a risk level based on the damage information.

[0070] The damage information includes a damage position and a damage parameter.

[0071] Based on the technical scheme, in the underwater structure damage identification method based on optical polarization imaging provided in the application, in a high-turbidity underwater environment, a traditional visual detection method is difficult to effectively identify a tiny crack due to serious scattering and low contrast, and an existing deep learning model lacks physical interpretability and is easily disturbed by noise, resulting in high false negative and false positive rates. To solve this core problem, the application provides an integrated technical scheme of "polarization imaging-physical attention-risk quantification". First, a polarization camera is carried by an ROV to synchronously collect 0°, 45°, 90° and 135° angle light intensity images at 40 FPS, ensuring that complete polarization observation without motion artifacts can still be obtained at a high-speed inspection speed of 2 m / s; second, a multi-dimensional physical quantity (such as PDI) is constructed based on a Stokes vector and input into an improved YOLOv7-AC model as a multi-channel input; the key innovation is to introduce a physical attention mechanism to convert the PDI into a first spatial attention mask, guide the model to focus on the real damage area with strong polarization response, and realize "tell AI where to look with optical laws"; in combination with a global attention mechanism for double-path feature modulation, the sensitivity and positioning accuracy of micro and multi-scale cracks are significantly improved; finally, crack parameters are extracted from the detection results, and a comprehensive risk index (IRI) is calculated by combining fracture mechanics and structure priori, realizing a closed loop from "detection" to "evaluation". The scheme not only breaks through the imaging bottleneck of turbid water, but also realizes high-precision, high-robustness and interpretable intelligent monitoring, meeting the stringent requirements of key infrastructure such as dams for safe operation.

[0072] In a possible implementation manner of the embodiment of the application, S202 can be specifically explained as follows:

[0073] Based on the light intensity image, the Stokes vector is calculated, including total light intensity , horizontal-vertical polarization difference , diagonal polarization difference , right-left circular polarization difference .

[0074] Wherein, , , , respectively represent the light intensity measured by the polarization plate angle of 0°, 90°, 45° and 135°. is the right circularly polarized light intensity, is the left circularly polarized light intensity.

[0075] It should be noted that the total light intensity represents the total energy of the incident light and does not contain polarization information, and is used for normalization or illumination compensation; the horizontal-vertical polarization difference represents the dominance of linear polarization in the horizontal direction, and if it is positive, the horizontal polarization is stronger; the diagonal polarization difference represents the polarization characteristics in the diagonal direction, and determines the direction of the polarization principal axis; the right-left circular polarization difference represents the rotation direction of the light wave.

[0076] Based on the Stokes vector, the polarization difference image PDI is calculated by the formula PDI = (I0 - I90) / (I0 + I90). The polarization difference image PDI is calculated.

[0077] It should be noted that the PDI image is a single-channel grayscale image, and the higher the pixel value, the stronger the linear polarization degree of the point; in the underwater environment, scattered light (from the water body) is usually non-polarized or weakly polarized light -> PDI value is low; the structural surface reflected light (such as the edge of the concrete crack) has strong polarization characteristics -> PDI value is high; therefore, PDI can effectively enhance the structural surface features and suppress the background scattering noise, and significantly improve the visibility of the micro crack.

[0078] Based on the above technical solution, the traditional method is limited by underwater optical scattering, turbidity interference and computing resource limitation, and it is difficult to realize sub-millimeter crack precision detection. The present application deeply fuses the multi-dimensional polarization information of the Stokes vector (I0, I90, I45, I135) 、 、 、 and the adaptive deep learning framework of YOLOv7-AC to construct a physical-guided intelligent detection system for high-turbidity water environment. By collecting light intensity images at four polarization angles (0°, 90°, 45°, 135°), the total light intensity , the horizontal-vertical polarization difference , the diagonal polarization difference and the circular polarization difference are calculated, and then the polarization difference image (PDI) is generated. This process fully utilizes the physical law that light preserves polarization characteristics when reflecting at different medium interfaces, effectively distinguishes between structural surface reflected light (strong polarization) and water body scattered light (weak polarization), and significantly improves the signal-to-noise ratio. As a key physical prior, PDI not only enhances the visibility of micro cracks, but also provides interpretable and robust input features for subsequent deep learning models. This method breaks through the bottleneck of traditional visual detection in turbid underwater environment, realizes the paradigm shift from "data-driven" to "physical enhancement", and is not only suitable for water conservancy facilities such as dams and reservoirs, but also can be extended to intelligent monitoring of key infrastructure such as bridge underwater structures and cooling towers of nuclear power plants, responding to the core demand of sustainable development goal (SDG9) for improving infrastructure resilience and intelligent operation and maintenance, and has wide technical popularization value and engineering application prospect.

[0079] In a possible implementation of the embodiment of the present application, S203 can be specifically implemented as shown in the following. Figure 3

[0080] The deep learning detection model comprises a feature fusion module, a polarization feature enhancement module and a multi-scale crack detection module.

[0081] The feature fusion module is configured to receive multi-channel input data and generate an initial feature map through a convolution layer, batch normalization and a nonlinear activation function.

[0082] For example, assuming that the input is an RGB image containing a building surface and a corresponding polarization difference image (PDI). The feature fusion module processes the two types of input respectively, and fuses them together through additional convolution operations to form a comprehensive feature representation.

[0083] The polarization feature enhancement module is configured to introduce a global attention mechanism; the global attention mechanism comprises a channel attention submodule and a spatial attention submodule.

[0084] It should be noted that when detecting the surface cracks of an underwater object, the channel attention can emphasize those features (such as edge or texture information) that are particularly useful for identifying cracks, and the spatial attention can help highlight the specific locations where cracks may exist.

[0085] The multi-scale crack detection module is configured to expand the structure of a backbone network or a feature pyramid network, and introduce a physically guided attention mechanism to enhance the feature response to potential structural damage areas.

[0086] It should be noted that in order to effectively detect various cracks from small to large, the multi-scale crack detection module can combine deep features at low resolution (for large crack detection) and shallow features at high resolution (for small crack detection). The physically guided attention mechanism can use physical quantities (such as the polarization information mentioned earlier) to improve the response to real cracks and reduce false positives.

[0087] Based on the above technical solution, in a high turbidity underwater environment, traditional visual methods are difficult to effectively detect the tiny cracks of water conservancy facilities such as dams, mainly limited by strong scattering interference, low image contrast, and blurred damage features, resulting in high missed detection rate and poor robustness. To solve this key technical bottleneck, the present application proposes an intelligent detection architecture that deeply couples the physical characteristics of polarized light imaging with an improved YOLOv7-AC model. This scheme uses Stokes vector (S) to represent the polarization information of the input image, and uses the improved YOLOv7-AC model to detect the cracks in the image. 、 、 ​) generates a polarization difference image (PDI) as a physical prior input feature fusion module to construct an initial feature map rich in structural information in cooperation with an RGB image; a polarization feature enhancement module introduces a global attention mechanism (GAM) to adaptively strengthen edge and texture features related to cracks using channel and spatial attention, and focuses on key areas; a multi-scale crack detection module further combines a physically guided attention mechanism to accurately respond to complex morphologies from millimeter-level micro-cracks to macro-cracks at different scales. This technical solution not only overcomes the constraints of turbid water on imaging quality, but also realizes the dual enhancement of "physical law + data-driven", significantly improving detection accuracy, generalization ability and real-time performance. Therefore, it is particularly suitable for high-reliability scenarios such as dam safety maintenance, hidden danger investigation, emergency response and structural health monitoring (SHM), and provides a practical and scalable technical path for intelligent operation and maintenance of underwater infrastructure.

[0088] In a possible implementation manner of the embodiment of the application, the S204 can be specifically described as follows:

[0089] The intermediate feature map is spatially modulated based on the physical attention mechanism and the global attention mechanism, and the specific modulation process is as follows:

[0090] The physical attention mechanism generates a first spatial attention mask based on the polarization difference image.

[0091] Specifically, the generation method of the first spatial attention mask is as follows:

[0092] The polarization difference image is processed through at least one lightweight convolution layer, and a Sigmoid activation function is used to generate a first spatial attention mask belonging to the interval [0, 1].

[0093] The global attention mechanism generates a second spatial attention mask based on the intermediate feature map itself by modeling the dependence of the channel and spatial dimensions.

[0094] Specifically, the generation method of the second spatial attention mask is as follows:

[0095] The channel attention submodule globally pools the intermediate feature map and processes it through a multi-layer perceptron to generate channel weights, and weights the intermediate feature map in the channel dimension based on the channel weights.

[0096] It should be noted that the original intermediate feature map is weighted and summed according to the channel weights to emphasize important feature channels.

[0097] The spatial attention submodule aggregates the intermediate feature map in the channel dimension and processes it through spatial convolution with a large kernel to generate spatial weights, and weights the channel-weighted feature map in the spatial dimension based on the spatial weights.

[0098] It should be noted that the feature map weighted by the channel is further weighted by the spatial dimension using the spatial weight.

[0099] The feature map weighted by the channel dimension and the spatial dimension generates a second attention mask.

[0100] It should be noted that assuming that the intermediate feature map is a multi-channel image, the channel attention sub-module can find that some channels are particularly important for identifying cracks, and therefore give these channels higher weights. The spatial attention sub-module can notice that certain specific regions are more likely to contain cracks than other regions, and therefore assign higher spatial weights to these regions.

[0101] The deep learning detection model further includes a learnable fusion module for weighted fusion of the first spatial attention mask and the second spatial attention mask to obtain a comprehensive attention mask, and element-wise multiplication of the comprehensive attention mask and the intermediate feature map is performed to spatially modulate the intermediate feature map.

[0102] Based on the above technical solution, in a high-turbidity underwater environment, the crack features are weak and easily covered by scattered noise, and a single attention mechanism cannot balance physical priori and data adaptability, resulting in insufficient robustness of the detection model and high false alarm rate. To solve this problem, the present application proposes a dual-path spatial modulation strategy that fuses a physical attention mechanism and a global attention mechanism (GAM). The physical attention mechanism generates a first spatial attention mask based on a polarization difference image (PDI), directly introduces the optical physics law of highlighting potential damage areas with strong polarization response, has clear interpretability and anti-interference ability; while the global attention mechanism starts from the intermediate feature map itself, emphasizes the feature channels (such as edges and textures) that are critical to crack identification through channel attention, and then focuses on high-probability damage locations through spatial attention, realizing data-driven context awareness. The two are integrated by a learnable fusion module to form a comprehensive attention mask that conforms to the physical law and adapts to the task semantics, and accurately modulates the feature map. This scheme effectively overcomes the generalization defects of purely data-driven methods in low signal-to-noise ratio scenarios, while avoiding the problem of insufficient adaptability of purely physical methods to complex backgrounds. Thus, the sensitivity and positioning accuracy of the model for small and multi-scale cracks in turbid water bodies are significantly improved, meeting the urgent needs of key infrastructure such as dams for high-reliability and high-robustness intelligent monitoring.

[0103] In a possible implementation manner of the embodiment of the present application, the S205 body can be specifically described as follows:

[0104] The equivalent length L and the spatial position of the crack are extracted from the damage information ), and combining the material degradation law, stress concentration effect and structure sensitivity prior, the comprehensive risk index IRI is calculated, which is defined as:

[0105] ;

[0106] where L is the equivalent length of the actual crack length or the length of the area equivalent to a straight line, is the reference length for normalization, m is the material degradation coefficient, m > 1; is the spatial orientation angle of the damage, and the stress concentration coefficient of the damage area under the material brittleness parameter k; β is the proportional coefficient, β [0, 1], which is used to control the influence weight of the stress concentration term on the total risk, to avoid over amplification; is the structure sensitivity weight field;

[0107] It should be pointed out that, is a function representing the stress concentration degree under the combined action of damage direction and material characteristics, ; wherein, is the half crack length, , is the crack tip curvature radius, is the principal stress direction, which is the direction of pure tension or pure compression in the material, and in the structure damage assessment, it is the key physical basis for judging whether the crack is in a high risk expansion state;

[0108] When , , is the maximum, the crack has a tendency to expand along the principal tensile direction, and the risk is higher; when is perpendicular to , , , the crack has little expansion tendency.

[0109] The structure sensitivity weight field is composed of prior knowledge, and its expression is:

[0110] ;

[0111] wherein, is the distance from the spatial point to the nearest key component, including the load-bearing beam, the anchoring area or the stress concentration area; λ is the attenuation coefficient, which is used to represent the rate of attenuation of structure sensitivity with distance.

[0112] If the damage is closer to the key component, W is larger and the risk is higher, and if the damage is farther from the key component, W is smaller and the risk is lower.

[0113] It is noted that the attenuation scale coefficient is pre-set according to the structure type, and for a concrete hydraulic structure, the value is 0.5 meters; for a steel structure bridge, the value is 0.3 meters. The value is determined based on the stress diffusion range in structural mechanics and engineering practice experience, and can also be further optimized through historical failure data or finite element simulation.

[0114] The IRI of the application is a mathematical expression of a comprehensive risk index, which is used to quantify the potential harm degree of the detected damage (such as cracks) in the underwater structure (such as dam, bridge, tunnel). It integrates four dimensions of material science, fracture mechanics, structural engineering and position sensitivity.

[0115] Map the IRI to discrete risk levels:

[0116] ;

[0117] wherein, , is a preset threshold, and .

[0118] Based on the above technical solutions, when evaluating the crack damage in the underwater structure such as dam, bridge or tunnel, the traditional single-dimensional analysis method often cannot fully reflect the complexity of the potential harm. Therefore, the application proposes the concept of comprehensive risk index (IRI), which integrates the material degradation law, stress concentration effect and structural sensitivity prior information to comprehensively quantify the risk of cracks and other damages. The equivalent length L and the spatial position As a basic input parameter, it provides information of the crack size and its location. Combined with the material degradation coefficient m, the additional risk caused by material aging can be adjusted. Next, The stress concentration coefficient considers the interaction between the crack direction and the material brittleness characteristics, especially when the crack direction is close to the principal stress direction, the expansion risk increases significantly. This method based on fracture mechanics can accurately capture the possibility and trend of crack propagation. Further, the structural sensitivity weight field ​, according to the distance between the crack and the key component to adjust the risk assessment. This method not only considers the influence of the local stress environment, but also takes into account the global characteristics of the structure, emphasizing the importance of key positions. For different types of structures, by setting appropriate attenuation coefficients λ, it can be flexibly adapted to the specific requirements of different types of engineering structures. The above factors are integrated into IRI and mapped to discrete risk levels, realizing the transformation from quantitative analysis to qualitative evaluation, which is convenient for decision-makers to quickly understand the severity of damage and develop appropriate maintenance plans. The advantage of this method is that it provides a multi-dimensional and systematic framework that not only considers physical factors but also incorporates practical considerations in structural safety assessment. In this way, not only does it improve the accuracy of risk assessment, but it also points out key areas for subsequent repair work, helping to optimize resource allocation, improve maintenance efficiency, and reduce potential safety hazards.

[0119] Some data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is obtained by software simulation of a large amount of collected data to obtain a formula closest to the actual situation. The preset parameters and preset thresholds in the formula are set by a person skilled in the art according to the actual situation or obtained by a large amount of data simulation.

Claims

1. An underwater structure damage identification method based on optical polarization imaging, characterized in that, The method comprises the following steps: acquiring light intensity images of an underwater target area under multiple preset polarization angles; calculating a multi-dimensional physical quantity representing a polarization state based on the light intensity images, and constructing multi-channel input data containing the multi-dimensional physical quantity; the multi-dimensional physical quantity comprises a Stokes vector and a polarization difference image; the method for obtaining the multi-dimensional physical quantity comprises: Based on the light intensity image, calculate Stokes vector, including total light intensity , horizontal-vertical polarization difference , polarization difference , right-left circular polarization difference ; Based on the Stokes vector, the polarization difference image PDI is calculated by the formula PDI = 2 * (S2 - S3) The polarization difference image PDI is calculated. inputting the multi-channel input data into a deep learning detection model, wherein the deep learning detection model is obtained based on a physical attention mechanism fused in a YOLOv7-AC architecture; the physical attention mechanism is an attention generation mechanism driven by external polarization imaging physical quantities and used for guiding the deep learning detection model to focus on a structure damage sensitive area; the physical attention mechanism generates a first spatial attention mask based on the polarization difference image, and modulates an intermediate feature map in the YOLOv7-AC architecture by using the first spatial attention mask; the deep learning detection model spatially modulates the intermediate feature map based on the physical attention mechanism and a global attention mechanism; the intermediate feature map is a multi-channel feature tensor generated by the deep learning detection model in a backbone network or a feature pyramid network; the deep learning detection model outputs damage information according to the spatially modulated intermediate feature map, and evaluates a risk level based on the damage information; the damage information comprises a damage position and a damage parameter. 2.The method according to claim 1, characterized in that, the deep learning detection model comprises a feature fusion module, a polarization feature enhancement module, and a multi-scale crack detection module; the feature fusion module is configured to receive the multi-channel input data and generate an initial feature map through a convolution layer, a batch normalization, and a nonlinear activation function; the polarization feature enhancement module is configured to introduce the global attention mechanism; the global attention mechanism comprises a channel attention submodule and a spatial attention submodule; the multi-scale crack detection module is configured to expand a structure of the backbone network or the feature pyramid network, and introduce a physically guided attention mechanism. 3.The method of claim 1, wherein, the deep learning detection model spatially modulates the intermediate feature map based on the physical attention mechanism and the global attention mechanism, which comprises: the physical attention mechanism generates a first spatial attention mask based on a polarization difference image calculated from multi-angle polarization light intensity images; the global attention mechanism generates a second spatial attention mask based on the intermediate feature map itself through modeling of dependence between channel and spatial dimensions; the deep learning detection model further comprises a learnable fusion module configured to weight and fuse the first spatial attention mask and the second spatial attention mask to obtain a comprehensive attention mask, and multiply the comprehensive attention mask and the intermediate feature map element by element to spatially modulate the intermediate feature map.

4. The method according to claim 1, wherein, the evaluation of the risk level based on the damage information comprises: extracting equivalent length L and spatial position of the crack from the damage information ), and combining material degradation law, stress concentration effect and structure sensitivity prior, calculating comprehensive risk index IRI, the IRI is defined as: ; wherein, is a reference length, m is a material degradation coefficient, m > 1, is a spatial direction angle of damage, is a stress concentration coefficient of the damage area under a material brittleness parameter k, , is a half crack length, , is a crack tip curvature radius, is a principal stress direction, is a structure sensitivity weight field, is a proportional coefficient, ; mapping the IRI to a discrete risk level; ; wherein , is a predetermined threshold value, and .​ 5. The method according to claim 4, wherein the method is characterized by, the structure sensitivity weight field comprises: the structure sensitivity weight field is composed of prior knowledge, and its expression is: ; wherein, for a spatial point distance to the nearest critical component, including a load-bearing beam, an anchoring zone or a stress concentration area; λ is an attenuation coefficient, used to characterize the rate of decay of the structural sensitivity with distance.

6. The method according to claim 3, wherein the method is characterized by, the method for generating the first spatial attention mask comprises The polarization difference image is processed by at least one lightweight convolution layer, and a first spatial attention mask belonging to the interval [0, 1] is generated using a Sigmoid activation function.

7. The method according to claim 3, wherein the method is characterized by, The method for generating the second spatial attention mask comprises: The channel attention sub-module performs global pooling on the intermediate feature map and processes it through a multi-layer perceptron to generate channel weights, and performs channel dimension weighting on the intermediate feature map based on the channel weights; The spatial attention sub-module performs channel dimension aggregation on the intermediate feature map and processes it through spatial convolution with a large kernel to generate spatial weights, and performs spatial dimension weighting on the channel-weighted feature map based on the spatial weights; The feature map weighted by the channel dimension and the spatial dimension generates a second attention mask.

8. An underwater structure damage identification system based on optical polarization imaging, operating based on the method of any one of claims 1-7, characterized in that, The detection module, the acquisition module and the evaluation module are connected. The acquisition module is configured to acquire light intensity images of an underwater target area at a plurality of preset polarization angles. A plurality of physical quantities representing polarization states are calculated based on the light intensity images, and multi-channel input data containing the plurality of physical quantities are constructed. The detection module is configured to input the multi-channel input data into a deep learning detection model, wherein the deep learning detection model is obtained based on a physical attention mechanism fused in a YOLOv7-AC architecture. The intermediate feature map is spatially modulated based on the physical attention mechanism and a global attention mechanism. The evaluation module is configured to output damage information from the deep learning detection model according to the spatially modulated intermediate feature map, and evaluate a risk level based on the damage information.

Citation Information

Patent Citations

  • Material identification method of polarization attention guiding mechanism

    CN118865338A

  • Vehicle point cloud wind resistance coefficient prediction method and system based on multi-scale learning and convolution

    CN120409269A