A method, system, device, and medium for offshore infrastructure surface defect detection

By using a hybrid radiation field model to extract features and fuse modeling of marine infrastructure image sequence data, the problem of low detection rate in the inspection of large facilities was solved, achieving efficient defect detection and 3D reconstruction, and improving the consistency and automation level of inspection.

CN120931647BActive Publication Date: 2026-02-06HAINAN RES INST OF ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511460832.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-02-06
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing technologies for defect detection in large-scale offshore infrastructure are limited by factors such as extreme viewing angles and shaking, resulting in low detection rates and a lack of consistency and interactivity.

Method used

A hybrid radiation field model is used to extract features from image sequence data. Geometric, color and semantic features are independently encoded and decoded to construct a 3D model with a specified spatial resolution. The model is then fused through a multi-level encoding module and an MLP decoding module to achieve end-to-end defect detection.

Benefits of technology

It improves the detection rate and consistency of the detection, realizes the automated operation of global detection, has the ability to view local details and global 3D reconstruction, reduces labor costs, and provides reliable protection in extreme environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931647B_ABST
    Figure CN120931647B_ABST
Patent Text Reader

Abstract

The application provides a marine infrastructure surface defect detection method, system, device and medium, and belongs to the technical field of defect detection. The method comprises the following steps: recognizing image sequence data of a marine infrastructure to obtain a defect segmentation result image sequence; performing feature extraction and feature space coding and decoding on the pose image sequence data and the defect segmentation result image sequence to obtain defect categories, global colors, local colors and depth information of the feature space; combining the spatial resolution of an actual scene with the defect categories, global colors, local colors and depth information of the independent feature space to extract a marine infrastructure surface three-dimensional model and obtain the position and quantity of marine infrastructure surface defects. The application can render a local defect detection image fused with global detection results by setting a global arbitrary view angle, and can extract facility three-dimensional semantic defect models with different precisions, position different types of defects and determine the quantity of defects by setting a spatial resolution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of defect detection, and particularly relates to a method, system, device and medium for detecting surface defects of offshore infrastructure. BACKGROUND

[0002] The surface defect detection of large offshore infrastructure such as dams is very important. With the increase of the service life, various types of defects often appear on the surface of the facility, which requires regular inspection and maintenance of the defects on the surface of the facility to ensure the smooth and safe operation of the infrastructure. With the development of computer vision technology, defect detection and three-dimensional reconstruction technology have made great progress. At present, the method for detecting defects on the surface of large infrastructure is as follows: a robot carries a vision sensor to collect surface image data, then a deep neural network is used to segment the collected pictures, and a multi-view geometry method such as SLAM or SFM method is used to realize three-dimensional reconstruction, and the results of defect semantic segmentation are registered to the three-dimensional reconstruction model.

[0003] The existing defect detection of large infrastructure such as large dams and bridges is a one-way process, still in the transition stage from theoretical research to practical application, and there are the following problems in the actual application process:

[0004] Due to the large scene, low accessibility and variable camera shooting angle of offshore infrastructure, the images collected in the actual scene of the surface of large infrastructure often have extreme view angle, shaking and other situations, and cannot be as standardized as the defects in the conventional data set, so the detection rate of the deep learning model for image defects in the actual scene of large infrastructure is low. SUMMARY

[0005] In order to overcome the deficiencies of the prior art, the present application provides a method for detecting surface defects of offshore infrastructure, comprising the following steps:

[0006] Obtain image sequence data of offshore infrastructure, and solve the camera pose of the image sequence data to obtain pose image sequence data;

[0007] Preliminarily identify the defects of offshore infrastructure in the image sequence data to obtain a defect segmentation result image sequence;

[0008] Extract features from the pose image sequence data and the defect segmentation result image sequence to obtain geometric features, color features and defect semantic features, and independently encode and decode the geometric features, color features and defect semantic features in a feature space to obtain defect categories, global color, local color and depth information in the independent feature space;

[0009] The spatial resolution of the actual scene is combined with the defect category, global color, local color and depth information of the independent feature space to output feature vectors under different scale mappings at a specified hierarchical resolution, and the feature vectors under different scale mappings are fused and modeled to obtain a marine infrastructure surface three-dimensional model at a specified spatial resolution, and the position and quantity of the marine infrastructure surface defects are extracted through the marine infrastructure surface three-dimensional model.

[0010] Preferably, the mixed radiation field model is used to obtain the pose image sequence data and the defect category, global color, local color and depth information of the defect segmentation result image sequence, and fusion modeling is performed to obtain a marine infrastructure surface three-dimensional model at a specified spatial resolution; the mixed radiation field model includes a multi-level geometric information explicit encoding module, a multi-level color information encoding module, a camera pose color spherical harmonic encoding module and a multi-head MLP decoding module;

[0011] The multi-head MLP decoding module includes local color MLP decoding, global color MLP decoding, depth MLP decoding and semantic MLP decoding.

[0012] Preferably, the pose image sequence data and the defect segmentation result image sequence are feature extracted to obtain geometric features, color features and defect semantic features, and the geometric features, color features and defect semantic features are independently encoded in a feature space, including:

[0013] The multi-level geometric information explicit encoding module is used to compress the real space coordinates of the pose image sequence data and the defect segmentation result image sequence into an implicit neural field range, the spatial resolution of the implicit neural field is hierarchically divided to obtain a cube space of each level, each vertex is assigned an independent index, the geometric information and defect information of each space point are extracted by linear interpolation of the vertices of multiple cube blocks of the cube space to obtain feature space encoding features of the geometric features and the defect semantic features; wherein the cube spaces of the same level are composed of cube blocks of the same size, and each cube block vertex has a feature encoding vector; the implicit neural field refers to a continuous cube space with a feature surface in a neural network;

[0014] The multi-level color information encoding module is used to extract the global color information of the pose image sequence data and the defect segmentation result image sequence to obtain feature space encoding features of the global color features;

[0015] The camera pose color spherical harmonic encoding module is used to extract the local camera pose view and local color information of the pose image sequence data and the defect segmentation result image sequence to obtain feature space encoding features of the local color features.

[0016] Preferably, the feature space encoding features are decoded, specifically:

[0017] Decode and render the feature space encoding features of the global color features using global color MLP decoding;

[0018] Decode and render the feature space encoding features of the local color features using local color MLP combined with camera pose color spherical harmonic encoding module, to obtain local color;

[0019] Set the spatial resolution of the mixed radiation field model output, use the depth MLP to decode the spatial depth information of the mixed radiation field, render the depth features of the three-dimensional model, and obtain the depth information of the spatial grid voxel corresponding to the resolution;

[0020] Set the spatial resolution of the mixed radiation field model output, use the semantic MLP to decode the spatial defect semantic information of the mixed radiation field, render the defect label information of the semantic image of the three-dimensional model at any viewing angle, and extract the spatial grid voxel semantic label information corresponding to the resolution as the defect category.

[0021] Preferably, the image sequence data of the offshore infrastructure includes RGB-D image data obtained by a color depth camera, image data obtained by a binocular camera, monocular camera image sequence with depth information obtained by a neural network, and point cloud and image data obtained by a calibrated laser radar and camera.

[0022] Preferably, the defect segmentation result image sequence includes crack images, erosion images, pitted surface images, or artificial repair category images.

[0023] Preferably, before obtaining the pose image sequence data and the defect category, global color, local color and depth information of the defect segmentation result image sequence by the mixed radiation field model, it further includes constructing a loss function, training the mixed radiation field model by the loss function, and the loss function is as follows:

[0024] ;

[0025] Wherein represents the sum of the loss function, represents the sum of the loss function in the global coarse optimization stage, represents the loss function in the local fine optimization stage, 、 and respectively represent the weights of the corresponding global color, semantic and local color loss, and the local color loss and the global color loss both adopt loss mode, and the semantic uses cross-entropy loss; wherein the geometric loss including surface normal eikonal loss and tagged surface point loss .

[0026] The application further provides a marine infrastructure surface defect detection system, comprising:

[0027] a data acquisition module, configured to acquire image sequence data of the marine infrastructure, and to obtain pose image sequence data by solving camera pose of the image sequence data;

[0028] an initial defect detection module, configured to preliminarily identify defects of the marine infrastructure in the image sequence data, and to obtain a defect segmentation result image sequence;

[0029] a feature rendering module, configured to extract geometric features, color features and defect semantic features from the pose image sequence data and the defect segmentation result image sequence, and to independently encode and decode the geometric features, the color features and the defect semantic features to obtain defect categories, global colors, local colors and depth information in independent feature spaces;

[0030] a defect scene reconstruction module, configured to combine spatial resolution of an actual scene with the defect categories, the global colors, the local colors and the depth information in the independent feature spaces, to output feature vectors under different scale mappings at a specified hierarchical resolution, and to fuse and model the feature vectors under the different scale mappings to obtain a marine infrastructure surface three-dimensional model at a specified spatial resolution, and to extract positions and quantities of marine infrastructure surface defects through the marine infrastructure surface three-dimensional model.

[0031] The application further provides a computer device, comprising a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the marine infrastructure surface defect detection method.

[0032] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is adapted to be loaded by a processor to execute the marine infrastructure surface defect detection method.

[0033] The defect detection method provided by the application has the following beneficial effects:

[0034] The application proposes that the mixed radiation field method is applied to the surface detection of large-scale infrastructures such as dams, feature extraction is performed on the pose image sequence data and the defect segmentation result image sequence, geometric features, color features and defect semantic features are obtained, and independent feature space coding and decoding are performed, the problems of poor consistency and weak interactivity of the original traditional detection and inspection method are solved, the phenomenon that the detection rate is lower than the data set due to the perspective and noise interference in the large-scale scene surface detection process is improved, the spatial resolution of the actual scene is set, the three-dimensional model of the offshore infrastructure surface with the specified spatial resolution is extracted, the position and quantity of the offshore infrastructure surface defects are obtained, an end-to-end interactive global detection scheme is formed, the image acquisition, defect segmentation, pose solving and overall model training can all realize automatic operation, the output result detection rate is improved, and the local detail viewing and global three-dimensional reconstruction capabilities are simultaneously provided, manpower cost is saved for the related industry, and reliable guarantee is provided for the detection and reconstruction tasks in extreme environments. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present application and the design scheme thereof, the drawings required by the present embodiments will be briefly introduced as follows. The drawings in the following description are only part of the embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0036] Figure 1 Explicit-implicit space coding schematic diagram for multi-level specified resolution features of the present application;

[0037] Figure 2 Structure and principle diagram of a double-branch space explicit-implicit neural field unified model;

[0038] Figure 3 Global-local training optimization flowchart of the explicit-implicit model;

[0039] Figure 4 End-to-end large-scale facility surface defect detection and reconstruction model usage flowchart;

[0040] Figure 5 Defect segmentation and artificial marking of the actual scene are compared, wherein, Figure 5 (a) of the first group of comparison charts, Figure 5 (b) of the second group of comparison charts, Figure 5 (c) of the third group of comparison charts, Figure 5 (d) of the fourth group of comparison charts, wherein Figure 5 (a1), (b1), (c1) and (d1) in (a), (b), (c) and (d) are original images of the four groups of comparison charts, respectively, Figure 5(a2), (b2), (c2), (d2) in FIG. 4 are four sets of artificial marking maps of the contrast graphs respectively, Figure 5 (a3), (b3), (c3), (d3) in FIG. 5 are four sets of network segmentation maps of the contrast graphs respectively;

[0041] Figure 6 For actual scene defect segmentation, rendering of defect semantic image and artificial marking comparison; wherein, Figure 6 (a) in FIG. 6 is the first set of contrast graphs, Figure 6 (b) in FIG. 7 is the second set of contrast graphs, Figure 6 (c) in FIG. 8 is the third set of contrast graphs, Figure 6 (d) in FIG. 9 is the fourth set of contrast graphs, wherein Figure 6 (a1), (a4), (b1), (b4), (c1), (c4), (d1), (d4) in FIG. 10 are four sets of original images of the contrast graphs respectively, Figure 6 (a2), (b2), (c2), (d2) in FIG. 11 are four sets of rendering maps after global mixed radiance field learning respectively, Figure 6 (a5), (b5), (c5), (d5) in FIG. 12 are four sets of rendering maps of U-net type neural network respectively, Figure 6 (a3), (b3), (c3), (d3) in FIG. 13 are four sets of local defect semantic images after global mixed radiance field learning respectively, Figure 6 (a6), (b6), (c6), (d6) in FIG. 14 are four sets of detection results output by U-net type neural network respectively;

[0042] Figure 7 is a generated grid model segmented by a deep neural network; wherein Figure 7 (a) in FIG. 15 is a global color model, Figure 7 (b) in FIG. 16 is a global semantic model;

[0043] Figure 8 is a collected dam flow passage image sequence; wherein Figure 8 (a) in FIG. 17 is a display of a collected surface image sequence, Figure 8 (b) in FIG. 18 is a display of a marked sequence image;

[0044] Figure 9 is a reconstructed three-dimensional model; wherein Figure 9 (a) in FIG. 19 is a color global model, Figure 9 (b) in FIG. 20 is a semantic global model;

[0045] Figure 10 is a sparse artificial marking driven mixed radiance field model rendering result; wherein, Figure 10 (a) in FIG. 21 is the first set of contrast graphs, Figure 10 (b) in FIG. 22 is the second set of contrast graphs, Figure 10(c) is a third group of contrast images, Figure 10 (d) is a fourth group of contrast images, wherein Figure 10 (a1), (a4), (b1), (b4), (c1), (c4), (d1), (d4) in (a) are original images of four groups of contrast images, respectively, Figure 10 (a2), (b2), (c2), (d2) in (b) are rendered images after learning of global mixed radiation field after training, respectively, Figure 10 (a5), (b5), (c5), (d5) in (c) are rendered images of U-net type neural network, respectively, Figure 10 (a3), (b3), (c3), (d3) in (d) are local defect semantic images after learning of global mixed radiation field after training, respectively, Figure 10 (a6), (b6), (c6), (d6) in (e) are initial labeling results, respectively. DETAILED DESCRIPTION

[0046] In order for those skilled in the art to better understand the technical solutions of the present application and to implement them, the present application will be described in detail below in conjunction with the drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and cannot be used to limit the protection scope of the present application.

[0047] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential" and the like indicate the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the technical solutions of the present application and simplifying the description, and is not intended to indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0048] In addition, the terms "first", "second" and the like are only for descriptive purposes and cannot be understood as indicating or implying relative importance. In the description of the present application, it should be noted that unless otherwise specified or limited, the terms "connected", "connected" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances. In the description of the present application, unless otherwise stated, the meaning of "multiple" is two or more, which will not be described in detail here.

[0049] Example 1:

[0050] This invention provides a method for detecting surface defects in marine infrastructure, specifically as follows: Figure 4 As shown, it includes the following steps:

[0051] Step S1: Acquire two-dimensional image data of the marine infrastructure, that is, acquire image sequence data of the marine infrastructure, and obtain pose image sequence data by camera pose calculation of the image sequence data.

[0052] The image sequence data includes RGB-D image data obtained by a color depth camera, image data obtained by a binocular camera, and point cloud and image data obtained by a calibrated LiDAR and camera.

[0053] Step S2: Input the two-dimensional image data into the neural network model to perform preliminary identification of defects in the marine infrastructure in the image sequence data, and obtain the defect segmentation result image sequence.

[0054] The defect segmentation result image sequence includes crack images, erosion images, pitted surface images, or artificially repaired images.

[0055] Step S3: Extract features from the pose image sequence data and the defect segmentation result image sequence to obtain geometric features, color features and defect semantic features. Then, encode and decode the geometric features, color features and defect semantic features independently in the feature space to obtain the defect category, global color, local color and depth information in the independent feature space.

[0056] The spatial resolution of the implicit neural field is hierarchically divided to obtain multiple hierarchical spaces. Voxel blocks are extracted from each hierarchical space, and each voxel block is assigned an independent index to obtain a mixed radiation field model.

[0057] The hybrid radiation field model is used to obtain the defect category, global color, local color and depth information of the pose image sequence data and defect segmentation result image sequence. The hybrid radiation field model includes a multi-level geometric information explicit encoding module, a multi-level color information encoding module, a camera pose color spherical harmonic encoding module and a multi-head MLP decoding module. The multi-head MLP decoding module includes local color MLP decoding, global color MLP decoding, depth MLP decoding and semantic MLP decoding.

[0058] To achieve unified 3D reconstruction (implicit 3D model representation) and local viewpoint generation, it is urgent to improve the ability of implicit neural fields to represent 3D models. This is because converting implicit neural fields into explicit 3D models (point clouds, occupied voxels, SDF (signed distance field) fields, etc.) requires hierarchical partitioning of the implicit field before extracting feature representations from the voxels.

[0059] The application introduces a mixed radiation field model, the multi-layer resolution setting of the implicit field does not change, but only part of the high-level resolution spatial voxel is assigned a learnable feature vector. On this basis, the resolution space of each layer is mapped to a one-dimensional Morton code using three-dimensional coordinates, and a hash table is constructed to assign independent indexes to each layer space voxel block (which refers to dividing the three-dimensional space into fixed size resolution cubic blocks) to realize one-to-one mapping relationship with the actual spatial voxel of corresponding resolution, and improve the representation ability of three-dimensional space. As shown in Figure 1 As can be seen from Figure 1 , by artificially setting a scaling ratio of the actual three-dimensional space to be constructed to the implicit space, the actual space point cloud is directly transformed to the implicit space with a scale of [-1, 1], which is also the 0th resolution of the structured point cloud. The green, orange, blue and red cubic voxel divisions in the figure represent the visualization of the resolutions from 0th to 3rd. The voxels of 0th and 1st resolutions are divided by solid lines, indicating that no vertex feature vector and hash index table are assigned to these two dimensions, and the 2nd and 3rd resolution voxels are divided by dashed lines, which are assigned with vertex features and hash indexes. The interpolation results of the vertex features of the points in each voxel are given by linear interpolation, and the feature is added level by level instead of cascaded, which can make the feature expression of each level have higher dimension. Finally, the feature vectors under different scale mappings can be output according to the specified resolution of each level, which can enable the operator to extract the three-dimensional model of the detection scene with different spatial resolutions.

[0060] The mixed radiation field model obtains the pose image sequence data, the defect category, the global color, the local color and the depth information of the defect segmentation result image sequence, and performs fusion modeling to obtain the surface three-dimensional model of the offshore infrastructure with a specified spatial resolution; the mixed radiation field model includes a multi-level geometric information explicit coding module, a multi-level color information coding module, a camera pose color spherical harmonic coding module and a multi-head MLP decoding module; the multi-head MLP decoding module includes local color MLP decoding, global color MLP decoding, depth MLP decoding and semantic MLP decoding

[0061] Figure 1 The color information high-resolution fusion and scale coding selection part in Figure 1Taking the example of color information encoding, this invention can choose to use the deepest encoded features to fuse blue and red resolutions, thereby improving the model's ability to represent color in space and from multiple perspectives. When applying geometric information (SDF) extraction, red resolution represents the upper limit of extraction accuracy. This invention can choose to extract blue features at a lower resolution and directly perform surface reconstruction, which can greatly improve the speed of surface extraction. Simultaneously, appropriately reducing the resolution of ensemble information extraction helps to better fuse high-level geometric voxel information across multiple frames.

[0062] like Figure 2 As shown, this invention uses an implicit spatial multi-level specified resolution feature encoding method to encode geometric features, semantic features, and color features separately using independent feature spaces. This reduces the interference of semantic and depth information on high-resolution color information during loss backpropagation. Figure 2 As can be seen, spatial encoding uses an explicit-implicit multi-level resolution mapping approach. Spatial semantic features (referring to feature vectors encoded by a deep MLP) and SDF features are decoded through their respective MLPs (Multilayer Perceptrons), outputting corresponding class labels and normalizing SDF (signed distance field) values. Global and local colors use their own separate explicit-implicit multi-level resolution encoding spaces. For local color generation, the camera pose is encoded using the SH spherical harmonic function, and the resulting spherical harmonic function feature coefficients are concatenated with the spatial feature encoding of the implicit location points in the local color space as pose representation information. For global color, a direct-connected MLP is used for decoding, ultimately outputting normalized global and local color information respectively.

[0063] 1) The encoding process specifically involves: extracting features from the pose image sequence data and the defect segmentation result image sequence to obtain geometric features, color features, and defect semantic features; and then performing independent feature space encoding on the geometric features, color features, and defect semantic features, including:

[0064] The real-world coordinates of pose image sequence data and defect segmentation result image sequence are compressed into the implicit neural field using a multi-level geometric information explicit encoding module. The spatial resolution of the implicit neural field is hierarchically divided to obtain a cubic space for each level. Each vertex is assigned an independent index. Linear interpolation is performed using the vertices of multiple cubic blocks in the cubic space to extract the geometric and defect information of each spatial point, resulting in feature space encoding features of geometric features and defect semantic features. The cubic space of the same level consists of cubic blocks of the same size, and each cubic block vertex has a feature encoding vector. The implicit neural field refers to a continuous cubic space with feature surfaces in the neural network.

[0065] The global color information of the pose image sequence data and the defect segmentation result image sequence is extracted using a multi-level color information encoding module to obtain feature space encoding features of global color features.

[0066] The local camera pose perspective and local color information of the pose image sequence data and the defect segmentation result image sequence are extracted using a camera pose color spherical harmonic encoding module to obtain feature space encoding features of local color features.

[0067] 11) The decoding process is specifically as follows: the feature space encoding features are decoded, specifically:

[0068] The feature space encoding features of the global color features are decoded and rendered using a global color MLP to extract color information of the global three-dimensional model, i.e., to obtain local color; the feature space encoding features of the local color features are decoded and rendered using a local color MLP combined with the camera pose color spherical harmonic encoding module to extract a color image of the local perspective, i.e., to obtain local color; the depth information of the global three-dimensional model is rendered using a depth MLP: the spatial resolution output by the mixed radiation field model is set, the spatial depth information of the mixed radiation field is decoded using the depth MLP, the depth features of the three-dimensional model are rendered, and the spatial grid voxel depth information corresponding to the resolution, i.e., the position information of the spatial points relative to the imaging origin of the camera, is obtained; the defect semantic categories of the global three-dimensional model are rendered using a semantic MLP: the spatial resolution output by the model is set, the spatial defect semantic information of the mixed radiation field is decoded using the semantic MLP, the defect label information of the semantic image of the three-dimensional model at any perspective is rendered, and the spatial grid voxel semantic label information corresponding to the resolution is extracted as the defect categories.

[0069] Figure 2 Two points need to be noted in the application. First, the application unifies the defect semantic feature encoding and the SDF encoding, and shares the same spatial features, because the semantic information of the surface is spatially unique and does not change with different perspectives. Then, the application separates the rendering of the color into two MLPs, the information of the local color is represented in a way of cascading based on the directional spherical harmonic function encoding and spatial encoding, and the global color is represented using a separate spatial feature. This way can improve the fine degree of rendering of the local image color and avoid the influence of the global color fusion on the local image rendering. The geometry encoding and the color encoding are mapped to the distance field and the color field space using the sigmoid function, and the semantic encoding is mapped to the semantic label space for iterative learning (fusion) through the softmax.

[0070] Step S4: set the spatial resolution of the actual scene corresponding to the mixed radiation field model, combine the spatial resolution of the actual scene with the defect category, global color, local color and depth information of the independent feature space, output the feature vectors under different scale mappings at a specified hierarchical resolution, and fuse the feature vectors under different scale mappings to obtain a marine infrastructure surface three-dimensional model at a specified spatial resolution, and extract the position and quantity of the marine infrastructure surface defects through the marine infrastructure surface three-dimensional model.

[0071] In combination with the constructed three-dimensional model, feature extraction can be performed on any hierarchical level with feature encoding in the implicit space, and then a plurality of functional MLP modules are used to solve different rendering tasks. Although the NeRF implicit neural field method can achieve high-quality expression of local image details, since the rendering equation of the radiation field is used as the basis for image generation, the density values of the dense sampling points in the space along the ray direction need to be fitted and regressed. This method causes the density of the corresponding space point position of the implicit field to not completely converge to the zero set of the SDF (distance field). Therefore, the present application uses Figure 3 The left side is a method of randomly sampling near-surface points by a TSDF-like method. This method is more accurate than the sampling of the original NeRF AABB bounding box.

[0072] The marine infrastructure surface three-dimensional model includes a global defect three-dimensional model and a global color three-dimensional model.

[0073] The training process of the mixed radiation field model is shown in Figure 3 and includes four steps:

[0074] According to the camera pose and the image sequence, an initialized mixed radiation field model is constructed by using the marching cube method.

[0075] First, global coarse training is performed, and the input sequence depth, color and semantic image are used to iteratively train the mixed radiation field space encoding feature and the MLP network parameter.

[0076] Then, when the depth loss converges, the geometric MLP network parameter is fixed, and the global three-dimensional grid voxel set is obtained by using the geometric MLP decoding.

[0077] Next, in order to perform fine learning on the local key frame, the intersection of the local voxel to be finely learned and the global voxel in the second step is solved by using the inverse projection method of the local image.

[0078] Finally, the spatial color encoding model of local intersection voxels, global color MLP, and local camera pose spherical harmonic encoding parameters are fine trained to make the model have stronger local image detail rendering capability.

[0079] From Figure 3 It can be seen that the input of the system is the image sequence after camera pose solving (using VO / SFM method), including depth, color and defect semantic map. The defect semantic is detected using U-net and other semantic segmentation models, including common concrete defects (cracks, spalling, pitting, rain erosion, artificial repair).

[0080] The rendering of the explicit-implicit multi-level specified resolution block is divided into two links: first, the global optimization link, which uses the sampling point attributes (color, depth and part of the semantic label) of the RGB-D input to supervise the learning of the sampling points in the coarse global stage. The sampling points are divided into near-measurement surface dense sampling points and far-measurement surface sparse sampling points, and the surface measurement data generated by the depth image is sampled through the set parameters, as shown in the following table. The sampled points are encoded in space and decoded by MLP to output the color, semantic and SDF value of the specified resolution block and perform supervised learning optimization, which saves the operation of the ray rendering equation; the second is the fine color optimization link in the local view cone space, which includes Figure 3 The process in circle 2-circle 4. Circle 2 can be considered to have learned the spatially fixed SDF and semantic label features after a certain number of iterations of global feature representation learning, and the geometric space encoding feature parameters The MLP and semantic MLP of SDF are fixed parameters, and the spatial surface extraction resolution is set to extract the SDF value, and then MarchingCubes is used to extract the surface mesh. It should be particularly pointed out that when the implicit field resolution level of the extracted global surface is slightly lower than the highest level of the input resolution ( ), the feature fusion effect of the implicit voxel can be achieved.

[0081] After obtaining the surface mesh, circle 3 can realize batch ray tracing of the local surface point set through camera pose inverse projection, and finally circle 4 realizes fine local color supervision learning through the highest resolution level of the local point set, and the index of the local color MLP output is obtained by cascading the SH encoding of the camera pose, and the fine local color value is outputted, and the fine local color supervision learning is completed. At the same time, the color encoding space and the global color MLP decoder are also optimized and learned.

[0082] Table 1 Explicit-implicit model resolution and sampling related parameter table

[0083]

[0084] Maximum resolution level in Table 1 is the actual scale of the highest resolution Correspondingly, both can determine the actual explicit space representation range of the implicit space coding As shown in the following formula.

[0085] ;

[0086] The corresponding feature level needs to be less than or equal to The sampling of the measured surface in the table adopts the method of distance threshold, and dense sampling is carried out in the distance range of the depth image measurement value Outside this range, the present application needs to expand the sampling range to increase the robustness of the measurement uncertainty, so that the adaptive sampling distance control with a proportion of is adopted for each light, and the starting value of the sampling range is , wherein represents the picture pixel coordinates Corresponding normalized spatial point measurement depth value, The end value of the interval sampling range is .

[0087] Since the model finally realizes the rendering of global color, local color, semantics and depth, before obtaining the pose image sequence data and the defect class, global color, local color and depth information of the defect segmentation result image sequence through the mixed radiation field model, it is necessary to establish a loss function respectively. At the same time, the feature learning process is divided into local and global two parts, therefore, the change of the loss function also has two stages, through the loss function to train the mixed radiation field model, the specific form is shown in the following formula.

[0088] ;

[0089] Wherein represents the sum of the loss function, represents the sum of the loss function in the global coarse optimization stage, represents the loss function in the local fine optimization stage, , and respectively represent the weight of the corresponding global color, semantic and local color loss, the local color loss and the global color loss and both adopt loss mode, the semantic uses cross entropy loss; wherein, the geometric loss Including surface normal Eikonal loss and labeling surface point loss , as shown below.

[0090] ;

[0091] ;

[0092] in, This represents the depth information obtained after MLP decoding. By using the sigmoid function to negate the value of 's', the SDF value is compressed to the interval [0,1], thus realizing a probability representation of the likelihood of a surface point. This indicates the size of the SDF; when the SDF equals 0, The output probability value is 1, achieving unbiasedness. This represents the actual SDF measurement at that point. The final geometric loss. .

[0093] The multi-type rendering results of local and global are represented by the following combination formula, where global color, semantics, distance field, and local color are represented by their respective... The branch proceeds to decode the spatially encoded features. Spatial encoding uses the same spatial encoding model for both the global SDF and semantic information. Using an alternative spatial coding model in color coding tasks Simultaneously, attitude spherical harmonic function encoding from a local camera perspective is incorporated. Image depth values ​​are directly estimated using the distance between the intersection point of the view frustum and the surface extracted by global optimization. For view pixels that do not intersect with the global mesh model, this invention uses bilinear interpolation to estimate their rendering values.

[0094] ;

[0095] The training model parameters are as follows:

[0096] According to the explicit-implicit model resolution and sampling related parameter table, set the corresponding parameters ( Set to 0.05 meters. Set to 0.1 meters. (Set to 0.9), 4096 points are randomly selected for training in each round, and after 5000 rounds (coarse), the geometric feature space is fixed. And semantically related features and corresponding MLPs, and then on the color feature space The fine training is performed for 15000 rounds, and the global coarse training takes 1 hour, and the fine training takes 10-11 hours. In the experiment, after the global grid model is constructed, the intersection of the local image view volume and the global pixel can reach 99.7% of the total number of image pixels, and after the bilinear interpolation, the integrity of the rendered color image can be ensured. The loss weights ɑ , β , γ , and are set to 0.5, 0.5, 0.2, 0.02 and 0.1 respectively.

[0097] Finally, by iteratively training the color, depth and semantic mixed radiation field model of the sequence defect semantic image, after training, by setting the TSDF (Truncated Signed Distance Field) grid scale (greater than or equal to the minimum scale of the mixed radiation field coding) artificially, the corresponding large environment surface color, defect semantic three-dimensional model can be extracted, and the ability to generate corresponding depth, color and semantic image after inputting a specific local perspective is possessed. The brief description of the implementation steps is shown in Figure 4 .

[0098] Figure 4 The orange box in the above figure represents the function of the model, and it can be seen that through the training of the mixed radiation field, the model has the ability of traditional three-dimensional semantic reconstruction, and can generate the detection result and the original image of the local perspective.

[0099] Embodiment 2

[0100] In order to verify the defect detection effect of the method of the present application, the present application uses the existing data set for testing, and applies it to the detection and reconstruction of the side wall of the dam flow channel.

[0101] First, the defect semantic detection process of the depth neural network (using the U-net model to train on the artificially marked good perspective data set) of the collected image of the flow channel surface is performed, and the detection rate of each defect and the typical detection graph are output as shown in Figure 5 and Figure 6 .

[0102] Table 2 Actual scene defect detection rate

[0103]

[0104] Figure 7 The global color and semantic grid model is generated by using the semantic image sequence segmented by the depth neural network, the depth image and the original color image. The spatial resolution set in the mixed radiation field three-dimensional grid extraction process is 1 cm.

[0105] Embodiment 3:

[0106] The artificial sparse labeling result is used as input, and a deep neural network is used as a defect detection tool for extreme view angles, and the accuracy and precision of the professional personnel self-labeling are not currently superior, but the artificial labeling is time-consuming and laborious, and the scene of large infrastructure such as a dam is large, and the picture sequence is large, and each label is not realistic. However, due to the excellent spatial color, semantic and depth corresponding coding ability of the mixed radiation field, only a few pictures can be labeled to ensure that the sparse sampling image sequence covers the entire area, so that the global fusion of the semantic label and the local accuracy can be generated in the full coverage case. Through experiments, in 1000 surface images, only 33 of them are labeled (artificial labeling rate 3.3%), and a global high-precision three-dimensional semantic model can be constructed to render the defect local labeling images of all 1000 images with high defect detection rate under the condition of ensuring global coverage.

[0107] As shown in Figure 8 , the dam flow surface size of image acquisition is 17m*15m, the artificial labeling area can cover the entire area through the camera, and the three-dimensional model output is as shown in Figure 9 , from Figure 9 It can be seen intuitively that 3.3% of the artificial labeling rate can still generate a clear global defect detection three-dimensional model, and the mixed radiation field after training is used to render all local image defect semantic pictures and calculate the detection rate of various defects, as shown in Table 3:

[0108] Table 3: Calculated detection rate of various defects

[0109]

[0110] It can be seen that the semantic mixed radiation field model can generate a complete large scene surface defect detection three-dimensional semantic model under the condition of sparse sequence high-precision labeling, and render high-quality local defect semantic images with high detection rate.

[0111] At the same time, the model can also render the defect area that cannot be accurately identified under the limitation of extreme view angle in the artificial labeling process. The rendered local semantic test image is compared as shown in Figure 10 , the first row is the rendering result after training, and the second row is the initial labeling result, and it can be seen that the defects in the circle are not easy to confirm by the human eye under the extreme view angle.

[0112] The application also provides a defect detection system, which comprises a data acquisition module, an initial defect detection module, a feature rendering module and a defect scene reconstruction module, specifically:

[0113] The data acquisition module is configured to acquire image sequence data of the offshore infrastructure, and obtain pose image sequence data by solving camera pose of the image sequence data; the initial defect detection module is configured to preliminarily identify defects of the offshore infrastructure in the image sequence data to obtain a defect segmentation result image sequence; the feature rendering module is configured to extract geometric features, color features and defect semantic features from the pose image sequence data and the defect segmentation result image sequence, and independently encode and decode the geometric features, the color features and the defect semantic features to obtain defect categories, global colors, local colors and depth information in independent feature spaces; and the defect scene reconstruction module is configured to combine spatial resolution of an actual scene with the defect categories, the global colors, the local colors and the depth information in the independent feature spaces to output feature vectors under different scale mappings at a specified resolution, and fuse the feature vectors under the different scale mappings to obtain an offshore infrastructure surface three-dimensional model at a specified spatial resolution, and extract positions and quantities of offshore infrastructure surface defects from the offshore infrastructure surface three-dimensional model.

[0114] The present application also provides a computer device, comprising a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the defect detection method.

[0115] According to the disclosed embodiments, the computer device can communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth communication, etc.), or communicate with any device (such as a router, a demodulator, etc.) that enables the computer device to communicate with one or more other computer devices.

[0116] The present application also provides a computer readable storage medium, which stores a computer program, and the computer program is adapted to be loaded by a processor to execute the defect detection method.

[0117] According to the disclosed embodiments, the storage medium can be a non-volatile computer readable storage medium, which can include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution system, device or apparatus.

[0118] The above embodiments are only the preferred embodiments of the present application, and the protection scope of the present application is not limited thereto, and any simple change or equivalent replacement of the technical solutions which can be obviously obtained by those skilled in the art within the technical range disclosed by the present application shall belong to the protection scope of the present application.

Claims

1. A method for detecting surface defects in marine infrastructure, characterized in that, Includes the following steps: Acquire image sequence data of marine infrastructure, and obtain pose image sequence data by processing the image sequence data through camera pose calculation; Preliminary identification of defects in marine infrastructure in image sequence data is performed to obtain defect segmentation result image sequences; Feature extraction is performed on the pose image sequence data and the defect segmentation result image sequence to obtain geometric features, color features and defect semantic features. The geometric features, color features and defect semantic features are then independently encoded and decoded in the feature space to obtain the defect category, global color, local color and depth information in the independent feature space. Combining the spatial resolution of the actual scene with the defect category, global color, local color and depth information of the independent feature space, feature vectors under different scale mappings are output at a specified level of resolution, and feature vectors under different scale mappings are fused and modeled to obtain a three-dimensional model of the marine infrastructure surface with a specified spatial resolution. The location and number of defects on the marine infrastructure surface are extracted through the three-dimensional model of the marine infrastructure surface. The hybrid radiation field model is used to obtain the defect category, global color, local color and depth information of the pose image sequence data and defect segmentation result image sequence, and then perform fusion modeling to obtain a three-dimensional model of the surface of marine infrastructure with a specified spatial resolution. The hybrid radiation field model includes a multi-level geometric information explicit encoding module, a multi-level color information encoding module, a camera attitude color spherical harmonic encoding module and a multi-head MLP decoding module. The multi-head MLP decoding module includes local color MLP decoding, global color MLP decoding, depth MLP decoding, and semantic MLP decoding; The step involves extracting features from the pose image sequence data and the defect segmentation result image sequence to obtain geometric features, color features, and defect semantic features, and then independently encoding these geometric features, color features, and defect semantic features in a feature space, including: A multi-level geometric information explicit encoding module is used to compress the real-world spatial coordinates of pose image sequence data and defect segmentation result image sequence into the implicit neural field. The spatial resolution of the implicit neural field is hierarchically divided to obtain a cubic space for each level. Each vertex is assigned an independent index. Linear interpolation is performed using the vertices of multiple cubic blocks in the cubic space to extract the geometric and defect information of each spatial point, resulting in feature space encoding features of geometric features and defect semantic features. Among them, the cubic space of the same level is composed of cubic blocks of the same size, and each cubic block vertex has a feature encoding vector. The implicit neural field refers to a continuous cubic space with feature surfaces in the neural network. The global color information of the pose image sequence data and the defect segmentation result image sequence is extracted using a multi-level color information encoding module to obtain the feature space encoding features of the global color features. The local camera pose viewpoint and local color information of the pose image sequence data and the defect segmentation result image sequence are extracted using the camera pose color spherical harmonic coding module to obtain the feature space coding features of the local color features.

2. The method for detecting surface defects in marine infrastructure according to claim 1, characterized in that, Decoding the encoded features in the feature space is specifically as follows: Global color MLP decoding is used to decode and render the feature space encoded features of global color features to extract the global color. Local color MLP is used in conjunction with camera pose color spherical harmonic coding module to decode and render the feature space encoded features of local color features to obtain local color; Set the spatial resolution of the mixed radiation field model output, use depth MLP to decode the spatial depth information of the mixed radiation field, render the depth features of the 3D model, and obtain the depth information of the spatial mesh voxels corresponding to the resolution. Set the spatial resolution of the mixed radiation field model output, use semantic MLP to decode the spatial defect semantic information of the mixed radiation field, render the defect label information of the semantic image of the 3D model from any viewpoint, and extract the spatial mesh voxel semantic label information of the corresponding resolution as the defect category.

3. The method for detecting surface defects in marine infrastructure according to claim 1, characterized in that, The image sequence data of the marine infrastructure includes RGB-D image data obtained by a color depth camera, image data obtained by a binocular camera, image sequences with depth information obtained by a monocular camera through a neural network, and point cloud and image data obtained by a calibrated lidar and camera.

4. The method for detecting surface defects in marine infrastructure according to claim 1, characterized in that, The defect segmentation result image sequence includes crack images, erosion images, pitted images, or images of artificial repair categories.

5. The method for detecting surface defects in marine infrastructure according to claim 1, characterized in that, Before obtaining the defect category, global color, local color, and depth information of the pose image sequence data and defect segmentation result image sequence through the hybrid radiation field model, the method further includes constructing a loss function and training the hybrid radiation field model using the loss function, as follows: ; in This represents the sum of the loss functions. This represents the sum of the loss functions in the global coarse-finding stage. This represents the loss function during the local fine-tuning stage. , and These represent the weights of the corresponding global color, semantic, and local color losses, respectively. (Local color loss...) and global color loss and All adopted Loss patterns, semantics Use cross-entropy loss; where geometric loss... Including surface normal Eikonal loss and labeling surface point loss .

6. A surface defect detection system for marine infrastructure, characterized in that, include: The data acquisition module is used to acquire image sequence data of marine infrastructure and to obtain pose image sequence data by processing the image sequence data through camera pose calculation. The initial defect detection module is used to perform preliminary identification of defects in marine infrastructure in image sequence data, and obtain defect segmentation result image sequences; The feature rendering module is used to extract features from pose image sequence data and defect segmentation result image sequence to obtain geometric features, color features and defect semantic features. It also performs independent feature space encoding and decoding on geometric features, color features and defect semantic features to obtain defect category, global color, local color and depth information in independent feature space. The defect scene reconstruction module is used to combine the spatial resolution of the actual scene with the defect category, global color, local color and depth information of the independent feature space, output feature vectors under different scale mappings at a specified resolution, and fuse and model the feature vectors under different scale mappings to obtain a three-dimensional model of the marine infrastructure surface with a specified spatial resolution, and extract the location and number of defects on the marine infrastructure surface through the three-dimensional model of the marine infrastructure surface. The hybrid radiation field model is used to obtain the defect category, global color, local color and depth information of the pose image sequence data and defect segmentation result image sequence, and then perform fusion modeling to obtain a three-dimensional model of the surface of marine infrastructure with a specified spatial resolution. The hybrid radiation field model includes a multi-level geometric information explicit encoding module, a multi-level color information encoding module, a camera attitude color spherical harmonic encoding module and a multi-head MLP decoding module. The multi-head MLP decoding module includes local color MLP decoding, global color MLP decoding, depth MLP decoding, and semantic MLP decoding; The step involves extracting features from the pose image sequence data and the defect segmentation result image sequence to obtain geometric features, color features, and defect semantic features, and then independently encoding these geometric features, color features, and defect semantic features in a feature space, including: A multi-level geometric information explicit encoding module is used to compress the real-world spatial coordinates of pose image sequence data and defect segmentation result image sequence into the implicit neural field. The spatial resolution of the implicit neural field is hierarchically divided to obtain a cubic space for each level. Each vertex is assigned an independent index. Linear interpolation is performed using the vertices of multiple cubic blocks in the cubic space to extract the geometric and defect information of each spatial point, resulting in feature space encoding features of geometric features and defect semantic features. Among them, the cubic space of the same level is composed of cubic blocks of the same size, and each cubic block vertex has a feature encoding vector. The implicit neural field refers to a continuous cubic space with feature surfaces in the neural network. The global color information of the pose image sequence data and the defect segmentation result image sequence is extracted using a multi-level color information encoding module to obtain the feature space encoding features of the global color features. The local camera pose viewpoint and local color information of the pose image sequence data and the defect segmentation result image sequence are extracted using the camera pose color spherical harmonic coding module to obtain the feature space coding features of the local color features.

7. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to perform a method for detecting surface defects of marine infrastructure as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to execute a method for detecting surface defects in marine infrastructure as described in any one of claims 1-5.

Citation Information

Patent Citations

  • TFT-LCD mura defect detection method based on ICA learning and multichannel fusion

    CN105913419A

  • Automobile leather defect detection method and system based on visual detection

    CN120655628A