Safety intelligent inspection robot system for ocean engineering coating construction

By integrating multiple sensors and AI algorithms into the intelligent inspection robot system, the problems of low efficiency and safety hazards in manual methods of coating construction and inspection of marine engineering equipment have been solved. It has achieved efficient and safe coating condition monitoring and fault early warning, and reduced operation and maintenance costs.

CN120927697APending Publication Date: 2025-11-11CNOOC CHANGZHOU PAINT & COATINGS IND RES INST +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510917538.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In the current technology, the coating construction and inspection of marine engineering equipment mainly rely on manual methods, which has problems such as low efficiency, difficulty in detecting internal defects, and great safety hazards, especially in high-risk areas such as confined compartments.

Method used

Develop an intelligent inspection robot system for safety of marine engineering coating construction. The system integrates LiDAR, binocular camera and infrared camera sensors, and is equipped with AI vision algorithms and 5G communication to achieve full-scene operation, monitor coating status and environment in real time, generate risk heat maps, and support digital twin linkage.

Benefits of technology

It achieves a 95% reduction in human exposure time in high-risk scenarios, a 40% reduction in coating maintenance costs, and an advance warning of sudden failures of more than 72 hours, providing a standardized intelligent operation and maintenance solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120927697A_ABST
    Figure CN120927697A_ABST
Patent Text Reader

Abstract

The invention provides an ocean engineering coating construction safety intelligent inspection robot system, which comprises a motion control module used for all-terrain passing motion to adapt to a complex ship environment; a laser radar sensor, a binocular camera sensor and an infrared camera sensor are integrated on the intelligent sensing module, and the intelligent sensing module is used for collecting environment data; and the data processing module is used for realizing real-time data analysis and three-dimensional modeling, and realizing the functions of personnel safety monitoring, coating state evaluation and environment three-dimensional reconstruction through cooperative work of the motion control module, the intelligent sensing module and the data processing module. The method has the beneficial effects that the manual exposure time of a high-risk scene is shortened by 95%, the maintenance cost of the coating is reduced by 40%, and the early warning amount of sudden failures reaches 72 hours or above. The modular design supports expansion of a mechanical arm / VOCs sensor and the like, and a standardized solution is provided for intelligent operation and maintenance of ocean engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent inspection of coatings on large ships and marine engineering equipment, and more specifically to an intelligent inspection robot system for the safety of marine engineering coating construction. Background Technology

[0002] The rapid development of marine engineering equipment has led to a significant increase in the demand for coating construction. However, with the increasing service life of equipment and the impact of factors such as marine environmental corrosion, mechanical damage, and material aging, the coatings of some marine engineering equipment have begun to show safety hazards, such as peeling, blistering, and corrosion, gradually entering the maintenance period. According to the development pattern of marine engineering, my country has entered a crucial stage of marine engineering equipment maintenance, shifting the focus from large-scale new construction to the maintenance of existing equipment. Timely and accurate detection of coating defects is the first step in achieving safe operation and maintenance. Since 2019, through coating repair and protective reinforcement, critical parts of the coatings on many ships and offshore platforms have been maintained. Some defects are visible on the coating surface, while others are hidden within the coating, posing potential threats. These defects are difficult to detect and gradually weaken the protective performance of the coating, increasing the risk of sudden equipment failure. Therefore, timely detection and maintenance of the coating's appearance and internal condition are crucial for ensuring the safe operation of marine engineering equipment.

[0003] Currently, safety inspections of marine equipment coating construction primarily rely on manual visual inspections: construction personnel check the coating surface condition (observing defects such as peeling, blistering, and cracks), measure coating thickness using a thickness gauge, verify the implementation of ventilation / explosion-proof measures in the construction area, and supervise the wearing of protective equipment and the proper storage of hazardous chemicals. However, this method suffers from inefficiency, reliance on personnel experience, subjectivity, difficulty in detecting internal defects, and safety hazards in special environments such as confined spaces. It is recommended to combine intelligent detection equipment with standardized procedures to improve safety management.

[0004] With the development of intelligent marine engineering equipment, traditional manual safety inspection methods are facing severe challenges in scenarios such as ships and offshore platforms. Construction personnel need to enter high-risk areas such as enclosed compartments and high-altitude sides to monitor coatings. They not only face risks such as exposure to toxic gases and falls from heights, but also the low efficiency of manual handheld monitoring equipment makes it difficult to fully detect defects such as coating peeling and blistering.

[0005] Therefore, there is an urgent need to develop an intelligent safety inspection robot system with full-scenario operation capabilities. This system can penetrate high-risk areas such as ballast tanks and fuel tanks to replace manual inspections, and integrates explosion-proof gas sensors and temperature and humidity monitoring modules to provide real-time warnings of hazardous environments. At the same time, it is equipped with AI vision algorithms to monitor the PPE wearing status of construction personnel in real time, and immediately alarm for violations. It also uses infrared thermal imaging technology to detect abnormal coating thickness (accuracy ±0.01mm) and internal corrosion defects. Based on 5G communication, it realizes real-time transmission of inspection data, automatically generates risk heat maps and maintenance suggestions, and supports digital twin linkage to provide decision support for managers. Summary of the Invention

[0006] This invention overcomes the shortcomings of the prior art and provides an intelligent inspection robot system for the safety of marine engineering coating construction.

[0007] The objective of this invention is achieved through the following technical solution.

[0008] A smart inspection robot system for safety inspection of marine engineering coating construction includes:

[0009] The motion control module is used for all-terrain maneuvering to adapt to the complex environment of a ship.

[0010] The intelligent sensing module integrates a lidar sensor, a binocular camera sensor, and an infrared camera sensor to collect environmental data.

[0011] The data processing module is used to perform real-time data analysis and 3D modeling.

[0012] Through the coordinated operation of the motion control module, intelligent sensing module, and data processing module, the system achieves functions such as personnel safety monitoring, coating status assessment, and three-dimensional environmental reconstruction.

[0013] The specific working steps of the intelligent inspection robot system for safety inspection of marine engineering coating construction are as follows:

[0014] S1. Data is collected through the intelligent sensing module;

[0015] S2. The data processing module preprocesses the data collected by the intelligent sensing module, and uses a multimodal feature-level fusion method to perform spatiotemporal synchronization on the data collected by the intelligent sensing module. The synchronized data is then subjected to data cleaning, point cloud denoising, and multi-sensor calibration extrinsic parameter matrix fusion processing in sequence.

[0016] S3. The processed data is further processed through the data processing module;

[0017] The coating defect detection unit in the data processing module performs multispectral fusion of infrared and laser data, defect classification, and NACE report generation;

[0018] The RGB-D data is processed by the construction safety monitoring unit of the depth vision in the data processing module. The safety status data and dynamic obstacle avoidance instructions are obtained by using YOLOv11 model inference, PPE recognition and dangerous behavior detection. The safety status data and dynamic obstacle avoidance instructions are then transmitted to the real-time alarm unit.

[0019] The point cloud data and IMU data are processed by the intelligent modeling unit of SLAM in the data processing module. The BIM model is obtained by using SLAM 3D reconstruction, dynamic obstacle removal and BIM mapping. The BIM model is then input into the 3D semantic model for compression and storage.

[0020] The binocular camera sensor has a binocular vision ranging module. In S1, the binocular vision ranging module performs distortion correction and epipolar correction on the image data, and performs disparity calculation and depth calculation on the data after distortion correction and epipolar correction.

[0021] Distortion correction uses the Brown-Conrady model to correct lens distortion, and the correction formula is as follows:

[0022]

[0023] Where k1, k2, and k3 are radial distortion coefficients, used to control the strength and order of barrel or pincushion distortion; p1 and p2 are tangential distortion coefficients, used to represent the offset caused by the imperfect coplanarity of the lens and sensor; x and y are normalized coordinates; Xcorrecet and Ycorrecet are the distorted coordinate values; and r is the radius to the imaging center.

[0024] Epipolar correction uses the Bouguet algorithm to calculate homography matrices H_1 and H_2, making the left and right images coplanar, and matching points only need to be searched along the horizontal direction.

[0025] Disparity calculation employs a semi-global matching SGBM algorithm to optimize the matching cost function.

[0026]

[0027] Where w(p,q) is the adaptive weight; ρ is the robust penalty function; I L I R , respectively, are the function values ​​of the left and right images; C(p,d) is the matching cost function;

[0028] The formula for depth calculation is as follows:

[0029]

[0030] Where z is the final depth, d is the parallax, b is the baseline distance, and f is the focal length in pixels.

[0031] The infrared camera sensor contains an infrared thermal imaging module. The specific operating steps of the infrared thermal imaging module in S1 include:

[0032] The radiation signal is converted, and the detector output voltage signal V of the infrared camera sensor is converted. out The relationship with the target's radiated energy is as follows:

[0033]

[0034] Where R is the detector responsivity (A / W), ∈(λ) is the target emissivity, and M λ (T) is Planck's formula for blackbody radiation, n th Thermal noise;

[0035] Noise equivalent temperature difference (NETD) calibration was performed using a blackbody source. The calculation formula is as follows:

[0036]

[0037] Where σnoise is the noise standard deviation, S is the responsivity, and ΔT is the blackbody temperature change.

[0038] Temperature inversion is performed, and the target temperature T is calculated based on the Stefan-Boltzmann law.

[0039]

[0040] Where E is the radiation energy and σ is the Stefan-Boltzmann constant.

[0041] The lidar has a built-in lidar point cloud processing module. The specific steps of the lidar point cloud processing module in S1 include:

[0042] For ToF ranging, the relationship between the laser pulse round-trip time Δt and the target distance d is as follows:

[0043]

[0044] Where c is the speed of light, n air n is the refractive index of air. vacuum The refractive index is the vacuum refractive index.

[0045] To compensate for system errors, the actual ranging time needs to be reduced by the system delay t. system_delay ,

[0046] Δt actual =Δt measured -tsystem_delay

[0047] Where, Δt actual Δt represents the actual distance measurement time. measured For measuring time; t system_delay This is due to system latency;

[0048] MEMS scanning control is performed, and the deflection angle θ(t) of the micromirror is controlled by a sinusoidal drive signal.

[0049] θ(t)=θ max ·sin(2πf scan t)

[0050] Among them, f scan The scanning frequency determines the point cloud sampling rate; θ max This represents the maximum deflection angle of the micromirror.

[0051] The specific steps of the multimodal feature-level fusion method in S2 include:

[0052] A1. Perform external parameter calibration, and obtain the transformation matrix T between the lidar sensor and the binocular camera sensor by calibrating the target. lidar→camera The transformation matrix T between the infrared camera sensor and the stereo camera sensor Thermal→camera ;

[0053] A2. Time synchronization is performed using the PTP precise time protocol to achieve μs-level synchronization and align the data from each sensor set on the intelligent sensing module with time.

[0054] A3. Perform feature fusion and dynamically adjust the weights of multimodal features using an attention mechanism.

[0055] F fusion =α·CNN(I RGB )+β·GNN(P Lidar )+γ·Grad(T Thermal )

[0056] Where α, β, and γ are adaptive weight coefficients; F is the fused comprehensive feature representation used in subsequent tasks; CNN(I RGB ) represents the RGB image features extracted by a convolutional neural network; I RGB The input is an RGB color image; GNN(P) Lidar P represents the LiDAR point cloud features extracted using a graph neural network. Lidar 3D point cloud data acquired by LiDAR; Grad(T) Thermal This is used to extract gradients or structural features from thermal imaging data to emphasize target and texture details; T Thermal This is thermal imaging infrared image data.

[0057] The data processing module adopts the YOLO11 model, replacing the original ELAN structure in the YOLO11 model with the MobileOne-S4 module. It constructs an efficient feature extraction unit using 3×3 and 1×1 convolutions combined with BN and SiLU activation, and introduces an Identity branch.

[0058] The formula for the MobileOne-S4 module is as follows:

[0059] W fused =W 3×3 +W 1×1 +I

[0060] Among them, W 3×3 For a 3×3 convolution kernel, W 1×1 The kernel is a 1×1 convolution, and I is the Identity mapping;

[0061] The channel shuffling mechanism of ShuffleNetV2 is introduced. The formula for the channel shuffling mechanism of ShuffleNetV2 is as follows:

[0062] Shuffle(x) i,j =x π(i),j ,π(i)=(i+j)modC

[0063] Where Shuffle(x) i,j This represents the channel shuffling operation; x represents the characteristics of the channel before shuffling; i is the channel index; j is the spatial position; C is the number of channels; π(i) is the shuffling index;

[0064] In the neck feature fusion stage of the YOLO11 model, the ViT-CSPNeck architecture is introduced. It combines the Transformer self-attention structure to extract global features and works with CSPNeck to complete local feature fusion. The feature selector module FSMFeatureSelectionModule realizes the dynamic weight allocation of features of different scales and channels.

[0065] The ViT branch divides the input image into 16×16 patches and extracts global features through a Transformer encoder. The key mathematical expression of the Transformer encoder module is as follows.

[0066]

[0067] Where z is the output vector, representing the result after multi-head self-attention and residual connection; z lThe input vector represents the output or initial input of the previous layer; LN(·) is layer normalization; MSA(·) is multi-head self-attention, which calculates self-attention on the normalized input to extract contextual information; z l-1 For residual connections, the original input information is preserved to alleviate the vanishing gradient problem; Q is the query matrix, used to calculate the similarity with the keys and determine which parts to focus on; K is the key matrix, used to interact with the query matrix to generate attention weights; V is the value matrix, used to sum the values ​​according to the attention weights to generate the final output; d k The dimension of the key vector is used as a scaling factor to prevent the dot product result from becoming too large and causing gradient instability; QK T For query-key interactions, the similarity between the query and the key is calculated, i.e., the unscaled attention score; Softmax(·) is a normalization function used to transform the attention score into a probability distribution.

[0068] YOLOv11 deep features are extracted using CSPNeck for multi-scale local features, and then dynamically weighted and fused with the ViT output using FSM.

[0069] The dynamic weighted fusion formula for the multi-branch feature fusion module is as follows:

[0070] F fusion =α·ViT(x)+(1-α)·CSPNeck(x),

[0071] α=σ(MLP([F vit ,F csp ]))

[0072] Where σ is the Sigmoid function, with an output range of [0,1]; F frusion The fused feature output is ViT(x); ViT(x) is the feature extraction of input x by VisionTransformer; CSPNeck(x) is the CNN feature extraction module based on CSP structure; α is the state fusion weight, generated by MLP and Sigmoid; [F vit ,F csp [] This is a concatenation operation used to merge ViT features and auxiliary features;

[0073] The YOLO11 model uses Focal-CIoU loss in the detection head stage, with the specific formula as follows.

[0074]

[0075] Where ρ is the Euclidean distance, c is the length of the diagonal of the minimum bounding rectangle, and 1-IoU is the IoU loss term, which is used to measure the degree of overlap between the predicted box and the ground truth box. Here, b is the normalized distance term for the center point, and b is the center of the predicted bounding box.gt ρ is the center of the true bounding box, ρ is the Euclidean distance, and αv is the aspect ratio consistency penalty term. The aspect ratio angle of the actual bounding box is used to adjust the aspect ratio of the actual bounding box. Convert to angle value; The aspect ratio angle of the prediction box is used to adjust the aspect ratio of the prediction box. Convert to angle values ​​and compare with the actual bounding box angles.

[0076] In the label matching and loss calculation stage of YOLO11 model training, dynamic label allocation is adopted to increase the number of positive samples to 3-5 / GT;

[0077] In the real-time tracking optimization phase of the YOLO11 model, the ByteTrack tracker is used, employing a track management and GIoU tracking matching strategy to achieve integrated detection and track tracking.

[0078] The Kalman filter state equation for the ByteTrack tracker is:

[0079]

[0080] The matching cost matrix uses GIoU as a metric, and the specific formula is as follows.

[0081]

[0082] Where C is the minimum closure region of A and B, IoU is the intersection-union ratio, and the overlap between the predicted box A and the ground truth box B is calculated; C\(A∪B) is the part of region C that does not belong to A or B.

[0083] The motion control module is an all-terrain robot platform, including a body, a head, and wheels. The body is equipped with a battery and a high-computing-power module, which serves as the hardware carrier for the data processing module. The head is located at the front of the body and is equipped with a lidar sensor, a binocular camera sensor, and an infrared camera sensor for the intelligent sensing module. The body has four wheels, and the bottom of the body has support pads.

[0084] The beneficial effects of this invention are: reducing human exposure time in high-risk scenarios by 95%, reducing coating maintenance costs by 40%, and providing early warning of sudden failures by more than 72 hours. Its modular design supports expansion with robotic arms / VOCs sensors, providing a standardized solution for intelligent operation and maintenance of marine engineering projects. Attached Figure Description

[0085] Figure 1 This is a schematic diagram of the backbone network structure;

[0086] Figure 2 This is a schematic diagram of the motion control module.

[0087] Figure 3 This is a side view of the motion control module;

[0088] In the picture: 1. Body; 2. Head; 3. Wheels; 4. Battery; 5. High-performance computing module; 6. Indicator light; 7. Support pad. Detailed Implementation

[0089] Example

[0090] A smart inspection robot system for safety inspection of marine engineering coating construction includes:

[0091] The motion control module is used for all-terrain maneuvering to adapt to the complex environment of a ship.

[0092] The intelligent sensing module integrates a lidar sensor, a binocular camera sensor, and an infrared camera sensor to collect environmental data.

[0093] The data processing module is used to perform real-time data analysis and 3D modeling.

[0094] Through the coordinated operation of the motion control module, intelligent sensing module, and data processing module, the system achieves functions such as personnel safety monitoring, coating status assessment, and three-dimensional environmental reconstruction.

[0095] In this embodiment, the intelligent sensing module includes an XT-16 LiDAR, a 4D LiDAR L1, an Intel RealSense D435i binocular camera, and an infrared thermal imaging camera.

[0096] The XT-16 lidar employs a mechanical rotating scanning principle and is equipped with 16 detection units, achieving omnidirectional coverage of a 360° horizontal field of view and a 30° vertical field of view (-15° to +15°), enabling precise capture of stereo information in complex scenes. Its ranging performance is excellent, covering a range of 0.05–120m. Under 10% reflectivity conditions, the maximum ranging distance for channels 5–12 reaches 80m, while channels 1–4 and 13–16 achieve 50m, meeting long-range detection requirements. Ranging accuracy is outstanding, with typical values ​​of ±1cm (accuracy) and 0.5cm (precision, 1σ), and standard values ​​of ±2cm (accuracy) and 2cm (precision), ensuring highly reliable spatial modeling.

[0097] The radar supports multiple scanning frame rates of 5Hz / 10Hz / 20Hz, with a maximum horizontal angular resolution of 0.09° (at 5Hz) and a vertical angular resolution of 2°, flexibly adapting to the dynamic detection needs of different scenarios. Employing a dual-echo mode, it achieves a maximum point frequency of 640,000 points / second (dual-echo mode) and a point cloud data transmission rate of 24.56Mbps, providing dense point cloud support for high-speed moving targets such as vehicles and drones.

[0098] The L1 4D LiDAR features a 360°×90° ultra-wide-angle scanning design, enabling hemispherical 3D spatial detection. It can detect overhead obstacles and identify even small steps as low as 5cm on the ground. Its ranging performance is excellent, with a minimum detection distance of only 0.05m, nearing zero blind zone, and a maximum range of 30m (under 90% reflectivity conditions), achieving an accuracy of ±2cm within a 10m range. 21,600 effective samples per second (total sampling rate 43,200 points / second) ensure rich environmental information acquisition. The radar outputs 3D point cloud (XYZ coordinates) and 1D grayscale data (reflection intensity), which can be used to distinguish different materials, such as glass and metal. It operates stably even in 100klux strong light environments and can identify highly transparent objects such as glass windows. The built-in 3-axis accelerometer and gyroscope (250Hz refresh rate) dynamically correct motion jitter, reducing point cloud distortion by 80%.

[0099] The Intel RealSense D435i binocular camera provides the system with accurate depth perception capabilities. This camera employs an infrared binocular + RGB combination scheme, based on the principle of active stereo vision. It projects a fixed texture through an infrared emitter to assist in ranging in weakly textured environments, with a ranging range of 0.3-3m and an accuracy of ±2cm. Its field of view is 87°×58° (horizontal×vertical), the RGB camera resolution is 1920×1080@30fps, and the depth map resolution is 1280×720@90fps. The camera supports hardware-level synchronization between the depth map and the color image, with a timestamp alignment error of less than 1ms, facilitating RGB-D data fusion. The built-in 6-DOF IMU (accelerometer + angular velocity) provides motion compensation in dynamic scenes.

[0100] The infrared thermal imaging camera has a resolution of 640×512, a temperature measurement range of -20℃ to 550℃, and can detect temperature differences as small as 0.07℃ (thermal sensitivity NETD). This camera supports multispectral fusion of thermal imaging data and RGB images, highlighting high-temperature areas.

[0101] The data processing module uses the NVIDIA Jetson Orin NX computing platform, which is a high-performance computing module installed on the machine.

[0102] The specific working steps of the intelligent inspection robot system for safety inspection of marine engineering coating construction are as follows:

[0103] S1. Data is collected through the intelligent sensing module;

[0104] S2. The data processing module preprocesses the data collected by the intelligent sensing module, and uses a multimodal feature-level fusion method to perform spatiotemporal synchronization on the data collected by the intelligent sensing module. The synchronized data is then subjected to data cleaning, point cloud denoising, and multi-sensor calibration extrinsic parameter matrix fusion processing in sequence.

[0105] S3. The processed data is further processed through the data processing module;

[0106] The coating defect detection unit in the data processing module performs multispectral fusion of infrared and laser data, performs defect classification and NACE report generation, and transmits the defect report to the risk heat map unit, which then issues a defect re-inspection request.

[0107] The RGB-D data is processed by the construction safety monitoring unit of the depth vision in the data processing module. The safety status data and dynamic obstacle avoidance instructions are obtained by using YOLOv11 model inference, PPE recognition and dangerous behavior detection. The safety status data and dynamic obstacle avoidance instructions are then transmitted to the real-time alarm unit.

[0108] The point cloud data and IMU data are processed by the intelligent modeling unit of SLAM in the data processing module. The BIM model is obtained by using SLAM 3D reconstruction, dynamic obstacle removal and BIM mapping. The BIM model is then input into the 3D semantic model for compression and storage.

[0109] The construction safety monitoring unit based on depth vision receives instructions from the NVIDIA Jetson Orin NX computing platform (100TOPS) and returns the collected construction safety data for processing.

[0110] The specific workflow is as follows: After receiving monitoring instructions, the NVIDIA Jetson Orin NX computing platform moves along a preset inspection path. Simultaneously, through attitude adjustment, the onboard Intel RealSense D435i binocular camera (1920×1080@30fps) and 4D LiDAR L1 (43200 points / second) achieve the optimal observation angle. RGB-D images and 3D point cloud data collected synchronously by multiple sensors are transmitted to the NVIDIA Jetson Orin NX computing platform in real time via Gigabit Ethernet. The NVIDIA Jetson Orin NX computing platform performs spatiotemporal registration on the data (time synchronization <10μs, spatial registration error <0.5mm) and then calls the YOLOv11-based risk identification model for processing. The system outputs a comprehensive safety assessment report that includes personnel PPE status (identification accuracy of safety helmets, protective clothing, etc. ≥99.2%), dangerous behaviors (identification delay of climbing, crossing boundaries, etc. <200ms), and high-temperature contact risks (fusion of infrared thermal imaging 0.07℃ sensitivity data). This report is then transmitted to the monitoring terminal in real time via 5G / WiFi 6 and other networks.

[0111] A multispectral coating defect detection unit, in collaboration with VIDIA Jetson OrinNX computation, performs coating quality assessment. During system operation, the robot moves along the bulkhead, automatically maintaining a distance of 30±5cm between the intelligent sensing module and the surface; the LiDAR L1 identifies macroscopic defects (peeling, cracks) using grayscale data (reflection intensity 0-255), with a detection sensitivity of 1mm. 2 The RGB images acquired by the binocular camera are used by a defect recognition network to classify micro-defects (bubbling, cracking, etc., with a recognition accuracy of 98.5%). An infrared thermal imager (640×512@25Hz) detects internal voids (minimum detectable size 3cm) through temperature difference analysis (resolution 0.07℃). 2 (Defects); After feature-level fusion of multi-source data on the NVIDIA Jetson OrinNX computing platform, a digital report conforming to the NACESP0169 standard is automatically generated, which includes 12 quantitative indicators such as defect type, size, and location, and is stored in the BIM system through Octree compression technology (compression ratio of 70%).

[0112] SLAM-based intelligent modeling units enable 3D digital reconstruction of ship cabins.

[0113] Once the system is started, the LiDAR (range accuracy ±1cm@10m) and the binocular camera (depth map resolution 1280×720) synchronously acquire environmental data. The NVIDIA Jetson Orin NX computing platform fuses point cloud and depth information in real time through a tightly coupled SLAM algorithm (15Hz update rate) to construct a 3D semantic model with centimeter-level accuracy (texture mapping accuracy 2048×2048). The model automatically marks high-risk areas according to OSHA standards, supports design drawing comparison (error <3mm) and maintenance path planning (A* algorithm optimization). The modeling process uses dynamic voxel filtering (grid size 5cm) and KD tree acceleration to ensure real-time processing (latency <100ms) on the NVIDIA Jetson Orin NX computing platform.

[0114] Furthermore, the NVIDIA Jetson Orin NX computing platform adopts a layered processing architecture for edge computing platforms, with the framework layers including:

[0115] Sensor layer: Ensures data acquisition timing consistency through hardware synchronization triggering (PTP protocol);

[0116] Preprocessing layer: Performs point cloud denoising (statistical outlier removal), image distortion correction (Zhang Zhengyou calibration method), and other operations;

[0117] The core processing layer is divided into three parallel threads: a security monitoring thread (highest QoS priority, using 30% of computing power), a defect analysis thread (accelerated with TensorRT enabled, using 40% of computing power), and a modeling thread (accelerated with CUDA, using 30% of computing power), which are used to complete the processing tasks of the corresponding modules.

[0118] Output layer: Processing results are transmitted via 5G (end-to-end latency <50ms) or WiFi 6 (throughput >1Gbps). The system supports OTA remote upgrades, which can dynamically update the algorithm model without interrupting the inspection operation.

[0119] Furthermore, the specific steps of this solution to achieve automated building structure inspection include:

[0120] First, the robot receives remote commands and initializes task parameters, loads the environment map, and plans the optimal path;

[0121] Secondly, by collecting environmental data collaboratively through multiple sensors, a 3D model with centimeter-level accuracy is constructed in real time.

[0122] During the inspection phase, the robot adopts a tiered inspection strategy: first, it uses a depth camera to quickly screen for safety risks to construction workers and surface defects in the coating; once an anomaly is detected, it activates a high-precision inspection mode and adjusts the robot's posture to perform a detailed scan; at the same time, it uses infrared equipment to detect internal defects along a predetermined route.

[0123] The robot is equipped with a dynamic obstacle avoidance system to ensure stable operation in complex environments. Detection data is transmitted back in real time via 5G / WiFi 6 dual-mode, processed by deep learning algorithms to generate a quantitative report containing 12 parameters, and accurately mapped onto a 3D model.

[0124] Finally, inspectors can monitor the operation status in real time through a multi-terminal interactive interface and flexibly adjust the inspection strategy. Actual testing showed that the system improves inspection efficiency by 320% compared to manual inspection, achieves a positioning accuracy of ±3mm, and has a single-charge runtime of 4.5 hours, making it suitable for automated inspection needs of various building structures.

[0125] Furthermore, the binocular camera sensor is equipped with a binocular vision ranging module. In S1, the binocular vision ranging module performs distortion correction and epipolar correction on the image data, and performs disparity calculation and depth calculation on the data after distortion correction and epipolar correction.

[0126] Distortion correction uses the Brown-Conrady model to correct lens distortion, and the correction formula is as follows:

[0127]

[0128] Where k1, k2, and k3 are radial distortion coefficients, used to control the strength and order of barrel or pincushion distortion; p1 and p2 are tangential distortion coefficients, used to represent the offset caused by the imperfect coplanarity of the lens and sensor; x and y are normalized coordinates; Xcorrecet and Ycorrecet are the distorted coordinate values; and r is the radius to the imaging center.

[0129] Epipolar correction uses the Bouguet algorithm to calculate homography matrices H_1 and H_2, making the left and right images coplanar, and matching points only need to be searched along the horizontal direction.

[0130] Disparity calculation employs a semi-global matching SGBM algorithm to optimize the matching cost function.

[0131]

[0132] Where w(p,q) is the adaptive weight; ρ is the robust penalty function; I L I R , respectively, are the function values ​​of the left and right images; C(p,d) is the matching cost function;

[0133] The formula for depth calculation is as follows:

[0134]

[0135] Where z is the final depth, d is the parallax, b is the baseline distance, and f is the focal length, in pixels.

[0136] Furthermore, in this embodiment, the infrared camera sensor uses a mercury cadmium telluride (HgCdTe) detector, and the infrared camera sensor includes an infrared thermal imaging module. The specific working steps of the infrared thermal imaging module in S1 include:

[0137] The radiation signal is converted, and the detector output voltage signal V of the infrared camera sensor is converted. out The relationship with the target's radiated energy is as follows:

[0138]

[0139] Where R is the detector responsivity (A / W), ∈(λ) is the target emissivity, and M λ (T) is Planck's formula for blackbody radiation, n th Thermal noise;

[0140] Noise equivalent temperature difference (NETD) calibration was performed using a blackbody source. The calculation formula is as follows:

[0141]

[0142] Where σnoise is the noise standard deviation, S is the responsivity, and ΔT is the blackbody temperature change.

[0143] Temperature inversion is performed, and the target temperature T is calculated based on the Stefan-Boltzmann law.

[0144]

[0145] Where E is the radiation energy, and σ is the Stefan-Boltzmann constant (5.67 × 10⁻⁶). -8 W / m 2 K 4 ).

[0146] Furthermore, the lidar is equipped with a lidar point cloud processing module. The specific steps of the lidar point cloud processing module in S1 include:

[0147] For ToF ranging, the relationship between the laser pulse round-trip time Δt and the target distance d is as follows:

[0148]

[0149] Where c is the speed of light, n air Given the refractive index of air (1.0003), n vacuum The refractive index in vacuum is 1.0;

[0150] To compensate for system errors, the actual ranging time needs to be reduced by the system delay t. system_dealy :

[0151] Δt actual =Δt measured -t system_delay

[0152] Where, Δt actual Δt represents the actual distance measurement time. measured For measuring time; t system_delay This is due to system latency;

[0153] MEMS scanning control is performed, and the deflection angle θ(t) of the micromirror is controlled by a sinusoidal drive signal.

[0154] θ(t)=θ max ·sin(2πf scan t)

[0155] Among them, f scan The scanning frequency determines the point cloud sampling rate, where the point cloud sampling rate is 43200 points / second; θ max This represents the maximum deflection angle of the micromirror.

[0156] Furthermore, the specific steps of the multimodal feature-level fusion method in S2 include:

[0157] A1. Perform external parameter calibration, and obtain the transformation matrix T between the lidar sensor and the binocular camera sensor by calibrating the target. lidar→camera The transformation matrix T between the infrared camera sensor and the stereo camera sensor Thermal→camera ;

[0158] A2. Time synchronization is performed using the PTP precise time protocol to achieve μs-level synchronization and align the data from each sensor set on the intelligent sensing module with time.

[0159] A3. Perform feature fusion and dynamically adjust the weights of multimodal features using an attention mechanism.

[0160] F fusion =α·CNN(I RGB )+β·GNN(P Lidar )+γ·Grad(T Thermal )

[0161] Where α, β, and γ are adaptive weight coefficients; F is the fused comprehensive feature representation used in subsequent tasks; CNN(I RGB ) represents the RGB image features extracted by a convolutional neural network; I RGB The input is an RGB color image; GNN(P) Lidar P represents the LiDAR point cloud features extracted using a graph neural network. Lidar 3D point cloud data acquired by LiDAR; Grad(T) Thermal This is used to extract gradients or structural features from thermal imaging data to emphasize target and texture details; T Thermal This is thermal imaging infrared image data.

[0162] In this embodiment, the multimodal feature-level fusion method is driven by the NVIDIA Jetson OrinNX computing platform (100 TOPS computing power), connecting each sensor via Gigabit Ethernet / USB 3.0 to achieve high-precision fusion with time synchronization accuracy <10μs and spatial calibration error <0.5mm. The fusion algorithm includes: SLAM mapping function, which fuses LiDAR point cloud (30m range) with binocular depth map (3m range) to construct a 3D semantic map with centimeter-level accuracy (15Hz update rate) and supports dynamic object culling; dynamic obstacle avoidance function, which combines the long-range detection of LiDAR (30m), the close-range details of binocular camera (0.3m), and infrared thermal source data to achieve a fast response capability with obstacle avoidance response latency <50ms.

[0163] Furthermore, as shown in the figure, the data processing module adopts the YOLO11 model, and its performance has been enhanced through improvements to the YOLO11 model. Details are as follows.

[0164] The data processing module adopts the YOLO11 model, replacing the original ELAN structure in the YOLO11 model with the MobileOne-S4 module. It constructs an efficient feature extraction unit using 3×3 and 1×1 convolutions combined with BN and SiLU activation, and introduces an Identity branch.

[0165] The formula for the MobileOne-S4 module is as follows:

[0166] W fused =W 3×3 +W 1×1 +I

[0167] Among them, W 3×3 For a 3×3 convolution kernel, W 1×1 The kernel is a 1×1 convolution, and I is the Identity mapping;

[0168] The channel shuffling mechanism of ShuffleNetV2 is introduced. The formula for the channel shuffling mechanism of ShuffleNetV2 is as follows:

[0169] Shuffle(x) i,j =x π(i),j ,π(i)=(i+j)modC

[0170] Where Shuffle(x) i,j This represents the channel shuffling operation; x represents the characteristics of the channel before shuffling; i is the channel index; j is the spatial position; C is the number of channels; π(i) is the shuffling index;

[0171] In the neck feature fusion stage of the YOLO11 model, the ViT-CSPNeck architecture is introduced. It combines the Transformer self-attention structure to extract global features and works with CSPNeck to complete local feature fusion. The feature selector module FSMFeatureSelectionModule realizes the dynamic weight allocation of features of different scales and channels.

[0172] The ViT branch divides the input image into 16×16 patches and extracts global features through a Transformer encoder. The key mathematical expression of the Transformer encoder module is as follows.

[0173]

[0174] Where z is the output vector, representing the result after multi-head self-attention and residual connection; z1 is the input vector, representing the output of the previous layer or the initial input; LN(·) is layer normalization; MSA(·) is multi-head self-attention, representing the calculation of self-attention on the normalized input to extract contextual information; +z1 is the residual connection, used to retain the original input information and alleviate the gradient vanishing problem; Q is the query matrix, used to calculate the similarity with the keys and determine which parts to focus on; K is the key matrix, used to interact with the query matrix to generate attention weights; V is the value matrix, used to sum the values ​​according to the attention weights to generate the final output; d k The dimension of the key vector is used as a scaling factor to prevent the dot product result from becoming too large and causing gradient instability; QK T For query-key interactions, the similarity between the query and the key is calculated, i.e., the unscaled attention score; Softmax(·) is a normalization function used to transform the attention score into a probability distribution.

[0175] YOLOv11 deep features are extracted using CSPNeck to create multi-scale local features, which are then dynamically weighted and fused with the ViT output using FSM. The dynamic weighting fusion formula for the multi-branch feature fusion module is as follows.

[0176] F fusion =α·ViT(x)+(1-α)·CSPNeck(x),

[0177] α=σ(MLP([F vit ,F csp ]))

[0178] Where σ is the Sigmoid function, with an output range of [0,1]; F frusion The fused feature output is ViT(x); ViT(x) is the feature extraction of input x by VisionTransformer; CSPNeck(x) is the CNN feature extraction module based on CSP structure; α is the state fusion weight, generated by MLP and Sigmoid; [F vit ,F csp [] This is a concatenation operation used to merge ViT features and auxiliary features;

[0179] The YOLO11 model uses Focal-CIoU loss (α = 0.5, γ = 2) in the detection head stage.

[0180]

[0181] Where ρ is the Euclidean distance, c is the length of the diagonal of the minimum bounding rectangle, and 1-IoU is the IoU loss term, which measures the degree of overlap between the predicted box and the ground truth box. The smaller the value, the higher the degree of overlap. Here, b is the normalized distance term for the center point, and b is the center of the predicted bounding box. gt ρ is the center of the true bounding box; c is the diagonal length of the minimum bounding box, normalized distance; αv is the aspect ratio consistency penalty term. The aspect ratio angle of the actual bounding box is used to adjust the aspect ratio of the actual bounding box. Convert to angle values ​​to avoid scale sensitivity; The aspect ratio angle of the prediction box is used to adjust the aspect ratio of the prediction box. Convert to angle values ​​and compare with the actual bounding box angles.

[0182] In the label matching and loss calculation stage of YOLO11 model training, dynamic label allocation (OTA algorithm) is adopted to increase the number of positive samples to 3-5 / GT;

[0183] In the real-time tracking optimization phase of the YOLO11 model, the ByteTrack tracker is used, employing a track management and GIoU tracking matching strategy to achieve integrated detection and track tracking.

[0184] The Kalman filter state equation for the ByteTrack tracker is:

[0185]

[0186] The matching cost matrix uses GIoU (Generalized IoU) as a metric, and the specific formula is as follows:

[0187]

[0188] Where C is the minimum closure region of A and B, IoU is the Intersection over Union (IoU), which calculates the overlap between the predicted bounding box A and the ground truth bounding box B; C\(A∪B) is the portion of region C that does not belong to A or B. Through these improvements, the tracking accuracy can be increased to 97.1%.

[0189] Furthermore, such as Figure 2 and Figure 3 As shown, the motion control module all-terrain robot platform includes a body 1, a head 2, and wheels 3. The body is equipped with a battery 4 and a high-computing-power module 5, which serves as the hardware carrier for the data processing module. The head 2 is located at the head position of the body 1. The head 2 is equipped with a laser radar sensor, a binocular camera sensor, and an infrared camera sensor for the intelligent sensing module. The body is equipped with four wheels 3, and the bottom of the body 1 is equipped with a support pad 7.

[0190] The main body 1 is equipped with a battery 4 and a high-performance computing module 5, with the battery 4 powering the various functional components. The head unit 2 has indicator lights 6. Various sensors are installed on the head unit 2 to collect information. A support pad 7 is used to support the main body 1.

[0191] In this embodiment, the motion control module is a high-strength torso module, which adopts an aerospace-grade hard aluminum alloy and a high-strength engineering plastic body. The dimensions are 70cm×43cm×50cm. It integrates the main control unit, dual battery compartments and equipment interfaces, and has an IP68 protection rating. The bottom is equipped with a modular load-bearing area, which supports 8kg dynamic counterweight adjustment to adapt to the balance requirements of different sensor combinations.

[0192] The motion control module employs a wheel-foot coordinated control algorithm, equipped with 16 high-precision joints. Each joint features a 45 N·m torque motor and 7-inch inflatable tracked feet, supporting wheeled gliding at 2.5 m / s and foot walking at 1.8 m / s. Through reinforcement learning, it optimizes 10 adaptive gait modes (including inverted / tripod modes), achieving dynamic balance by combining an IMU (0.001g accuracy) and foot force sensors. This enables obstacle-crossing capabilities up to a 70cm vertical drop / 35° slope, adapting to the complex environment of ships.

[0193] In summary, the main effective effects of the present invention are as follows:

[0194] 1. Enhanced capabilities for all-scenario operations

[0195] Based on an all-terrain robot platform, the system employs a wheel-foot hybrid design, integrating 16 high-precision joints with 7-inch inflatable tracked feet. It supports wheeled gliding at 2.5 m / s and legged walking at 1.8 m / s, overcoming limitations imposed by 35° slopes and 70cm vertical obstacles. Through reinforcement learning-optimized 10 adaptive gaits (including inverted / tripod modes), coupled with a 0.001g precision IMU system, it achieves dynamic balance in complex shipboard environments, improving efficiency by 320% compared to traditional manual inspections. Its IP68 protection rating and wide operating temperature range of -20℃ to 50℃ allow it to operate continuously for 4.5 hours in high-risk areas such as ballast tanks and fuel tanks, completely eliminating the safety hazards of manual entry into confined spaces.

[0196] 2. Multi-dimensional security monitoring innovation

[0197] A multi-sensor system integrating 4D LiDAR (43,200 points / second), XT-16 LiDAR (640,000 points / second), infrared thermal imaging (0.07℃ sensitivity), and an RGB-D camera (1920×1080@30fps) achieves collaborative monitoring of personnel safety and coating condition through a fusion algorithm with μs-level time synchronization and 0.5mm spatial registration error. An improved YOLOv11 model combined with DeepSeek-7B large model analysis achieves a 99.5% accuracy rate in identifying safety helmets / protective clothing, with a hazardous behavior recognition latency of <200ms, and supports graded early warning based on a probability threshold (>0.7). Real-world testing shows that the system can identify objects as small as 1mm in diameter. 2 Level 1 surface defects and 3cm 2 The internal hollow areas show a defect classification accuracy of 98.5%, far exceeding the limits of manual visual inspection.

[0198] 3. Reconstruction of Intelligent Decision-Making System

[0199] Leveraging the OrinNX platform with 100 TOPS computing power, the system constructs centimeter-level precision 3D semantic maps in real time at a frequency of 15Hz, achieving BIM mapping of defect locations with an error of <5mm through LiDAR SLAM. The automatically generated NACE standard report includes 12 quantization parameters and supports Octree compression (70% compression ratio) storage and AR visualization interaction. An innovative hierarchical processing architecture (QoS priority scheduling + TensorRT acceleration) ensures end-to-end latency of <50ms, allowing managers to remotely adjust detection strategies via 5G / WiFi 6 networks, enabling digital twin-based collaborative decision-making.

[0200] 4. Breakthrough in industrial standard compatibility

[0201] The system strictly adheres to ISO4628 and OSHA standards for defect classification and risk assessment. The infrared thermal imaging module achieves precise temperature measurement from -20℃ to 550℃ based on Planck's law, while the lidar identifies material characteristics through reflection intensity (0-255). Dynamic voxel filtering (5cm grid) and KD tree acceleration technology ensure that the modeling thread maintains a latency of <100ms with 30% computing power utilization, perfectly meeting the real-time requirements of ship maintenance.

[0202] The embodiments of the present invention have been described in detail above, but the content described is only a preferred embodiment of the present invention and should not be considered as limiting the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the patent coverage of the present invention.

Claims

1. A smart inspection robot system for safety of marine engineering coating construction, characterized in that, include: The motion control module is used for all-terrain maneuvering to adapt to the complex environment of a ship. The intelligent sensing module integrates a lidar sensor, a binocular camera sensor, and an infrared camera sensor to collect environmental data. The data processing module is used to perform real-time data analysis and 3D modeling. Through the coordinated operation of the motion control module, intelligent sensing module, and data processing module, the system achieves functions such as personnel safety monitoring, coating status assessment, and three-dimensional environmental reconstruction.

2. The intelligent inspection robot system for marine engineering coating construction safety according to claim 1, characterized in that, The specific working steps of the intelligent inspection robot system for safety inspection of marine engineering coating construction are as follows: S1. Data is collected through the intelligent sensing module; S2. The data processing module preprocesses the data collected by the intelligent sensing module, and uses a multimodal feature-level fusion method to perform spatiotemporal synchronization on the data collected by the intelligent sensing module. The synchronized data is then subjected to data cleaning, point cloud denoising, and multi-sensor calibration extrinsic parameter matrix fusion processing in sequence. S3. The processed data is further processed through the data processing module; The coating defect detection unit in the data processing module performs multispectral fusion of infrared and laser data, defect classification, and NACE report generation; The RGB-D data is processed by the construction safety monitoring unit of the depth vision in the data processing module. The safety status data and dynamic obstacle avoidance instructions are obtained by using YOLOv11 model inference, PPE recognition and dangerous behavior detection. The safety status data and dynamic obstacle avoidance instructions are then transmitted to the real-time alarm unit. The point cloud data and IMU data are processed by the intelligent modeling unit of SLAM in the data processing module. The BIM model is obtained by using SLAM 3D reconstruction, dynamic obstacle removal and BIM mapping. The BIM model is then input into the 3D semantic model for compression and storage.

3. The intelligent inspection robot system for marine engineering coating construction safety according to claim 2, characterized in that: The binocular camera sensor has a binocular vision ranging module. In S1, the binocular vision ranging module performs distortion correction and epipolar correction on the image data, and performs disparity calculation and depth calculation on the data after distortion correction and epipolar correction. Distortion correction uses the Brown-Conrady model to correct lens distortion, and the correction formula is as follows: Where k1, k2, and k3 are radial distortion coefficients, used to control the strength and order of barrel or pincushion distortion; p1 and p2 are tangential distortion coefficients, used to represent the offset caused by the imperfect coplanarity of the lens and sensor; x and y are normalized coordinates; Xcorrecet and Ycorrecet are the distorted coordinate values; and r is the radius to the imaging center. Epipolar correction uses the Bouguet algorithm to calculate homography matrices H_1 and H_2, making the left and right images coplanar, and matching points only need to be searched along the horizontal direction. Disparity calculation employs a semi-global matching SGBM algorithm to optimize the matching cost function. Where w(p,q) is the adaptive weight; ρ is the robust penalty function; I L I R , respectively, are the function values ​​of the left and right images; C(p,d) is the matching cost function; The formula for depth calculation is as follows: Where z is the final depth, d is the parallax, b is the baseline distance, and f is the focal length in pixels.

4. The intelligent inspection robot system for marine engineering coating construction safety according to claim 2, characterized in that, The infrared camera sensor contains an infrared thermal imaging module. The specific operating steps of the infrared thermal imaging module in S1 include: The radiation signal is converted, and the detector output voltage signal V of the infrared camera sensor is converted. out The relationship with the target radiated energy is as follows: Where R is the detector responsivity (A / W), ∈(λ) is the target emissivity, and M λ (T) is Planck's formula for blackbody radiation, n th Thermal noise; Noise equivalent temperature difference (NETD) calibration was performed using a blackbody source. The calculation formula is as follows: Where σnoise is the noise standard deviation, S is the responsivity, and ΔT is the blackbody temperature change. Temperature inversion is performed, and the target temperature T is calculated based on the Stefan-Boltzmann law. Where E is the radiation energy and σ is the Stefan-Boltzmann constant.

5. The intelligent inspection robot system for marine engineering coating construction safety according to claim 2, characterized in that, The lidar has a built-in lidar point cloud processing module. The specific steps of the lidar point cloud processing module in S1 include: For ToF ranging, the relationship between the laser pulse round-trip time Δt and the target distance d is as follows: Where c is the speed of light, n air n is the refractive index of air. vacuum The refractive index is the vacuum refractive index. To compensate for system errors, the actual ranging time needs to be reduced by the system delay t. system_delay , Δt actual =Δt measured -t system_delay Where, Δt actual Δt represents the actual distance measurement time. measured For measuring time; t system_delay This is due to system latency; MEMS scanning control is performed, and the deflection angle θ(t) of the micromirror is controlled by a sinusoidal drive signal. θ(t)=θ max ·sin(2πf scan t) Among them, f scan The scanning frequency determines the point cloud sampling rate; θ max This represents the maximum deflection angle of the micromirror.

6. The intelligent inspection robot system for safety of marine engineering coating construction according to claim 2, characterized in that, The specific steps of the multimodal feature-level fusion method in S2 include: A1. Perform external parameter calibration, and obtain the transformation matrix T between the lidar sensor and the binocular camera sensor by calibrating the target. lidar→camera The transformation matrix T between the infrared camera sensor and the stereo camera sensor Thermal→camera ; A2. Time synchronization is performed using the PTP precise time protocol to achieve μs-level synchronization and align the data from each sensor set on the intelligent sensing module with time. A3. Perform feature fusion and dynamically adjust the weights of multimodal features using an attention mechanism. F fusion =α·CNN(I RGB )+β·GNN(P Lidar )+γ·Grad(T Thermal ) Where α, β, and γ are adaptive weight coefficients; F is the fused comprehensive feature representation used in subsequent tasks; CNN(I RGB ) represents the RGB image features extracted by a convolutional neural network; I RGB The input is an RGB color image; GNN(P) Lidar P represents the LiDAR point cloud features extracted using a graph neural network. Lidar 3D point cloud data acquired by LiDAR; Grad(T) Thermal This is used to extract gradients or structural features from thermal imaging data to emphasize target and texture details; T Thermal This is thermal imaging infrared image data.

7. The intelligent inspection robot system for marine engineering coating construction safety according to claim 2, characterized in that: The data processing module adopts the YOLO11 model, replacing the original ELAN structure in the YOLO11 model with the MobileOne-S4 module. It constructs an efficient feature extraction unit using 3×3 and 1×1 convolutions combined with BN and SiLU activation, and introduces an Identity branch. The formula for the MobileOne-S4 module is as follows: W fused =W 3×3 +W 1×1 +I Among them, W 3×3 For a 3×3 convolution kernel, W 1×1 The kernel is a 1×1 convolution, and I is the Identity mapping; The channel shuffling mechanism of ShuffleNetV2 is introduced. The formula for the channel shuffling mechanism of ShuffleNetV2 is as follows: Shuffle(x) i,j =x π(i),j ,π(i)=(i+j)modC Where Shuffle(x) i,j For channel shuffling operation; x represents the characteristics of the channel before shuffling; i is the channel index; j is the spatial position; C is the number of channels; π(i) is the shuffling index.

8. The intelligent inspection robot system for marine engineering coating construction safety according to claim 7, characterized in that: In the neck feature fusion stage of the YOLO11 model, the ViT-CSPNeck architecture is introduced. It combines the Transformer self-attention structure to extract global features and works with CSPNeck to complete local feature fusion. The feature selector module FSMFeatureSelectionModule realizes the dynamic weight allocation of features of different scales and channels. The ViT branch divides the input image into 16×16 patches and extracts global features through a Transformer encoder. The key mathematical expression of the Transformer encoder module is as follows. z l =MSA(LN(z l-1 ))+z l-1 , Where z is the output vector, representing the result after multi-head self-attention and residual connection; z l The input vector represents the output or initial input of the previous layer; LN(·) is layer normalization; MSA(·) is multi-head self-attention, which calculates self-attention on the normalized input to extract contextual information; z l-1 For residual connections, the original input information is preserved to alleviate the vanishing gradient problem; Q is the query matrix, used to calculate the similarity with the keys and determine which parts to focus on; K is the key matrix, used to interact with the query matrix to generate attention weights; V is the value matrix, used to sum the values ​​according to the attention weights to generate the final output; d k The dimension of the key vector is used as a scaling factor to prevent the dot product result from becoming too large and causing gradient instability; QK T For query-key interactions, the similarity between the query and the key is calculated, i.e., the unscaled attention score; Softmax(·) is a normalization function used to transform the attention score into a probability distribution.

9. The intelligent inspection robot system for marine engineering coating construction safety according to claim 8, characterized in that: YOLOv11 deep features are extracted using CSPNeck to create multi-scale local features, which are then dynamically weighted and fused with the ViT output using FSM. The dynamic weighting fusion formula for the multi-branch feature fusion module is as follows. α=σ(MLP([F vit ,F csp ])) Where σ is the Sigmoid function, with an output range of [0,1]; F frusion The fused feature output is ViT(x); ViT(x) is the feature extraction of input x by VisionTransformer; CSPNeck(x) is the CNN feature extraction module based on CSP structure; α is the state fusion weight, generated by MLP and Sigmoid; [F vit ,F csp [] This is a concatenation operation used to merge ViT features and auxiliary features; The YOLO11 model uses Focal-CIoU loss in the detection head stage, with the specific formula as follows. Where ρ is the Euclidean distance, c is the length of the diagonal of the minimum bounding rectangle, and 1-IoU is the IoU loss term, which is used to measure the degree of overlap between the predicted box and the ground truth box. Here, b is the normalized distance term for the center point, and b is the center of the predicted bounding box. gt ρ is the center of the true bounding box, ρ is the Euclidean distance, and αv is the aspect ratio consistency penalty term. The aspect ratio angle of the actual bounding box is used to adjust the aspect ratio of the actual bounding box. Convert to angle value; The aspect ratio angle of the prediction box is used to adjust the aspect ratio of the prediction box. Convert to angle values ​​and compare with the actual bounding box angles; In the label matching and loss calculation stage of YOLO11 model training, dynamic label allocation is adopted to increase the number of positive samples to 3-5 / GT; In the real-time tracking optimization phase of the YOLO11 model, the ByteTrack tracker is used, employing a track management and GIoU tracking matching strategy to achieve integrated detection and track tracking. The Kalman filter state equation for the ByteTrack tracker is: The matching cost matrix uses GIoU as a metric, and the specific formula is as follows. Where C is the minimum closure region of A and B, IoU is the intersection-union ratio, and the overlap between the predicted box A and the ground truth box B is calculated; C\(A∪B) is the part of region C that does not belong to A or B.

10. The intelligent inspection robot system for marine engineering coating construction safety according to any one of claims 1-9, characterized in that: The motion control module is an all-terrain robot platform, including a body, a head, and wheels. The body is equipped with a battery and a high-performance computing module, which serves as the hardware carrier for the data processing module. The head is located at the front of the body and is equipped with a lidar sensor, a binocular camera sensor, and an infrared camera sensor for the intelligent sensing module. The body has four wheels, and the bottom of the body has support pads.