A multi-modal deep-sea net cage AI autonomous inspection and cleaning system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-08-11
AI Technical Summary
这类设备无法实时感知和识别网衣上的污损类型、程度及网衣破损情况,更无法根据作业现场复杂多变的水流环境动态调整自身姿态和清洗力度
[0014] Technical Effects: This invention solves the core problems of blind washing and insufficient anti-flow capability in existing technologies by equipping the robot with a multimodal fusion perception system integrating sonar, camera, lidar, and water flow sensors, and a six-degree-of-freedom motion control module containing an anti-flow adaptive controller, and deeply coupling the two through an AI decision module. Its innovative technical points are: the AI decision module performs nonlinear compensation and feature fusion on multi-source perception data to identify the soiling and damage status of the mesh in real time and generate operation decisions including paths and parameters; simultaneously, the anti-flow adaptive controller directly uses flow field data to calculate thruster commands, enabling the robot to resist water flow disturbances and maintain a stable operating posture. This allows the cleaning operation to no longer be a fixed-pattern mechanical movement, but an adaptive intelligent behavior based on environmental perception and its own state, achieving the goals of protecting the mesh and improving cleaning efficiency.
Smart Images

Figure CN122540347A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater operation tools, specifically to a multimodal deep-sea cage AI autonomous inspection and cleaning system. Background Technology
[0002] As deep-sea cage aquaculture expands into deeper waters, the cages, constantly submerged in seawater, become highly susceptible to fouling by marine organisms. This fouling leads to clogged mesh, impaired water exchange, and even serious incidents like fish escaping due to mesh damage. Therefore, regular inspection and cleaning of the cages are crucial for ensuring aquaculture safety. Currently, the inspection and cleaning of deep-sea cages primarily rely on manual diving or single-function remotely controlled cleaning robots. However, manual cleaning is inefficient, risky, and extremely costly. While existing cleaning robots can crawl and clean the mesh, their sensing and control capabilities are severely lacking, often resulting in blind cleaning or control based on simple thresholds. These devices cannot perceive or identify the type and extent of fouling or mesh damage in real time, nor can they dynamically adjust their posture and cleaning intensity according to the complex and changing water flow environment at the work site. When faced with multidirectional and turbulent currents common in the deep sea, the robot's posture is easily unstable. The contact force between the cleaning disc and the net is either too large and damages the net, or too small and cannot effectively remove dirt, resulting in poor cleaning effect, large damage to the net, and weak resistance to current.
[0003] Therefore, there is an urgent need for an intelligent inspection and cleaning system that can autonomously sense the status of netting and adaptively adjust its operation strategy in complex water flow environments. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a multimodal autonomous inspection and cleaning system for deep-sea cages, comprising a main frame, a multi-sensor fusion module, an AI decision-making module, an actuator module, and a communication module. The system also includes a six-degree-of-freedom motion control module, which comprises a forward / reverse drive system, a lateral drive system, a heel drive system, a roll angle drive system, a pitch angle drive system, and a yaw angle drive system. The six-degree-of-freedom motion control module also includes a flow-resistant adaptive controller. The multi-sensor fusion module includes a sonar sensor, a camera sensor, a lidar sensor, and a water flow sensor, which are used to output acoustic data, optical image data, point cloud data, and flow field data respectively. The output of the water flow sensor is connected to the AI decision module and the anti-flow adaptive controller, respectively. The flow field data serves as the environmental perception input of the AI decision module and the real-time control input of the anti-flow adaptive controller. The AI decision-making module is used to acquire the acoustic data, the optical image data, the point cloud data, and the flow field data, and to perform nonlinear compensation and feature fusion on the four types of data. Based on the fused perception features, it identifies the contaminated area of the net cage and the damaged area of the net, and generates an operation decision instruction containing the cleaning path and cleaning parameters. The actuator module is used to receive the operation decision instruction and adjust the working parameters of the cleaning disc and / or the gripping mechanism; The anti-flow adaptive controller is used to acquire the flow field data and calculate the thruster commands based on the flow field data to drive the various drive systems of the six-degree-of-freedom motion control module. The anti-flow adaptive controller is also used to feed back the actual motion state to the AI decision module, and the AI decision module updates the environmental perception parameters and operation decision instructions according to the actual motion state.
[0005] Preferably, the AI decision-making module includes an environmental perception layer, a feature extraction layer, a decision-making layer, and a behavior planning layer; The environmental perception layer is used to receive the acoustic data, the optical image data, the point cloud data, and the flow field data, and to model the nonlinear noise caused by environmental interference, modeling errors, and sensor noise through a nonlinear mapping function, calculate the compensation amount to cancel the nonlinear noise, and superimpose the compensation amount into the original control signal. The feature extraction layer is used to perform dimensionality reduction processing on the acoustic data, optical image data, point cloud data and flow field data after superimposing the compensation amount using the principal component analysis method, and outputs the dimensionality-reduced perception feature set; The decision layer integrates the lightweight compressed deep learning model. The decision layer is used to obtain the dimensionality-reduced perceptual feature set and identify the soiled areas of the net cage and the damaged areas of the netting. The behavior planning layer is used to plan the cleaning path based on the distribution of the soiled areas of the net cage and the damaged areas of the netting using a fuzzy C-means clustering algorithm, and to determine the cleaning parameters based on the type and degree of soiling of the net cage.
[0006] More preferably, the forward / reverse drive system, the lateral drive system, the heave drive system, the roll angle drive system, the pitch angle drive system, and the yaw angle drive system are each composed of a permanent magnet synchronous motor and a planetary gear reducer modular propeller, and each modular propeller is installed on the main frame through a standard electrical interface and a mechanical quick-connect structure.
[0007] More preferably, the lightweight compressed deep learning model is a network model obtained by channel pruning and weight quantization based on the ResNet-18 network structure. The lightweight deep learning model is deployed in the underwater embedded edge computing unit built into the AI decision module. The parameter compression rate of the network model is not less than 60%, and the inference frame rate of the network model after channel pruning and weight quantization in the AI decision module is not less than 10 frames / second.
[0008] More preferably, the nonlinear mapping relationship on which the environmental perception layer calculates the compensation amount, and the weights on which the feature extraction layer fuses the acoustic data, the optical image data, the point cloud data, and the flow field data, are all obtained through supervised learning training on multiple sets of sensor noise and cage status training data in a water environment with specific turbidity, temperature, and salinity parameter ranges. The preset limited interval is stored in the built-in storage unit of the AI decision module.
[0009] More preferably, the feature extraction layer uses principal component analysis to perform dimensionality reduction and fusion on the acoustic data, the optical image data, the point cloud data, and the flow field data, and dynamically determines the number of principal components to be retained after dimensionality reduction based on the water turbidity, temperature, and salinity parameters collected by the water flow sensor; the dynamic determination rule for the number of principal components to be retained is: when the water turbidity is higher than a first threshold, the number of retained principal components is increased by a first increment to maintain the effective feature dimension of the optical image data.
[0010] More preferably, the behavior planning layer sends the planned cleanup path and the cleanup parameters to the anti-flow adaptive controller. The anti-flow adaptive controller calculates the thruster command acting on the six-degree-of-freedom motion control module based on the flow field data and feeds back the actual motion state to the AI decision module. The AI decision module updates the nonlinear compensation model parameters of the environment perception layer based on the deviation between the actual motion state and the expected motion state, and re-outputs the updated compensation amount to the feature extraction layer.
[0011] More preferably, the AI decision module is also used to determine the effective coefficient of the sensing features based on the time delay of the acoustic data, the attenuation coefficient of the sound wave in the water, the time delay of the optical image data, the turbidity scattering coefficient of the water, the time delay of the point cloud data, and the scattering coefficient of suspended particles in the water, and to calculate the comprehensive resistance of the system during operation based on the seawater density, the system velocity relative to the water, the effective upstream area of the system, the local static strain of the net, the elastic modulus of the net, the ratio of the local deformation of the net to the original length of the deformed section, and the static water foundation resistance.
[0012] More preferably, the AI decision module is also used to determine the edge inference computing power ratio based on the effective coefficient of the perception features, the comprehensive resistance, the maximum design cruise speed of the system and the total rated drive power of the six-degree-of-freedom motion control module, and to determine the optimal cleaning operation force based on the edge inference computing power ratio, the quantification index of the cage fouling degree, the static pressure of the cleaning medium and the effective working area of the cleaning disc, and the actuator module is used to adjust the rotation speed and downward pressure of the cleaning disc according to the optimal cleaning operation force.
[0013] More preferably, the forward / reverse drive system, the yaw angle drive system, and the heave drive system all include a permanent magnet synchronous motor and a planetary gear reducer. The anti-current adaptive controller is embedded in the local controller corresponding to each of the forward / reverse drive system, the yaw angle drive system, and the heave drive system. The anti-current adaptive controller stores an adaptive PID parameter adjustment algorithm.
[0014] Technical Effects: This invention solves the core problems of blind washing and insufficient anti-flow capability in existing technologies by equipping the robot with a multimodal fusion perception system integrating sonar, camera, lidar, and water flow sensors, and a six-degree-of-freedom motion control module containing an anti-flow adaptive controller, and deeply coupling the two through an AI decision module. Its innovative technical points are: the AI decision module performs nonlinear compensation and feature fusion on multi-source perception data to identify the soiling and damage status of the mesh in real time and generate operation decisions including paths and parameters; simultaneously, the anti-flow adaptive controller directly uses flow field data to calculate thruster commands, enabling the robot to resist water flow disturbances and maintain a stable operating posture. This allows the cleaning operation to no longer be a fixed-pattern mechanical movement, but an adaptive intelligent behavior based on environmental perception and its own state, achieving the goals of protecting the mesh and improving cleaning efficiency. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall module architecture of the multimodal deep-sea cage AI autonomous inspection and cleaning system; Figure 2 This is a schematic diagram of the hardware physical layout of an AI-powered autonomous inspection and cleaning device for multimodal deep-sea cages. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] Traditional deep-sea cage cleaning equipment suffers from limited sensing dimensions and simple control methods, making it unable to adaptively adjust its operating strategy based on the actual fouling state of the nets in turbid and multi-directional water environments, resulting in limited net damage and cleaning efficiency.
[0018] Based on this, please refer to Figures 1-2 This embodiment provides a multimodal autonomous inspection and cleaning system for deep-sea cages. The system includes a main frame, a six-degree-of-freedom motion control module, a multi-sensor fusion module, an AI decision-making module, an actuator module, and a communication module. The connections between the modules are as follows: the output of the multi-sensor fusion module is connected to the input of the AI decision-making module; the first output of the AI decision-making module is connected to the input of the actuator module; the second output of the AI decision-making module is connected to the input of the current-resistant adaptive controller of the six-degree-of-freedom motion control module; the output of the current-resistant adaptive controller is connected to the control terminals of the advance / retreat drive system, yaw angle drive system, lateral movement drive system, pitch angle drive system, roll angle drive system, and heave drive system, respectively; the output of the flow sensor is simultaneously connected to both the AI decision-making module and the current-resistant adaptive controller, forming a dual-flow-field data dual-closed-loop coupling architecture. The perception compensation closed loop and the motion control closed loop operate in parallel and make collaborative decisions; the communication module is bidirectionally connected to the AI decision-making module.
[0019] The main frame consists of an upper mounting plate, a lower mounting plate, and longitudinal support rods connecting the two. The upper mounting plate is made of aluminum alloy sheet with a thickness of 5 to 10 mm, and the lower mounting plate is also made of aluminum alloy with the same thickness. The longitudinal support rods are made of stainless steel tubing with a diameter of 50 to 80 mm. Both the upper and lower mounting plates have counterweight mounting positions for installing counterweights. The counterweight mounting positions are rectangular grooves formed on the edge of the plates, and the shape of the grooves is adapted to the shape of the counterweights to facilitate the insertion and removal of the counterweights. The longitudinal support rods are equipped with a pitch angle adjustment mechanism for adjusting the pitch angle of the main frame. This pitch angle adjustment mechanism uses an electric screw lifting structure with a lifting stroke of 0 to 100 mm. A stepper motor drives the screw to rotate, causing the nut seat that mates with the screw to move along the axis of the longitudinal support rod, thereby changing the relative angle between the upper and lower mounting plates. When the pitch angle adjustment mechanism is in the zero position, the upper mounting plate is parallel to the lower mounting plate; when the lead screw pushes one side of the upper mounting plate to rise, the main frame as a whole generates a pitch angle, ranging from 0 degrees to 15 degrees.
[0020] By adjusting the fixed position of the counterweight in its mounting position and the lifting stroke of the pitch angle adjustment mechanism, combined with the thrust vector distribution of the six-degree-of-freedom motion control module, the overall center of buoyancy and center of gravity of the system can be aligned vertically, and a stable horizontal posture can be achieved when the center of gravity is lower than the center of buoyancy. When switching between bottom net cleaning and side net cleaning modes, the mounting position of the counterweight at the top or bottom is reconfigured. At the same time, the pitch angle adjustment mechanism is used to change the attitude angle of the main frame, and the thrust distribution matrix of each thruster is adjusted, so that the robot can stably adhere to the net surface in either a horizontal bottom-adhering or vertical wall-adhering posture.
[0021] For bottom net cleaning, all counterweights are installed in the counterweight mounting positions on the lower mounting plate, bringing the robot's center of gravity close to the bottom and placing it in a horizontally suspended state with the cleaning disc facing downwards. The thrust distribution matrix of the six-degree-of-freedom motion control module automatically switches to bottom net mode: the two transverse propellers of the lateral drive system provide the main downward pressure against the net, the vertical propeller of the yaw drive system assists in maintaining attitude balance, and the longitudinal propeller of the forward / backward drive system is responsible for propulsion along the net surface.
[0022] For side net cleaning, remove the robot from the water surface, move the counterweight from the lower mounting plate and install it in the counterweight mounting position on the upper mounting plate, and adjust the pitch angle adjustment mechanism to make the robot stand upright with the cleaning disc facing the side net. The robot is then attached to the side net surface in a vertical posture.
[0023] The thrust distribution matrix automatically switches to side net mode: the central vertical propeller of the yaw angle drive system provides the main net-attaching pressure, the forward and backward longitudinal propellers of the advance / reverse drive system are responsible for vertical movement along the side net, and the lateral propellers of the lateral drive system are responsible for lateral translation. This dual-mode attitude switching mechanism, which combines mechanical counterweight adjustment with electrical thrust distribution, enables the same robot to adapt to two completely different working conditions—bottom net and side net—without requiring additional replacement of mechanical components.
[0024] The six-degree-of-freedom motion control module includes an anti-current adaptive controller, a forward / backward drive system, a yaw drive system, a lateral drive system, a pitch drive system, a roll drive system, and a heave drive system. The forward / backward drive system consists of two propellers arranged front and rear, mounted at the front and rear ends of the main frame, respectively. The axes of the two propellers are parallel to the robot's longitudinal axis, achieving forward and backward movement through forward and reverse rotation. The yaw drive system consists of three vertically oriented propellers, one mounted at the geometric center of the main frame, and the other two mounted at the front and rear, respectively. The axes of the three propellers are perpendicular to the horizontal plane, achieving in-situ turning and heading maintenance by coordinating the thrust of the three propellers. The lateral drive system consists of two laterally arranged propellers, mounted on the left and right sides of the main frame, respectively. The propeller axes are parallel to the robot's lateral axis, used to control the robot's left and right lateral movement. The pitch angle drive system is installed on the top front side of the main frame, the roll angle drive system is installed on the top left side of the main frame, and the heave drive system is installed on the top rear side of the main frame. All three are small-thrust propellers used to precisely adjust the robot's pitch, roll, and depth attitude.
[0025] Each drive system employs a modular propeller consisting of a permanent magnet synchronous motor and a planetary gear reducer. The permanent magnet synchronous motor is a BL57 series model, with a rated power of 200 to 400 watts, a rated speed of 3000 rpm, and a peak torque of 1.2 Nm. The planetary gear reducer has a reduction ratio of 20 to 50. The input shaft is connected to the output shaft of the permanent magnet synchronous motor via a coupling, and the output shaft is connected to the propeller hub via a key. The propeller adopts a three-bladed design, with blades made of fiber-reinforced nylon, a blade diameter of 150 to 250 mm, and a pitch ratio of 0.8 to 1.2. Under rated operating conditions, a single propeller can provide 80 to 160 Newtons of thrust.
[0026] The anti-current adaptive controller is embedded in the local controller of each thruster and stores an adaptive PID parameter adjustment algorithm. The local controller of each thruster uses an STM32F407 series microcontroller with a main frequency of 168 MHz, a built-in floating-point unit, a 12-bit analog-to-digital converter, and a controller area network interface.
[0027] The adaptive PID parameter adjustment algorithm employs a feedforward-feedback dual-loop control architecture: the outer loop is a flow field feedforward compensation loop, and the inner loop is an attitude feedback PID control loop. The feedforward loop rapidly calculates the compensation thrust based on real-time flow field data from the flow sensor, suppressing the initial impact of flow disturbances; the feedback loop precisely corrects the deviation between the robot's actual attitude and the target attitude, eliminating residual errors from the feedforward compensation. The dual-loop operation ensures both rapid anti-flow response and precise attitude control.
[0028] The flow sensor is mounted at the center of the bottom of the main frame, 30 mm from the surface of the bottom mounting plate, and is fixed with a stainless steel bracket. The flow sensor employs an acoustic Doppler current profiler, operating at a frequency of 600 kHz, with a flow velocity measurement range of 0 to 3 m / s, a measurement accuracy of ±0.01 m / s, a flow direction measurement accuracy of ±2 degrees, and an output frequency of 10 Hz. The flow sensor collects flow field data including flow velocity and direction, and transmits it via a CAN bus at an update rate of 1 kHz to the anti-flow adaptive controller and AI decision module.
[0029] After acquiring the flow field data, the anti-flow adaptive controller processes it according to the feedforward-feedback dual-loop control architecture as follows: The first step is flow field vector decomposition and feedforward compensation calculation. The anti-flow adaptive controller decomposes the velocity vector in the flow field data into three components along the robot's longitudinal, lateral, and vertical directions. Each component is multiplied by a preset feedforward gain coefficient to obtain the corresponding feedforward compensation thrust value. The feedforward gain coefficient is calibrated through a water tank towing test. Different attitude modes correspond to different gain matrices to adapt to the differences in water flow forces under different upstream areas.
[0030] The second step is attitude deviation feedback calculation. The anti-current adaptive controller calculates the feedback correction thrust value based on the deviation between the robot's current actual attitude and the target attitude using a PID controller. The PID controller parameters employ an adaptive adjustment mechanism: the proportional coefficient is dynamically adjusted according to the rate of change of the flow velocity in the flow field data. For every 0.1 m / s² increase in the rate of change, Increase by 5; dynamically adjust the integral coefficient based on the integral value of the robot's position deviation. For every 0.1 m / s increase in the integral value, Decrease by 1; dynamically adjust the differential coefficients based on the differential value of the position deviation. For every 0.05 meters per second increase in the differential value, Add 2. This adjustment rule is implemented through a fuzzy control table, eliminating the need for online solving of complex optimization problems and meeting millisecond-level real-time requirements.
[0031] The third step is thrust distribution calculation. The feedforward compensation thrust value and the feedback correction thrust value are added to obtain the total thrust requirement in each direction. The total thrust requirement is then distributed according to the thruster spatial layout matrix to calculate the thrust value that each thruster should output. Finally, the thrust value is converted into a speed command based on the thrust-speed characteristic curve of the thruster. The thrust distribution matrix is a 6×6 pseudo-inverse matrix, corresponding to the mapping relationship between the six thrusters and the six degrees of freedom, ensuring fault-tolerant control even in the event of a partial thruster failure.
[0032] Each drive system generates reverse thrust to counteract the impact of the water flow. An adaptive PID algorithm updates the flow field data and motor encoder feedback signal in real time at a frequency of 1 kHz, dynamically adjusting the proportional, integral, and derivative parameters. When the ocean current velocity and direction change abruptly, the feedforward compensation is applied to the thruster output immediately to suppress robot pose deviation.
[0033] Example 1: At a depth of 40 meters on the seabed, a robot is cleaning a bottom net. The current sensor measures a flow velocity of 1.5 m / s, with the flow direction forming a 30-degree angle with the robot's longitudinal axis. The current-resistant adaptive controller first enters the feedforward loop: decomposing the flow velocity into a longitudinal component of 1.3 m / s and a lateral component of 0.75 m / s. Based on the feedforward gain coefficient matrix in the bottom net mode, the longitudinal feedforward compensation thrust is calculated to be 104 Newtons, and the lateral feedforward compensation thrust is calculated to be 45 Newtons. Simultaneously, it enters the feedback loop: due to the impact of the water flow causing a 2 cm positional deviation and a 0.5-degree attitude angle deviation, the adaptive PID controller outputs a longitudinal feedback correction thrust of 30 Newtons, a lateral feedback correction thrust of 15 Newtons, and a pitch correction torque of 5 Newtons. The total thrust requirement is 134 Newtons longitudinally, 60 Newtons laterally, and a pitch torque of 5 Newtons. After thrust distribution calculation, the forward and backward drive system outputs 70 Newtons from the front propeller and 64 Newtons from the rear propeller. The lateral drive system outputs 32 Newtons from the left propeller and 28 Newtons from the right propeller. The pitch drive system outputs 5 Newtons of thrust to generate a corrective torque. Each thruster adjusts its rotational speed accordingly to ensure the robot maintains a positional deviation of no more than 3 centimeters and an attitude angle deviation of no more than 0.3 degrees in a crossflow of 1.5 meters per second.
[0034] The multi-sensor fusion module includes a sonar sensor, a camera sensor, a lidar sensor, and the water flow sensor. Each sensor achieves time synchronization via a hardware trigger signal, with a synchronization accuracy better than 1 millisecond. The AI decision module assigns a unified timestamp and spatial pose label to each sensor's data frame, ensuring precise alignment of multimodal data in the spatiotemporal dimensions and providing a reliable foundation for subsequent feature fusion.
[0035] The sonar sensor employs a compressed ultrasonic sensor with a center frequency of 100-200 kHz and a bandwidth of 40 MHz. It is mounted on the top front of the main frame and secured by a gimbal bracket, allowing for ±30 degrees of elevation angle adjustment. The sonar sensor emits an ultrasonic beam angle of 3 degrees, with a range resolution of 5 mm and a detection range of 50 meters. In turbid water, although the sound waves are attenuated, they can still penetrate the suspended particle layer to obtain information on the macroscopic outline of the net cage and the distribution of fouling. The sonar sensor operates in active mode, emitting a single ultrasonic pulse with a pulse width of 0.1 ms each time, then switching to receiving mode to receive echoes reflected from the net cage and fouling organisms. The echo signals are bandpass filtered, envelope detected, and logarithmically amplified before being converted into digital signals by an analog-to-digital converter at a sampling rate of 2 MHz, forming acoustic data. The acoustic data includes a distance-intensity sequence for each detection angle, with a data rate of approximately 200 kilobytes per second.
[0036] The camera sensor is a 2-megapixel marine-environment-adaptive camera, mounted on the lower front of the main frame, with the lens facing forward and a field of view of 90 degrees horizontally and 70 degrees vertically. The camera is equipped with an infrared illumination module consisting of eight 850nm infrared LEDs with a total radiant power of 10 watts, providing effective illumination up to 5 meters in complete underwater darkness. The camera uses a global shutter CMOS sensor with a pixel size of 3.75μm x 3.75μm, a sensitivity of 2000 mV / lux·s, and supports automatic exposure and automatic white balance. The camera outputs 1920 x 1080 pixel color images at a frame rate of 30 frames per second and a data rate of approximately 180 megabytes per second. Image data is transmitted to a signal processor via a MIPI interface. The signal processor performs preprocessing on the image, including dehazing, sharpening, and color correction. The preprocessed optical image data is then sent to the AI decision-making module via Gigabit Ethernet.
[0037] The lidar sensor uses a 2D lidar model, Lidar-L3HR, mounted at the top center of the main frame and protected by a transparent hemispherical shield. The lidar achieves 360-degree omnidirectional scanning through mechanical rotation at a speed of 10 Hz, completing one scan every 0.1 seconds. The laser emits infrared pulsed laser light with a wavelength of 905 nm, a pulse repetition frequency of 20 kHz, and a single pulse energy of 1 microjoule. The lidar has an angular resolution of 0.5 degrees, meaning it can acquire 720 measurement points per rotation, each containing distance and reflection intensity information. In clear water, the lidar's maximum detection range is 30 meters; in water with a turbidity of 2 NTU, the detection range decreases to 10 meters. The lidar sensor is used to generate localized high-precision 3D point clouds. By combining the 2D scanning plane with the robot's lifting motion, the 3D topographic data of the mesh surface is obtained. The point cloud data has a data rate of approximately 50 kilobytes per second, and the data format includes XYZ coordinates and intensity values for each point.
[0038] The spatial installation positions of each sensor are precisely calibrated, and the calibration parameters are stored in the calibration parameter file of the AI decision module. The calibration process includes: using known marker points on the cage as references, solving the transformation matrix between the coordinate systems of each sensor and the robot's base coordinate system using a hand-eye calibration algorithm. The calibration accuracy is better than 2 millimeters, ensuring accurate spatial registration of multimodal data.
[0039] Each sensor is waterproof and equipped with a signal processor. The sonar sensor's transducer is polyurethane-encapsulated for waterproofing, with a pressure resistance depth of 100 meters; the camera sensor's housing is made of 316L stainless steel, and the viewing window is made of sapphire glass, sealed with O-rings and screws; the lidar sensor's protective cover is made of polycarbonate, and waterproofing is achieved between it and the base using a silicone sealing ring. Each signal processor filters, reduces noise, and formats the raw data, outputting acoustic data, optical image data, point cloud data, and the aforementioned flow field data, which are then aggregated via a gigabit Ethernet switch and sent to the AI decision module.
[0040] The AI decision-making module runs on an embedded industrial control computer equipped with an Intel Core i7 processor and 16GB of DDR4 memory, running the Ubuntu operating system and the ROS robot operating system framework. Internally, the AI decision-making module is divided into a four-layer architecture: environmental perception, feature extraction, decision-making, and behavior planning. These layers interact through standardized interfaces, and each layer is bound to a specific combination of algorithms, forming a complete intelligent chain from perception to extraction, decision-making, and planning. This deep integration of the four-layer architecture with specific algorithms is not a simple modular stacking, but a customized design for the specific scenario of deep-sea cage inspection and cleaning. The algorithms at each layer collaborate and exchange parameters to achieve efficient and accurate intelligent decision-making.
[0041] The environmental perception layer receives the acoustic data, optical image data, point cloud data, and flow field data. Because suspended particles in turbid deep-sea water attenuate sound waves and scatter laser light, and water flow fluctuations and robot vibrations introduce sensor noise, the environmental perception layer models this nonlinear noise using a pre-trained nonlinear mapping function. This nonlinear mapping function is implemented using a three-layer fully connected neural network. The number of neurons in the input layer equals the feature dimension of the sensor data, the hidden layer contains 128 neurons using the ReLU activation function, and the output layer has the same number of neurons as the input layer. The total number of network parameters is approximately 10,000. This network takes the original sensor data, the current turbidity parameter, and the robot's motion state as input and outputs a set of compensation values. These compensation values are superimposed on the original control signal to compensate for signal distortion introduced by environmental interference, modeling errors, and sensor noise.
[0042] For acoustic data compensation, the environmental perception layer extracts the echo intensity sequence for each detection angle from the acoustic data. Based on the current water turbidity parameters, it consults a pre-calibrated sound attenuation compensation curve and performs distance and attenuation compensation on the echo intensity, making the compensated echo intensity approach the theoretical value under clear water conditions. For optical image data compensation, the environmental perception layer performs dark channel prior dehazing on the image, estimates the water transmittance map and atmospheric light values, and uses the transmittance map to restore the color of each pixel, thereby eliminating image whitening and contrast reduction caused by water turbidity. For point cloud data compensation, the environmental perception layer uses a random sampling consensus algorithm to remove outliers from the point cloud. Outliers are mainly caused by the instantaneous reflection of large suspended particles in the water. A density clustering filter with a radius of 0.1 meters and a minimum number of points of 10 is used for secondary filtering. For flow field data compensation, the environmental perception layer uses a Kalman filter algorithm to smooth the flow velocity measurements from the water flow sensor, eliminating high-frequency fluctuation components caused by water turbulence.
[0043] The nonlinear mapping relationship was obtained through supervised learning. The training data covered aquatic environments with specific turbidity, temperature, and salinity parameters, and included multiple sets of labeled samples with known sensor noise and cage conditions. During training, sensor data collected in turbid water was used as input, and sensor data collected in the same scenario in clear water was used as labels. The network weights of the nonlinear mapping function were iteratively updated by minimizing the mean square error between the compensated data and the labeled data. The training dataset consisted of 5000 samples, covering combinations of 5 turbidity levels, 3 temperatures, and 3 salinities.
[0044] The feature extraction layer performs dimensionality reduction on the multimodal data after superimposed compensation. Principal component analysis (PCA) is used to project the high-dimensional acoustic, optical, and point cloud data into a low-dimensional subspace, retaining the principal components with the highest variance contribution rate to form the dimensionality-reduced perceptual feature set. The dimensionality reduction process preserves key information such as the mesh cage's geometric contour, mesh texture direction, and material reflectivity, while suppressing noise.
[0045] For acoustic data, the 256 distance sampling points per scan line were reduced to 16 dimensions; for optical image data, the original 1920x1080 image was first scaled to 224x224 pixels, then flattened into a 50176-dimensional vector, reducing it to 128 dimensions; for point cloud data, the 2880-dimensional vector of XYZ coordinates and intensity values of 720 points per frame was reduced to 32 dimensions. The feature vectors after dimensionality reduction for the three modalities were concatenated to form a 176-dimensional perceptual feature set.
[0046] The weights used by the feature extraction layer when fusing different modalities of data are also determined through supervised learning in a simulated environment. During training, multimodal dimensionality reduction features are used as input, and the degree of cage fouling is labeled as a supervision signal. The linear weighting coefficients of each modality feature are determined through cross-validation. The physical meaning of the weight coefficients is that the contribution of each modality of data to fouling identification varies under different turbidity environments. Optical images have the highest weight in clear water environments, while acoustic data have a higher weight in turbid environments. This dynamic weighting mechanism enables multimodal fusion to adapt to different water quality environments and maintain stable recognition performance.
[0047] The decision layer is loaded with a lightweight compressed deep learning model. Based on the ResNet-18 network structure, this model employs a two-stage compression strategy of channel pruning and weight quantization to compress the network, achieving a parameter compression rate of no less than 60% while maintaining recognition accuracy loss of no more than 3%. This specific two-stage compression scheme is optimized for the computational constraints of embedded industrial control computers, balancing model accuracy and inference speed.
[0048] The specific method for channel pruning is as follows: Calculate the L1 norm of each channel in each convolutional layer of ResNet-18, sort them by norm from smallest to largest, and prune 30% of the channels with the smallest L1 norm. After pruning, fine-tune the network training to restore accuracy. A pruning ratio of 30% is the optimal value determined through multiple experiments—below 30%, the compression effect is not significant, and above 30%, the accuracy drops too quickly. The specific method for weight quantization is as follows: Quantize the 32-bit floating-point weight parameters into 8-bit integer weight parameters. The quantization uses a symmetric linear quantization scheme, and the quantization step size is determined by dividing the maximum absolute value of the weights in each layer by 127. After quantization, the model's storage space is reduced to 1 / 4 of the original, the inference speed is improved by about 2 times, and the accuracy loss is controllable.
[0049] After channel pruning and weight quantization, the network model's parameter compression rate reaches over 60%. The original model size is approximately 45 megabytes, while the compressed size is approximately 17 megabytes. This reduction in model storage space and inference computation allows it to run directly on embedded processors. An inference frame rate of at least 10 frames per second (fps) is achieved. This is because channel pruning reduces the number of convolutional layer channels by approximately 30%, lowering the multiply-accumulate operation cost per forward inference by approximately 55%. Weight quantization converts 32-bit floating-point operations to 8-bit integer operations. Utilizing the embedded processor's NEONSIMD instruction set and hardware acceleration unit, the single-frame inference time is reduced from approximately 250 milliseconds in the original model to approximately 80 milliseconds. Measured against a target inference frame rate of 10 fps (i.e., 100 milliseconds per frame), the compressed model's inference speed meets real-time requirements, with a measured inference frame rate reaching 12 fps. This ensures that the AI decision-making module can process multimodal perception data and generate operational decision instructions in a timely manner. The decision layer obtains the dimensionality-reduced perception feature set as input and outputs the identification results of contaminated areas and damaged areas of the net cage through forward inference. The identification results include the type of fouling organism, the degree of coverage, and the location and size of the puncture. For punctures larger than 5 cm by 5 cm, the identification rate is no less than 92%. Fouling types can be classified into common attached organisms such as algae and shellfish. The degree of coverage is divided and quantified using a 0.1 m by 0.1 m grid, with each grid outputting a coverage value between 0 and 1.
[0050] Example 2: In one inspection, the perceptual feature set output by the feature extraction layer is a 176-dimensional vector. The decision layer inputs this vector into a compressed ResNet-18 network. After 18 layers of convolution and fully connected operations, the network generates a 20x20 heatmap at the output layer. Each heatmap unit corresponds to a 0.1m x 0.1m grid area on the mesh. The value in the 5th row, 8th column of the heatmap is 0.92, indicating a 92% probability of algal fouling in this grid area; the value in the 12th row, 15th column is 0.85, indicating an 85% probability of shellfish fouling in this grid area; the value in the 3rd row, 4th column is 0.98, and the geometric features of this grid area match the damage feature template by 0.95, thus identifying it as mesh damage. The damage size, calculated by fitting the point cloud data, is 8cm x 6cm. The inference time for this frame is 80ms, meeting the requirements for online processing. Compared to the uncompressed original ResNet-18 model, the compressed model improves inference speed by 2.3 times, reduces memory usage by 62%, and only decreases recognition accuracy by 2.1%, fully meeting the real-time requirements of embedded devices.
[0051] The behavior planning layer employs fuzzy C-means clustering to cluster discrete soiled and damaged points into several regions based on their spatial location. The membership index of the fuzzy C-means clustering algorithm is set to 2, the termination tolerance is set to 0.001, and the number of regions is adaptively determined by the algorithm based on the spatial distribution density of the soiled points, ranging from 2 to 8 regions. After clustering, a shortest cleaning path covering all soiled points is planned for each region. Path planning uses a grid-based wavefront propagation algorithm to generate a serpentine traversal path while avoiding damaged regions.
[0052] The behavior planning layer determines the cleaning parameters for each area based on the type and degree of fouling, including the initial settings for the cleaning disc rotation speed and downward pressure. The cleaning parameters are quantified using an index representing the degree of fouling in the net cage. Based on this, the index is output by the decision-making level after a weighted comprehensive assessment of fouling coverage and type. For algal fouling, the weighting coefficient is 0.6; for shellfish fouling, the weighting coefficient is 1.0. Coverage rate is the percentage of fouled grid cells in the area, with a value range of 0 to 1. Quantitative Index of Cage Fouling Degree The value is the product of coverage and type weight coefficient, with an upper limit of 1. This quantification mechanism allows cleaning parameters to precisely match the severity of soiling, avoiding a one-size-fits-all cleaning strategy. It ensures cleaning effectiveness while preventing damage to the mesh caused by over-cleaning.
[0053] Example 3: The decision layer identified three fouled areas and one damaged area in the bottom mesh. Area 1 contained 50 fouled meshes, of which 35 were algal fouled and 15 were shellfish fouled. The weight of algae was 0.6, the weight of shellfish was 1.0, and the coverage was 0.7. The weighted calculation showed that this area... The behavior planning layer plans a serpentine path for area one, with a total length of 8 meters, based on... The initial rotation speed of the cleaning disc was set at 100 revolutions per minute, and the downward pressure at 300 Newtons. Area two consisted of pure algae fouling with a coverage rate of 0.4. This corresponds to a washing disc rotation speed of 80 revolutions per minute and a downward pressure of 150 Newtons. Area three is purely shellfish fouling, with a coverage rate of 0.8. This corresponds to a washing disc rotation speed of 150 revolutions per minute and a downward pressure of 500 Newtons. This differentiated washing parameter setting allows for gentle washing of algae areas to protect the mesh, while powerful washing of shellfish areas ensures effective removal, achieving an optimal balance between cleaning effectiveness and mesh protection.
[0054] The actuator module includes a cleaning mechanism and a gripping mechanism. The cleaning disc of the cleaning mechanism is made of nylon, with a diameter of 1 to 2 meters and a thickness of 20 millimeters. Its surface is covered with anti-slip textures, consisting of a diamond-shaped cross pattern 3 millimeters deep and spaced 10 millimeters apart. The center of the back of the cleaning disc is connected to the output shaft of the cleaning drive unit via a flange. The cleaning drive unit includes a cleaning drive motor and a gear transmission mechanism. The cleaning drive motor is a permanent magnet synchronous motor with a rated power of 500 watts and a rated speed of 1500 rpm. It drives the cleaning disc to rotate through a planetary gear reducer with a reduction ratio of 10, and the cleaning disc speed can be adjusted from 50 to 200 rpm. The cleaning drive unit is mounted on a linear slide, which is driven by an electric push rod. This linear slide can move the cleaning disc vertically with a travel distance of 100 millimeters. The downward pressure applied by the cleaning disc to the mesh is indirectly controlled by controlling the thrust of the electric push rod, and the downward pressure can be adjusted from 50 to 800 Newtons.
[0055] The gripping mechanism features a telescopic handle, constructed from an inner and outer tube, both made of stainless steel. The outer tube has an outer diameter of 20 mm, and the inner tube has an outer diameter of 16 mm. The telescopic range is 50 mm to 300 mm. The front end of the gripping handle has a two-part gripper with a serrated gripping surface on the inner side. The maximum opening of the gripper is 50 mm. The gripping drive unit includes a gripper drive motor and a telescopic drive motor. The gripper drive motor is a stepper motor that drives the gripper to open and close via a worm gear mechanism, with an adjustable gripping force of 20 to 100 Newtons. The telescopic drive motor is a DC geared motor that drives the inner tube to extend and retract via a rack and pinion mechanism. The execution module receives work decision commands generated by the AI decision module, adjusting the rotation speed and downward pressure of the cleaning disc and the gripping posture of the gripping mechanism to perform cleaning and maintenance operations.
[0056] During the operation, the behavior planning layer sends the planned cleaning path and cleaning parameters to the anti-flow adaptive controller. The anti-flow adaptive controller calculates the thruster commands acting on the six-degree-of-freedom motion control module based on the flow field data and feeds back the robot's actual motion state to the environment perception layer of the AI decision module. The AI decision module and the anti-flow adaptive controller perform collaborative calculations based on a series of mathematical models to achieve quantitative adjustment of computing resources, motion energy consumption, and cleaning force output. This dynamic collaborative mechanism of perception, computing power, and cleaning force is one of the core innovations of this invention—it links perception quality, computing resources, and operation effect through a quantitative model, achieving optimal allocation of system resources.
[0057] The system uses the effective coefficient of the perceived features Quantify the reliability of multi-sensor fusion data under current water turbidity conditions. Effective coefficient of sensing features. Solve using the following formula: ; In the formula, Sonar sensor data delay refers to the time it takes for a sound wave to travel from transmission and reception to the data packet transmission to the fusion module; a typical value is 0.05 seconds. The attenuation coefficient of sound waves in turbid water is used to characterize the sound energy loss per unit distance. It is calibrated based on the concentration and size distribution of suspended particles in seawater, with a typical value of 0.2. This represents the data latency of the camera sensor, including exposure, image compression, and transmission delays, typically 0.08 seconds. The turbidity scattering coefficient reflects the ability of water to scatter visible light and is positively correlated with the turbidity NTU value, with a typical value of 0.3. This represents the data latency of the lidar sensor, encompassing the laser pulse flight time and point cloud preprocessing latency, with a typical value of 0.02 seconds. This is the scattering coefficient of suspended particles in water, describing the intensity of laser scattering by suspended particles; a typical value is 0.25. The sampling and processing cycle for sensor fusion is set by the system and is usually 0.1 seconds. This is the standard detection range for sonar sensors, reaching up to 50 meters in clean water environments. This is the actual detection distance under the current turbidity level. Due to the influence of turbidity, it typically decreases to 20 meters.
[0058] The derivation of this formula is based on the following logic: multiplying the data delay of each sensor by the corresponding environmental attenuation coefficient, and then dividing by a unified fusion processing cycle, represents the proportion of information loss caused by delay and attenuation in each channel. Adding 1 to the denominator to the sum of the loss proportions of all channels, and taking the reciprocal, results in... Between 0 and 1, when the delay and attenuation of each channel are both 0. When the time delay and attenuation approach 1, the perceived quality is optimal; as time delay and attenuation increase... The square root factor reflects the additional reduction in the sensing signal-to-noise ratio caused by range compression. The overall effectiveness of multimodal perception in turbid water bodies was comprehensively quantified, providing a basis for allocating inference depth for AI decision-making modules.
[0059] Example 4: In a certain operation, the turbidity of the water was 2 NTU, and the time delay and attenuation coefficient of each sensor were as follows: Second, , Second, , Second, Integration cycle Seconds, standard sonar detection range meters, current actual detection distance Meters. First, calculate the sonar channel loss ratio: Camera channel loss ratio: LiDAR channel loss ratio: The sum of the proportions of the three losses is The denominator is Taking the reciprocal gives The distance attenuation factor is Multiply the two together, This value indicates that, under the current turbidity, the overall effectiveness of multimodal perception is approximately 45.5% of that in a clean water environment. Based on this, the AI decision-making module adjusts the confidence threshold and feature weight allocation during inference—increasing the weight of acoustic data and decreasing the weight of optical data, while raising the recognition confidence threshold from 0.8 to 0.9 to reduce the false detection rate.
[0060] Comprehensive resistance during system operation Solve using the following formula: ; In the formula, The density of seawater is taken as 1025 kg per cubic meter. The velocity of the system relative to the water body is calculated by a Doppler velocimeter or inertial navigation system. The effective frontal area of the system in the direction of water flow is calculated based on the robot's projected area and drag coefficient, with a typical value of 0.5 square meters. The local axial static strain of the mesh is defined as the deformation of the mesh after being subjected to force. With the original length The ratio, that is Its value is affected by the pressure under the washing disc, the tension of the mesh, and the material properties, with typical values ranging from 0.01 to 0.05. For nylon mesh, the elastic modulus of the mesh material is... It is approximately 2.5 gigapascals. For the local deformation of the mesh. This is the original length of the locally deformed segment of the mesh. The basic resistance of the system in still water is obtained by experimental towing measurement, with a typical value of 50 Newtons. , , The operating condition correction coefficients were determined through flow field simulation and fitting of sea trial data, and were set to 1.2, 0.8, and 1.0, respectively.
[0061] The theoretical basis of this formula is the superposition principle of fluid dynamics and elasticity. The first term is the fluid dynamic pressure resistance, which is related to the dynamic pressure. Proportional This integrates the drag coefficient and the half-factor. The second term is the elastic drag caused by the deformation of the mesh. According to Hooke's Law, stress is proportional to strain, and the elastic force equals the elastic modulus multiplied by the strain and then multiplied by the area of the force application. multiplied by The deformation drag is obtained after correction. The third term is the basic drag correction. The sum of the three terms constitutes the total drag that the robot must overcome during the net-laying operation, providing a basis for the thruster thrust budget and cleaning pressure setting.
[0062] Example 5: Seawater density during bottom net cleaning operations kilograms per cubic meter, robot relative speed to water flow meters per second, effective frontal area Square meters. The mesh is made of nylon with a modulus of elasticity of [missing information]. Pa, original length of the partially deformed segment of the mesh The rice grains deform under the pressure of the washing pan. Meters, then static strain Static water foundation resistance Newton. Operating condition correction factor. , , Calculate the first term, hydrodynamic drag: Newton. Calculate the second term, elastic resistance: After simplification, it becomes That is, 900 Newtons. Calculate the third term, the fundamental resistance correction: Newton. Combined resistance. Newton. This result indicates that under this operating condition, the deformation elastic drag of the mesh accounts for 81.5% of the total drag, which is the main load the thruster needs to overcome. Therefore, the downforce in the cleaning parameters needs to be appropriately reduced to decrease deformation drag. This finding directly guides the optimization strategy for cleaning force—by moderately reducing the downforce, the total drag can be significantly reduced, thereby saving thruster energy consumption and extending operation time.
[0063] Edge inference computing power ratio Solve using the following formula: ; In the formula, The system's computing power is assigned a weighting coefficient, which is set by the user according to the task priority, with a typical value of 0.8. The effective coefficients of the aforementioned perceptual features. This constitutes comprehensive resistance. The maximum design cruising speed of the system is set to 0.5 meters per second. The total rated drive power of the six-degree-of-freedom motion control module is given by the sum of the rated power of all thrusters, typically 1200 watts. The idle operating power consumption of the embedded AI decision module is typically 20 watts. This is the baseline computing power consumption for the embedded AI decision-making module, typically 60 watts.
[0064] The derivation of this formula follows this approach: the total power budget of the system is divided into two parts: motion consumption and computational consumption. Motion consumption power equals the product of resistance and velocity. Its total rated power The ratio represents the proportion of power used in motion. This represents the percentage of remaining power available for the AI module. This remaining percentage is multiplied by the effective coefficient of the perceived features. and weighting coefficients This reflects that the better the perceived quality, the more worthwhile it is to invest computing power. (Logarithmic terms) This characterizes the computational power elasticity that the AI module can unleash under the condition of idle power consumption relative to baseline power consumption; the higher the ratio, the higher the upper limit of available computational power. Ultimately... The output is a dimensionless coefficient between 0 and 1, which guides the AI decision-making module to dynamically adjust the model accuracy and batch size during inference, thereby maximizing inference capability under power constraints.
[0065] Example 6: Based on the data from the previous two examples, the overall resistance... Newton, effective coefficient of perceptual features Maximum design cruising speed meters per second, total rated drive power Watts, no-load power consumption Watts, reference power consumption Watts, weighting coefficient Calculate the power consumed during exercise: Watts. Power consumption ratio for motion: Remaining power ratio: Calculate the logarithmic term: , Multiply all terms together: The results indicate that under the current high-drag, low-perceived-quality operating conditions, only about 5.6% of the system power budget is available for AI inference. The AI decision-making module needs to automatically switch to a lightweight inference mode—reducing the input image resolution from 224×224 to 160×160, disabling some convolutional layers, and lowering model accuracy to adapt to this computational constraint. However, under clear water, low-drag operating conditions… A score of 0.3 or higher can be achieved, at which point the AI module can run a full-precision model and obtain a higher recognition accuracy. This dynamic computing power allocation mechanism enables the system to automatically balance perception accuracy and motion capability under different operating conditions, thereby optimizing overall performance.
[0066] Optimal cleaning operation force Solve using the following formula: ; In the formula, The cleaning force distribution coefficient is set to 0.9 based on the type of cleaning disc. This represents the percentage of edge inference computing power. The index is used to quantify the degree of fouling of fish cages. It is determined by the decision-makers based on the type of fouling and the coverage area. The value range is from 0 to 1, where 0 indicates no fouling and 1 indicates full coverage and stubborn fouling. To determine the static pressure of the cleaning medium, the seawater pressure at the depth of the water is taken. The effective contact area between the cleaning disc and the mesh is determined by multiplying the square of the cleaning disc radius by the contact angle coefficient, with a typical value of 0.3 square meters. The hydrodynamic influence coefficient is obtained through CFD simulation, with a typical value of 1.5. The angle of attack coefficient of the cleaning disc on the mesh, when the cleaning disc is perpendicular to the mesh. It is 1, and less than 1 when tilted. The dynamic pressure of the water flow on the system, by calculate. The area of the cleaning disc facing the water flow direction depends on the diameter of the cleaning disc and the projection of the flow velocity direction. The maximum pressure force that the system can provide with its maximum thrust is obtained by subtracting its own drag from the thrust of the thruster at full power, with a typical value of 800 Newtons. This constitutes comprehensive resistance.
[0067] The theoretical basis of this formula is static equilibrium. The first term represents the active cleaning force component, which is related to the degree of soiling. and the percentage of available computing power Positive correlation: The more abundant the computing power and the more severe the contamination, the greater the applied static pressure. The reference static pressure is given, multiplied by After adjusting the quantitative index, it is transformed into targeted cleaning force. The second item is the water flow disturbance compensation component, dynamic pressure. Multiply by the frontal area The impact force of the water flow is obtained, and then through the coefficient and The corrected value is the effect of water flow on the stability of the cleaning contact force, multiplied by the ratio of the limiting force to the overall resistance. Normalization scaling is performed to ensure that the compensation amount does not exceed the system capacity. The algebraic sum of the two terms represents the optimal cleaning force that the actuator module should output. The AI decision-making module is based on The system generates commands to adjust the overall downward pressure and rotation speed of the cleaning disc via a proportional pressure reducing valve and a servo motor, thereby applying force as needed.
[0068] Example 7: Continuing with the data from the previous examples, the system operates at a water depth of 40 meters, and the static pressure of the cleaning medium is... Pa. Effective working area of the cleaning disc. Square meters. Quantitative index of soiling level. Edge inference computing power ratio Cleaning force distribution coefficient Calculate the first active cleaning force component: Calculate step by step: , , , Newton. Dynamic pressure of water. Pa. Washing disc frontal area square meters, hydrodynamic influence coefficient Angle of attack coefficient Extreme pressure on the net Newton, total resistance Newton, liberty Calculate the second component of the flow compensation: Calculate step by step: , , Newton. Optimal cleaning force. Newtons. This result indicates that under high-resistance conditions at a water depth of 40 meters, the cleaning disc needs to apply approximately 3139 Newtons of downward pressure to effectively remove fouling and resist water flow disturbance. The actuator module adjusts the electric actuator thrust to 3139 Newtons accordingly, and simultaneously... The cleaning disc rotation speed is set at 120 revolutions per minute. It's worth noting that the first item, "Active Cleaning Power," includes the percentage of computing power. The impact—when When the temperature is low, the system will appropriately reduce the cleaning force to save energy; when When the pressure is high, the system will increase the cleaning force to enhance the cleaning effect. This linkage mechanism between computing power and cleaning force enables the system to adapt to different operating conditions and achieve the optimal balance between energy efficiency and effectiveness.
[0069] The overall workflow is illustrated using a complete bottom net cleaning operation as an example. Before the operation begins, the operator installs the counterweight on the lower mounting plate and adjusts the pitch adjustment mechanism to keep the robot horizontal. The robot is then lowered into the water, and the 4G wireless communication unit of the communication module establishes a connection with the shore-based control center, transmitting robot status data back. The operator then issues inspection and cleaning task instructions through the shore-based control center.
[0070] The six-degree-of-freedom motion control module is activated, and the forward / backward drive system and yaw drive system propel the robot to the starting position at the bottom of the cage. The GPS / BeiDou dual-mode positioning unit provides the initial position coordinates, and the inertial navigation unit continuously outputs the robot's attitude angles. Once the robot reaches the predetermined starting position, the lateral drive system stops, and the vertical propeller of the yaw drive system reverses to provide downforce, keeping the cleaning disc in close contact with the surface of the bottom netting.
[0071] The multi-sensor fusion module of the AI decision-making module begins collecting environmental data. Each sensor achieves time synchronization via hardware trigger signals, with a synchronization accuracy better than 1 millisecond. The sonar sensor emits ultrasonic waves and receives the echoes, generating acoustic data; the camera sensor captures visible light images of the netting, which are pre-processed by a signal processor to form optical image data; the lidar sensor rotates and scans at a frequency of 10 Hz, generating point cloud data; and the water flow sensor continuously measures flow velocity and direction, generating flow field data. All four data streams are converged to the AI decision-making module via a gigabit Ethernet switch.
[0072] After receiving four data streams, the environmental perception layer performs distance and attenuation compensation on the acoustic data, performs dark channel prior dehazing on the optical image data, filters out outliers from the point cloud data, and smooths the flow field data using Kalman filtering. During the compensation process, the environmental perception layer uses the real-time calculated effective coefficients of the sensing features. Dynamically adjust the compensation intensity for each mode. The lower the value, the greater the compensation intensity, in order to restore the perceived quality as much as possible. The compensated data is fed into the feature extraction layer, and the dimensionality is reduced by principal component analysis. The acoustic data is reduced to 16 dimensions, the optical image data to 128 dimensions, and the point cloud data to 32 dimensions, which are then stitched together to form a 176-dimensional perceptual feature set.
[0073] The compressed ResNet-18 network loaded in the decision layer acquires the sensory feature set. After forward inference, it identifies three fouled areas and one damaged area on the output heatmap. Area 1 is a mixed fouling of algae and shellfish, area 2 is purely algae fouling, and area 3 is purely shellfish fouling. The damaged area is located in the northeast corner of the bottom net, and the size of the damaged opening, fitted by the point cloud, is 10 cm by 8 cm, which exceeds the threshold of 5 cm by 5 cm, requiring repair.
[0074] The behavior planning layer uses a fuzzy C-means clustering algorithm to group contaminated points into three regions based on their spatial location, and plans a serpentine full-coverage cleaning path for each region. The contamination level in region one is quantified by an index. The cleaning disc rotation speed was set at 100 revolutions per minute, and the downward pressure at 300 Newtons; the algae coverage rate in area two was 0.4. The washing disc rotation speed was set at 80 revolutions per minute, and the downward pressure at 150 Newtons; the shellfish coverage rate in area three was 0.8. The cleaning disc rotation speed was set at 150 revolutions per minute and the downward pressure at 500 Newtons.
[0075] The behavior planning layer sends the cleanup path and cleanup parameters to the flow-resistant adaptive controller. Simultaneously, the AI decision-making module, based on the current water flow and resistance conditions, collaboratively solves for the effective coefficients of the sensing features using a mathematical model. Comprehensive resistance Edge inference computing power ratio and optimal cleaning power The solution will be found The algorithm compares the cleanup parameters with those preset by the behavior planning layer, and takes the larger value as the final execution value to ensure the cleanup effect. This decision-making mechanism, which combines behavior planning and quantitative optimization, ensures both the rationality of the path planning and the optimality of the cleanup parameters.
[0076] The actuator module receives the final operation decision command. The cleaning drive motor drives the cleaning disc to rotate at the commanded speed, and the electric push rod presses the cleaning disc against the netting with the commanded thrust. The robot advances in a serpentine pattern from south to north along the planned path to clean area one. During the cleaning process, the water flow sensor continuously monitors changes in the flow field. When the water flow velocity suddenly increases from 1.5 m / s to 2.0 m / s, the anti-flow adaptive controller detects the change in flow field data, updates the feedforward compensation in real time, and increases the thruster thrust to maintain the robot's netting-attached posture. At the same time, the AI decision module recalculates the overall resistance. and optimal cleaning power The updated downward pressure command is sent to the actuator module, and the electric push rod adjusts the thrust in real time to prevent the robot from being washed away from the surface of the net by the water flow.
[0077] After cleaning Zone 1 is completed, the robot moves to Zone 2 and performs cleaning according to the updated cleaning parameters. After cleaning Zone 2 is completed, the robot moves to Zone 3 to perform cleaning. After all three zones are cleaned, the robot moves to the damaged area. The gripping handle of the gripping mechanism extends to the edge of the damaged area. The telescopic drive motor adjusts the inner tube to a suitable extension length according to the distance between the damaged area and the robot. The gripper drive motor drives the grippers to close, clamping the netting on one side of the damaged area. After the gripping is stable, the gripping handle slowly retracts, pulling the netting on both sides of the damaged area closer together, providing an operational basis for subsequent repair work.
[0078] For side net cleaning, the operator lifts the robot out of the water, moves the counterweight from the lower mounting plate to its mounting position on the upper mounting plate, and simultaneously adjusts the pitch adjustment mechanism to make the robot stand upright. After being lowered back into the water, the robot is attached to the side net surface in a vertical posture, with the cleaning disc's working surface facing the side net. At this time, the thrust distribution matrix of the six-degree-of-freedom motion control module automatically switches to side net mode, and the feedforward gain matrix of the anti-current adaptive controller also switches synchronously. The subsequent perception, decision-making, and control processes are the same as for bottom net cleaning; the forward and backward drive system drives the cleaning disc to move up and down along the side net, and the lateral drive system drives the robot to translate laterally.
[0079] The communication module employs a dual-mode communication approach combining 4G wireless communication and CAN bus wired communication. The 4G wireless communication unit uses a Quectel EC20 module, supporting the LTE Cat4 standard with a downlink rate of 150 Mbps and an uplink rate of 50 Mbps. It connects to the embedded computer of the AI decision module via a USB interface. The CAN bus wired communication unit uses an MCP2515 controller and a TJA1050 transceiver, with a communication rate of 1 Mbps. It connects to the AI decision module via an SPI interface. All sensors and controllers are mounted on the CAN bus, with pre-assigned node addresses: sonar sensor 0x01, camera sensor 0x02, lidar sensor 0x03, water flow sensor 0x04, anti-current adaptive controllers 0x10 to 0x15 corresponding to the local controllers of the six thrusters, and the actuator module 0x20.
[0080] The communication module integrates a GPS / BeiDou dual-mode positioning unit and an inertial navigation unit. The GPS / BeiDou dual-mode positioning unit uses the UBLOX-M8N module, simultaneously receiving GPS and BeiDou satellite signals, with a horizontal positioning accuracy of ±1 meter and an update frequency of 10 Hz. The inertial navigation unit uses the MPU9250 nine-axis sensor, integrating a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer, with an attitude angle measurement accuracy of ±0.5 degrees and an update frequency of 100 Hz.
[0081] Near the water surface, the 4G module uploads the operational status, contamination identification results, and damage location coordinates to the shore-based control center. During underwater operations, the CAN bus handles data exchange between modules, and the integrated navigation system provides the robot with position and attitude information, ensuring closed-loop control of the inspection path. The integrated navigation system uses GPS / BeiDou positioning data as the absolute position reference and inertial navigation data for relative position calculation. By fusing the two data sources through extended Kalman filtering, positioning accuracy reaches ±1 meter, and attitude accuracy reaches ±0.5 degrees. In underwater environments where GPS signals are lost, inertial navigation can independently maintain a position drift of no more than 5 meters within 10 minutes. Simultaneously, Doppler velocity data from the water flow sensor also participates in navigation fusion, further suppressing position drift and improving the accuracy and reliability of underwater navigation.
[0082] The forward / reverse drive system, the yaw angle drive system, and the lateral drive system all include a permanent magnet synchronous motor and a planetary gear reducer. The anti-current adaptive controller is embedded in the local controller corresponding to each of the forward / reverse drive system, the yaw angle drive system, and the lateral drive system. The anti-current adaptive controller stores an adaptive PID parameter adjustment algorithm. The initial parameters of the adaptive PID algorithm are set as follows: proportional coefficient... Integral coefficient Differential coefficients During algorithm execution, the flow velocity is dynamically adjusted based on the rate of change of the flow field data. For every 0.1 m / s² increase in the rate of change, Increase by 5; dynamically adjust based on the integral value of the robot's position deviation. For every 0.1 m / s increase in the integral value, Decrease by 1; dynamically adjust based on the differential value of the position deviation. For every 0.05 meters per second increase in the differential value, Add 2. This adjustment rule is implemented through a fuzzy control table, eliminating the need for online solving of complex optimization problems and meeting millisecond-level real-time requirements.
[0083] The cleaning tray is made of nylon with an anti-slip textured surface. The gripping mechanism has a retractable handle, and its gripping range covers 50 mm to 300 mm. The nylon material is PA6, with a tensile strength of 80 MPa, a flexural strength of 120 MPa, and a water absorption rate of 1.2%. It maintains dimensional stability and abrasion resistance even after long-term underwater immersion. The anti-slip texture is 3 mm deep, with a 60-degree angle and a 10 mm spacing. Both the inner and outer tubes of the gripping handle are made of 316L stainless steel, and a PTFE sliding bushing is provided between the outer wall of the inner tube and the inner wall of the outer tube, ensuring that the telescopic friction does not exceed 5 Newtons. A 3 mm thick silicone rubber pad is attached to the gripping surface of the claws to increase the coefficient of friction and prevent damage to the net rope.
[0084] The communication module employs a dual-mode communication approach combining 4G wireless communication and CAN bus wired communication, integrating a GPS / BeiDou dual-mode positioning unit and an inertial navigation unit. When the robot is on the surface or in shallow water, the communication module prioritizes 4G wireless communication to upload operational status, contamination identification results, damage location coordinates, and environmental data to the shore-based control center. When the robot dives to a depth exceeding 5 meters, the 4G signal weakens, and the communication module automatically switches to CAN bus mode. Sensors and controllers exchange data internally via the CAN bus. The AI decision-making module temporarily stores the data to be transmitted back to the local solid-state drive, and then uploads it in batches via the 4G module after the robot surfaces to shallow water. The integrated navigation system uses GPS / BeiDou positioning data as the absolute position reference and inertial navigation data for relative position calculation. By fusing the two data sources through extended Kalman filtering, the positioning accuracy reaches ±1 meter, and the attitude accuracy reaches ±0.5 degrees. In underwater environments where GPS signals are lost, the inertial navigation system can independently maintain a position drift of no more than 5 meters within 10 minutes.
[0085] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A multimodal deep-sea cage autonomous inspection and cleaning system, comprising a main frame, a multi-sensor fusion module, an AI decision-making module, an actuator module, and a communication module, characterized in that: The system also includes a six-degree-of-freedom motion control module, which comprises a forward / backward drive system, a lateral drive system, a heel drive system, a roll angle drive system, a pitch angle drive system, and a yaw angle drive system. The six-degree-of-freedom motion control module also includes a current-resistant adaptive controller. The multi-sensor fusion module includes a sonar sensor, a camera sensor, a lidar sensor, and a water flow sensor, which are used to output acoustic data, optical image data, point cloud data, and flow field data respectively. The output of the water flow sensor is connected to the AI decision module and the anti-flow adaptive controller, respectively. The flow field data serves as the environmental perception input of the AI decision module and the real-time control input of the anti-flow adaptive controller. The AI decision-making module is used to acquire the acoustic data, the optical image data, the point cloud data, and the flow field data, and to perform nonlinear compensation and feature fusion on the four types of data. Based on the fused perception features, it identifies the contaminated area of the net cage and the damaged area of the net, and generates an operation decision instruction containing the cleaning path and cleaning parameters. The actuator module is used to receive the operation decision instruction and adjust the working parameters of the cleaning disc and / or the gripping mechanism; The anti-flow adaptive controller is used to acquire the flow field data and calculate the thruster commands based on the flow field data to drive the various drive systems of the six-degree-of-freedom motion control module. The anti-flow adaptive controller is also used to feed back the actual motion state to the AI decision module, and the AI decision module updates the environmental perception parameters and operation decision instructions according to the actual motion state.
2. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 1, characterized in that, The AI decision-making module includes an environmental perception layer, a feature extraction layer, a decision-making layer, and a behavior planning layer; The environmental perception layer is used to receive the acoustic data, the optical image data, the point cloud data, and the flow field data, and to model the nonlinear noise caused by environmental interference, modeling errors, and sensor noise through a nonlinear mapping function, calculate the compensation amount to cancel the nonlinear noise, and superimpose the compensation amount into the original control signal. The feature extraction layer is used to perform dimensionality reduction processing on the acoustic data, optical image data, point cloud data and flow field data after superimposing the compensation amount using the principal component analysis method, and outputs the dimensionality-reduced perception feature set; The decision layer integrates the lightweight compressed deep learning model. The decision layer is used to obtain the dimensionality-reduced perceptual feature set and identify the soiled areas of the net cage and the damaged areas of the netting. The behavior planning layer is used to plan the cleaning path based on the distribution of the soiled areas of the net cage and the damaged areas of the netting using a fuzzy C-means clustering algorithm, and to determine the cleaning parameters based on the type and degree of soiling of the net cage.
3. The deep-sea cage autonomous inspection and cleaning system according to claim 2, characterized in that, The forward / backward drive system, the lateral drive system, the heave drive system, the roll angle drive system, the pitch angle drive system, and the yaw angle drive system are each composed of a permanent magnet synchronous motor and a planetary gear reducer modular propeller. Each modular propeller is installed on the main frame through a standard electrical interface and a mechanical quick-connect structure.
4. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 2, characterized in that, The lightweight compressed deep learning model is a network model obtained by channel pruning and weight quantization based on the ResNet-18 network structure. The lightweight deep learning model is deployed in the underwater embedded edge computing unit built into the AI decision module. The parameter compression rate of the network model is not less than 60%, and the inference frame rate of the network model after channel pruning and weight quantization in the AI decision module is not less than 10 frames / second.
5. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 2, characterized in that, The nonlinear mapping relationship on which the environmental perception layer calculates the compensation amount, and the weights on which the feature extraction layer fuses the acoustic data, the optical image data, the point cloud data, and the flow field data, are all obtained through supervised learning training on multiple sets of sensor noise and cage status training data in a water environment with specific turbidity, temperature, and salinity parameter ranges. The preset limited interval is stored in the built-in storage unit of the AI decision module.
6. The deep-sea cage autonomous inspection and cleaning system according to claim 2, characterized in that, The feature extraction layer uses principal component analysis to perform dimensionality reduction and fusion on the acoustic data, optical image data, point cloud data, and flow field data. The number of principal components to be retained after dimensionality reduction is dynamically determined based on the water turbidity, temperature, and salinity parameters collected by the water flow sensor. The dynamic determination rule for the number of principal components to be retained is as follows: when the water turbidity is higher than a first threshold, the number of retained principal components is increased by a first increment to maintain the effective feature dimension of the optical image data.
7. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 2, characterized in that, The behavior planning layer sends the planned cleanup path and cleanup parameters to the anti-flow adaptive controller. The anti-flow adaptive controller calculates the thruster command acting on the six-degree-of-freedom motion control module based on the flow field data and feeds back the actual motion state to the AI decision module. The AI decision module updates the nonlinear compensation model parameters of the environment perception layer based on the deviation between the actual motion state and the expected motion state, and re-outputs the updated compensation amount to the feature extraction layer.
8. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 7, characterized in that, The AI decision module is also used to determine the effective coefficient of the sensing features based on the time delay of the acoustic data, the attenuation coefficient of the sound wave in the water, the time delay of the optical image data, the turbidity scattering coefficient of the water, the time delay of the point cloud data, and the scattering coefficient of suspended particles in the water. It also calculates the comprehensive resistance of the system during operation based on the seawater density, the velocity of the system relative to the water, the effective upstream area of the system, the local static strain of the net, the elastic modulus of the net, the ratio of the local deformation of the net to the original length of the deformed section, and the static water foundation resistance.
9. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 8, characterized in that, The AI decision-making module is also used to determine the edge inference computing power ratio based on the effective coefficient of the perception features, the comprehensive resistance, the maximum design cruise speed of the system and the total rated drive power of the six-degree-of-freedom motion control module, and to determine the optimal cleaning operation force based on the edge inference computing power ratio, the quantification index of the degree of fouling of the cage, the static pressure of the cleaning medium and the effective working area of the cleaning disc. The actuator module is used to adjust the rotation speed and downward pressure of the cleaning disc according to the optimal cleaning operation force.
10. The multimodal deep-sea cage autonomous inspection and cleaning system according to claim 3, characterized in that, The forward / reverse drive system, the yaw angle drive system, and the heave drive system all include a permanent magnet synchronous motor and a planetary gear reducer. The anti-current adaptive controller is embedded in the local controller corresponding to each of the forward / reverse drive system, the yaw angle drive system, and the heave drive system. The anti-current adaptive controller stores an adaptive PID parameter adjustment algorithm.