Underwater robot control system for detection
By introducing spiking neural networks and multi-level denoising convolution techniques, optimizing sonar image processing, and combining with an extended PPO model, the problems of noise suppression and path planning for underwater robots in complex environments were solved, achieving efficient target detection and obstacle avoidance, and improving the autonomy and reliability of underwater robots.
Patent Information
- Application Number
- CN202511464131.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
How can existing underwater robot control systems achieve autonomous navigation, target recognition and obstacle avoidance, and path planning in complex environments? How can underwater robots achieve efficient target detection, obstacle avoidance, and path planning in complex underwater environments?
By introducing spiking neural networks, multi-level denoising networks, and multi-level denoising convolution techniques, the processing and noise suppression capabilities of underwater sonar images are optimized. By introducing spiking neural networks, multi-level denoising networks, and multi-level denoising suppression capabilities, the detection accuracy and robustness of underwater targets are optimized. Furthermore, by extending the PPO reinforcement learning model, the stability and convergence speed of obstacle avoidance path planning are enhanced.
It significantly improves the accuracy and robustness of underwater target detection, enhances the stability and convergence speed of obstacle avoidance path planning, and realizes the autonomy and reliability of efficient environmental monitoring, object detection, and seabed construction tasks.
Smart Images

Figure CN120972741A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, in particular to a control system of an underwater robot for exploration. BACKGROUND
[0002] With the deepening of underwater exploration tasks, underwater robots are increasingly widely used in underwater environment monitoring, object detection, ocean resource investigation, seabed construction and other fields. Underwater robots often need to deal with complex and uncertain environments, such as turbid water quality, complex underwater obstacles and extreme water depth conditions. Therefore, how to efficiently realize the autonomous navigation, target recognition and obstacle avoidance control of underwater robots has become an important issue in current underwater robot technology. Existing underwater robot control systems mostly rely on traditional sensor technologies such as sonar, depth sensors, inertial measurement units (IMU) to obtain underwater environment information and perform simple path planning. However, these systems have several problems. First, existing systems mostly rely on standard PPO models for obstacle avoidance path planning. Although this algorithm can better complete basic path planning tasks, in dynamic environments, especially in underwater environments with multiple unknown obstacles, the PPO model often shows slow convergence speed and low strategy stability, resulting in inaccurate path planning. Second, existing path planning algorithms have poor obstacle avoidance performance in dynamic environments and are difficult to cope with frequently changing obstacles and uncertainties in underwater environments. Finally, existing underwater robot control systems still have some gaps in real-time performance and accuracy, making it difficult to meet the requirements of high-precision navigation and complex tasks. SUMMARY
[0003] The present application provides a robot control system for underwater exploration, aiming to solve the problems of noise suppression, target detection accuracy, obstacle avoidance performance and path planning in complex underwater environments in existing systems. The system optimizes the processing and noise suppression ability of sonar images by introducing a spiking neural network (SNN) and a multi-stage denoising convolution technology, thereby significantly improving the detection accuracy and robustness of underwater targets. At the same time, an extended PPO reinforcement learning model is used to enhance the stability and convergence speed of obstacle avoidance path planning in dynamic environments by introducing a lower bound of strategy update, which can adjust the strategy in real time in rapidly changing underwater environments, ensuring the efficiency and accuracy of the system. Through these technological innovations, the present application realizes efficient target detection, three-dimensional obstacle positioning, real-time environment mapping and path planning in complex underwater environments, significantly improving the autonomy and reliability of underwater robots in performing tasks such as environment monitoring, object detection and seabed construction.
[0004] The present application provides a control system of an underwater robot for exploration, which comprises a carrier platform, a multi-sensor fusion module and a core controller.
[0005] A carrier platform for carrying a detection device, the detection device comprising an inertial measurement unit (IMU), a depth sensor, a sonar imaging unit and an underwater positioning unit;
[0006] A multi-sensor fusion module configured to collect robot pose, depth, sonar image and position data in real time through the detection device, obtain multi-source sensor data, and realize multi-source sensor data fusion through a Kalman filtering algorithm to output high-precision state estimation data;
[0007] A core controller connected to the multi-sensor fusion module and configured to:
[0008] The sonar image is processed by a pulse underwater YOLO model to extract image features of obstacles and targets, generate a target class probability map and a bounding box coordinate map, and combine the high-precision state estimation data to perform three-dimensional positioning of obstacles and generate spatial semantic information; the pulse underwater YOLO model is constructed based on a conventional YOLO model by introducing a pulse neural network and a multi-level denoising convolution kernel technology to optimize the feature extraction and noise suppression capability of the conventional YOLO model;
[0009] The spatial semantic information is combined to construct a real-time environment map using SLAM;
[0010] Based on the real-time environment map, an extended PPO model is used to generate robot control instructions to execute obstacle avoidance path planning; the extended PPO model is constructed by introducing a policy update lower bound to optimize the policy stability and convergence speed of the traditional PPO model.
[0011] Further, the generation process of the target class probability map and the bounding box coordinate map specifically includes the following steps:
[0012] Step S1: normalizing the pixel values of the sonar image to generate normalized image data;
[0013] Step S2: performing preliminary feature extraction on the normalized image data through a convolution layer and a batch normalization layer to generate a floating-point feature map, copying the floating-point feature map into a time step format and inputting it into an integrate-fire neuron (IF neuron) to generate a pulse feature map;
[0014] Step S3: presetting three groups of denoising convolution kernels, the three groups of denoising convolution kernels including a horizontal difference kernel (used to detect horizontal pulse interference), a vertical difference kernel (used to detect vertical pulse interference) and a Laplacian kernel (used to detect isolated noise points), performing a sliding window convolution operation on the pulse feature map using the three groups of denoising convolution kernels to obtain multi-channel convolution response results; introducing a threshold judgment mechanism, combining the multi-channel convolution response results, performing pulse point detection and pixel flipping processing, and outputting a noise-remedied pulse feature map;
[0015] Step S4: input the noise repair pulse feature map into a backbone network to generate a deep feature map, the backbone network is composed of pulse unit modules, the pulse unit modules are constructed by introducing a CSPNet residual structure, combining feature channel splitting, local convolution and cross-path activation preservation mechanism, for improving the deep neuron pulse firing rate, relieving the "pulse degradation" problem, enhancing the deep pulse firing rate, relieving the activation sparsity problem caused by the depletion of the integral-discharge type neuron membrane potential, and realizing stable extraction of multi-scale pulse features;
[0016] Step S5: maximum pooling and activation operation are performed on the deep feature map to generate a fusion feature map, and according to the fusion feature map, a target class probability map and a bounding box coordinate map are generated.
[0017] Further, the extended PPO model is used to generate robot control instructions to perform obstacle avoidance path planning and target tracking tasks, which specifically includes the following steps:
[0018] Step B1: collect historical trajectory data and combine real-time environment maps as the basis for strategy optimization of the extended PPO model, use generalized advantage estimation to obtain the original advantage value of the robot in the current state, calculate statistical indicators according to the original advantage value, and perform normalization processing to generate a modulated advantage value, the statistical indicators include L2 norm and standard deviation;
[0019] Step B2: use the modulated advantage value to construct a strategy loss function and iteratively update the strategy network parameters; introduce an extended strategy update lower bound to control the update amplitude of the strategy network and ensure the stability of the strategy optimization process;
[0020] Step B3: construct a training target according to the modulated advantage value to optimize and train the value function network, ensure that the training signals of the strategy network and the value function network are consistent, improve the generalization ability of the control strategy, and generate an optimized strategy distribution;
[0021] Step B4: sample actions according to the optimized strategy distribution, and map the corresponding thruster control parameters to specific thrust instructions, including thrust direction, amplitude and execution duration, to form robot control instructions;
[0022] Step B5: execute the robot control instructions to drive the thrusters to realize hovering, pitching, yawing and depth-keeping movements; perform obstacle avoidance path planning and target tracking tasks.
[0023] By using the above scheme, the application has the following beneficial effects:
[0024] The present application successfully solves the noise interference problem existing in the current underwater robot control system by introducing the pulse neural network (SNN) and multi-stage denoising convolution technology; the traditional system is difficult to efficiently process complex noise in the underwater environment, resulting in low target detection accuracy, and even affecting the subsequent path planning and obstacle avoidance effect; and the present application significantly improves the accuracy and robustness of target detection by optimizing the noise suppression capability of the sonar image, ensuring that the robot can accurately identify targets and obstacles in complex underwater environments, providing a more reliable basis for the robot to perform tasks.
[0025] At the same time, the system adopts an extended PPO reinforcement learning model, and by introducing a lower bound of policy update, the problem of slow convergence speed and poor stability of the existing path planning model in dynamic underwater environment is solved; the traditional PPO model often appears slow convergence and unstable path planning when dealing with rapidly changing environment, which is difficult to ensure the efficient operation of the underwater robot in the variable environment; and the present application enhances the convergence speed and stability of the strategy, so that the robot can adjust the path planning in real time in the underwater environment, ensuring the obstacle avoidance effect and efficiency of task execution of the robot in the multi-obstacle environment.
[0026] Through these technical innovations, the present application realizes efficient target detection, three-dimensional obstacle positioning, real-time environment mapping and path planning, significantly improving the autonomy and reliability of the underwater robot in tasks such as environmental monitoring, object detection and seabed construction; compared with the traditional system, the present application not only optimizes the algorithm performance, improves the system response speed and task execution accuracy, but also greatly enhances the adaptability and autonomy of the underwater robot in complex and dynamic underwater environment, can efficiently cope with various uncertain factors, thereby improving the success rate of task completion and the overall performance of the robot in complex tasks. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 The path planning and actual travel trajectory comparison chart in Example One;
[0028] Figure 2 The convergence curve chart of the extended PPO model training process proposed in Example Four.
[0029] Figure 1 The blue dotted circle trajectory represents the planning path, the red solid circle trajectory represents the actual path, and the gray arrow represents the offset error direction between the planning point and the actual point; the lower side is the X coordinate (m), and the left side is the Y coordinate (m);
[0030] Figure 2 In the figure, the blue curve: the cumulative reward (Reward) rises with the training round number, showing a gradual saturation trend, and the strategy performance converges; the red curve: the strategy loss (Loss) decreases with the round number, indicating that the strategy update gradually stabilizes. DETAILED DESCRIPTION
[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0032] In an embodiment, the present application provides an underwater robot control system for detection, comprising a carrier platform, a multi-sensor fusion module and a core controller. Figure 1
[0033] The carrier platform is used to carry a detection device, and the detection device comprises an inertial measurement unit (IMU), a depth sensor, an imaging sonar unit and an underwater positioning unit.
[0034] The detection device comprises:
[0035] The IMU is Xsens MTi-680G (three-axis acceleration + gyroscope + magnetometer).
[0036] The depth sensor is Blue Robotics Bar100 (±0.1m precision).
[0037] The imaging sonar is Teledyne BlueView P900-130, with a resolution of 480x640 and a frame rate of 10Hz.
[0038] The underwater positioning system is an LBL array + USBL fusion positioning system, with an error controlled within ±0.25m.
[0039] The multi-sensor fusion module collects robot posture, depth, sonar image and position data in real time through the detection device, obtains multi-source sensor data, and realizes multi-source sensor data fusion through a Kalman filtering algorithm to output high-precision state estimation data.
[0040] The core controller is connected to the multi-sensor fusion module and is configured to:
[0041] The pulse underwater YOLO model is used to process the sonar image, extract the image features of obstacles and targets, generate a target class probability graph and a bounding box coordinate graph, and combine the high-precision state estimation data to perform three-dimensional positioning of obstacles and generate spatial semantic information. The pulse underwater YOLO model is based on a conventional YOLO model, and is constructed by introducing a pulse neural network and a multi-level denoising convolution kernel technology to optimize the feature extraction and noise suppression capability of the conventional YOLO model.
[0042] The target is detected at t=8.2s , (category: rock):
[0043] Image coordinates (bbox): [x = 124, y = 280, w = 36, h = 42] ;
[0044] The actual coordinates after conversion are: [X = 7.4m, Y = 12.6m, Z = - 14.2m] ;
[0045] The three-dimensional positioning error is verified to be ±0.42m;
[0046] Combined with spatial semantic information, a real-time environment map is constructed by SLAM;
[0047] Based on the real-time environment map, a robot control instruction is generated by extending the PPO model to execute obstacle avoidance path planning; the extended PPO model is constructed by introducing a strategy update lower bound to optimize the strategy stability and convergence speed of the traditional PPO model;
[0048] Figure 1 The path planning generated by the robot in the embodiment when performing the obstacle avoidance path planning task is compared with the actual travel trajectory; the blue dotted circle point trajectory represents the planning path, the red solid circle point trajectory represents the actual path, and the gray arrow represents the offset error direction between the planning point and the actual point; the lower side is the X coordinate (m), and the left side is the Y coordinate (m).
[0049] Embodiment two, based on embodiment one, in this embodiment, the generation process of the target category probability graph and the bounding box coordinate graph specifically includes the following steps:
[0050] Step S1: normalize the pixel value of the sonar image to generate normalized image data;
[0051] Step S2: the normalized image data is preliminarily extracted through the convolution layer and the batch normalization layer to generate a floating-point feature map, and the floating-point feature map is copied into a time step format and input into an integral-firing neuron (IF neuron) to generate a pulse feature map;
[0052] Step S3: three groups of denoising convolution kernels are preset, the three groups of denoising convolution kernels include a horizontal difference kernel (used for detecting horizontal pulse interference), a vertical difference kernel (used for detecting vertical pulse interference) and a Laplace kernel (used for detecting isolated noise points), and the three groups of denoising convolution kernels are used to perform sliding window convolution operation on the pulse feature map respectively to obtain multi-channel convolution response results; a threshold judgment mechanism is introduced, and the multi-channel convolution response results are combined to perform pulse point detection and pixel flipping processing, and a noise repair pulse feature map is output, and the formula used is as follows:
[0053] The multi-kernel convolution response calculation formula is:
[0054]
[0055] wherein, denotes the denoising kernel number, including horizontal difference kernel, vertical difference kernel and Laplacian kernel; denotes the convolution response value of the i-th kernel at the image coordinate point , used for judging whether it is a noise point or not; denotes the kernel radius, denotes the double summation convolution operation, denotes the pixel value of the coordinate point offset by denotes the corresponding weight value of the i-th kernel, denotes the position index in the kernel (since the subscript starts from 0, it needs to be offset by +1); The pixel flipping and noise repair judgment rule is:
[0056]
[0057]
[0058] denotes the output image pixel value at the point denotes the noise response judgment threshold, denotes the maximum value in all kernel responses, denotes the original pixel value of the input image;
[0059] If the maximum response value is not less than the threshold, it is judged as a noise point, and the flipping processing is performed, otherwise the original value remains unchanged;
[0060] Step S4: input the noise repair pulse feature map into the backbone network to generate a deep feature map, the backbone network is composed of pulse unit modules, the pulse unit modules are constructed by introducing the CSPNet residual structure, combining feature channel splitting, local convolution and cross-path activation preservation mechanism, for improving the deep neuron pulse firing rate, relieving the “pulse degradation” problem, enhancing the deep pulse firing rate, relieving the activation sparsity problem caused by the depletion of the integral-discharge type neuron membrane potential, and realizing the stable extraction of multi-scale pulse features;
[0061] The pulse unit module specifically includes:
[0062] Channel splitting: The output of the upper-layer pulse unit module is divided into two groups of feature maps along the channel direction after convolution. This is used for main / sub-branch division, reflecting the cross-path activation preservation mechanism. By sharing and fusing information between the main and sub-branch paths, the information flow and activation signals between multiple paths can be maintained, thereby enhancing the effective activation of each layer in the network during processing. The formula used is as follows:
[0063] ;
[0064] in, Indicates the upper-level pulse unit module (the first) The output feature map of the layer. This represents the feature map (main channel) of the current layer's main branch. This represents the feature map (secondary channel) of the current layer's secondary branch. This represents the convolution operation. This indicates an equal division operation along the channel dimension;
[0065] Secondary trunk pulse activation and local convolution: This involves branch feature maps... The input is fed into the IF neuron, and the output is processed by time-step pulse encoding. Then, a set of local convolutions is used to extract deep pulse response features, preserving the detail intensity and spatial expansion capability of the pulse expression.
[0066] Fusion and Residual Enhancement: Improving the Original Backbone The results of the secondary processing are concatenated to form a deep output feature map. After fusing the two pulse expression modes, the residual enhancement path of the previous layer output is introduced to form the deep output feature map. The formula used is as follows:
[0067] ;
[0068] in, Indicates the current number The output of the layer pulse unit module, This represents the activation function of the IF spiking neuron. Membrane potential accumulation and emission are performed, and pulse coding characteristics are output. This indicates a splicing operation. Indicates to Perform convolution;
[0069] Step S5: Perform max pooling and activation operations on the deep feature map to generate a fused feature map. Based on the fused feature map, generate a target class probability map and a bounding box coordinate map.
[0070] Example 3, based on Example 1, specifically includes the following steps in the generation process of the target category probability map and bounding box coordinate map:
[0071] Step R1: normalize the sonar image pixel value to generate normalized image data;
[0072] Step R2: perform preliminary feature extraction on the normalized image data through the convolution layer and the batch normalization layer to generate a floating-point feature map, copy the floating-point feature map into a time step format and input it into the integrate-fire neuron (IF neuron) to generate a pulse feature map;
[0073] Step R3: preset three groups of denoising convolution kernels, including a horizontal difference kernel (for detecting horizontal pulse interference), a vertical difference kernel (for detecting vertical pulse interference) and a Laplace kernel (for detecting isolated noise points), perform a sliding window convolution operation on the pulse feature map using the three groups of denoising convolution kernels to obtain a multi-channel convolution response result; introduce a threshold judgment mechanism, combine the multi-channel convolution response result, perform pulse point detection and pixel flipping processing, and output a noise repaired pulse feature map;
[0074] Step R4: input the noise repaired pulse feature map into the backbone network to generate a deep feature map, the backbone network is composed of a neural network module, and the input data is processed layer by layer using a conventional feature map transmission and nonlinear transformation method to gradually extract more discriminative image features;
[0075] Step R5: perform maximum pooling and activation operations on the deep feature map to generate a fusion feature map, and generate a target class probability map and a bounding box coordinate map according to the fusion feature map.
[0076] In this embodiment, an extended PPO model is used to generate robot control instructions to perform obstacle avoidance path planning and target tracking tasks, which specifically includes the following steps: Figure 2
[0077] Step B1: collect historical trajectory data and combine a real-time environment map as the basis for strategy optimization of the extended PPO model, use generalized advantage estimation to obtain an original advantage value of the robot in the current state, calculate statistical indicators according to the original advantage value and perform normalization processing to generate a modulated advantage value, the statistical indicators include L2 norm and standard deviation;
[0078] Step B2: use the modulated advantage value to construct a strategy loss function and iteratively update the strategy network parameters; introduce an extended strategy update lower bound to control the update amplitude of the strategy network to ensure the stability in the strategy optimization process, the formula is as follows:
[0079] Strategy loss function formula:
[0080] ;
[0081] in, Represents the policy loss function. Indicates the current strategy. Indicates the policy network parameters, Indicates the sample index. E j [] Represents the mathematical expectation. This represents the probability ratio between the current strategy and the historical strategies. Indicates the modulated dominance value. Indicates the clipping threshold. Represents the clipping function make sure exist [1 - ε, 1 + ε] Within the range;
[0082] The formula for updating the lower bound using the extended strategy is:
[0083] ;
[0084] in, Indicates the current strategy Expected reward Representing historical strategies Expected reward Indicates the discount factor. Indicating in historical strategy Induced state-action distribution Below, for state-action pairs Expectations; Represents the dominance function. Indicating the dominance function The upper realm, Indicates the state of the new and old strategies. The total variational distance below;
[0085] The core function of the extended policy update lower bound is to ensure that the policy improves at least in each update. In reinforcement learning, it is generally desirable that each policy update improves the agent's performance, rather than just causing random fluctuations. The extended policy update lower bound provides a theoretical guarantee of minimum improvement, ensuring that the expected reward of the current policy is higher than the expected reward of the historical policy in each update, at least reaching an expected gain.
[0086] By introducing the total change distance and the magnitude of the maximum advantage value The extended policy update lower bound helps limit the magnitude of policy update, avoiding too large policy change at each update, which can cause unstable learning, or even make the agent rapidly deviate from a better state, resulting in non-convergence or slow convergence during training; by controlling the update magnitude, the extended policy update lower bound ensures that each update is within a reasonable range, preventing drastic fluctuations in policy;
[0087] Step B3: According to the modulated advantage value, a training target is constructed to optimize and train the value function network, ensuring that the training signals of the policy network and the value function network are consistent, improving the generalization ability of the control policy, and generating an optimized policy distribution;
[0088] In this embodiment, the Figure 2 : In the 0-100 training round, the convergence curve of the extended PPO model training process is shown. The blue curve shows that the cumulative reward (Reward) increases with the training round, showing a gradual saturation trend, and the policy performance converges. The red curve shows that the policy loss (Loss) decreases with the round, indicating that the policy update is gradually stable.
[0089] Step B4: According to the optimized policy distribution, action sampling is performed, and the corresponding thruster control parameters are mapped to specific thrust commands, including thrust direction, amplitude, and execution duration, to form robot control commands.
[0090] Step B5: Execute the robot control command to drive the thruster to achieve hovering, pitching, yawing, and depth-keeping motion; execute obstacle avoidance path planning and target tracking tasks.
[0091] In this embodiment, based on embodiment two, the process of generating robot control commands and executing obstacle avoidance path planning and target tracking tasks includes the following steps:
[0092] Step E1: Collect historical trajectory data and combine real-time environment maps as the basis for policy optimization of the extended PPO model, use generalized advantage estimation to obtain the original advantage value of the robot in the current state, calculate statistical indicators based on the original advantage value, and perform normalization to generate modulated advantage values. The statistical indicators include L2 norm and standard deviation.
[0093] Step E2: Use the modulated advantage value to build a policy loss function and iteratively update the policy network parameters.
[0094] Step E3: According to the modulated advantage value, a training target is constructed to optimize and train the value function network, ensuring that the training signals of the policy network and the value function network are consistent, improving the generalization ability of the control policy, and generating an optimized policy distribution.
[0095] Step E4: action sampling is performed according to the optimized strategy distribution, and the corresponding thruster control parameters are mapped to specific thrust instructions, including thrust direction, amplitude and execution duration, to form robot control instructions;
[0096] Step E5: execute the robot control instructions to drive the thrusters to realize hovering, pitching, yawing and depth-keeping movements; execute obstacle avoidance path planning and target tracking tasks.
[0097] The above describes the present application and its embodiments, which are not restrictive, and the drawings only show one of the embodiments of the present application, and the actual structure is not limited thereto; in general, if a person skilled in the art is inspired thereby, without departing from the purpose of the present application, similar structural modes and embodiments are not created by creative design, which should belong to the protection scope of the present application.
Claims
1. A control system for an underwater robot used for detection, characterized in that, The system includes: a carrier platform, a multi-sensor fusion module, and a core controller; The carrier platform carries the detection equipment; The multi-sensor fusion module collects robot posture, depth, sonar images and position data in real time through detection devices, and outputs high-precision state estimation data; The core controller, connected to the multi-sensor fusion module, is configured as follows: The pulsed underwater YOLO model is used to process sonar images, generate target category probability maps and bounding box coordinate maps, and combine them with high-precision state estimation data to perform 3D obstacle localization and generate spatial semantic information. By combining spatial semantic information, SLAM is used to construct a real-time environment map; Historical trajectory data is collected and combined with real-time environmental maps. An extended PPO model is used to generate robot control commands to perform obstacle avoidance path planning and target tracking tasks.
2. The underwater robot control system for detection according to claim 1, characterized in that: The pulsed underwater YOLO model consists of a backbone network, convolutional layers, and batch normalization layers.
3. The underwater robot control system for detection according to claim 1, characterized in that: The extended PPO model includes a policy network and a value function network.
4. The underwater robot control system for detection according to claim 2, characterized in that: The process of generating the target category probability map and bounding box coordinate map includes the following steps: Step S1: Normalize the pixel values of the sonar image to generate normalized image data; Step S2: Perform preliminary feature extraction on the normalized image data through convolutional layers and batch normalization layers to generate floating-point feature maps. Copy the floating-point feature maps into time-step format and input them into the integral-fire neuron to generate pulse feature maps. Step S3: Preset three sets of denoising convolution kernels, and use the three sets of denoising convolution kernels to perform sliding window convolution operation on the pulse feature map to obtain multi-channel convolution response results; introduce a threshold judgment mechanism, combine the multi-channel convolution response results, perform pulse point detection and pixel flipping processing, and output noise-repaired pulse feature map; Step S4: Input the noise-repaired impulse feature map into the backbone network to generate a deep feature map; Step S5: Perform max pooling and activation operations on the deep feature map to generate a fused feature map. Based on the fused feature map, generate a target class probability map and a bounding box coordinate map.
5. The underwater robot control system for detection according to claim 4, characterized in that: Three sets of denoising convolution kernels, including horizontal difference kernel, vertical difference kernel, and Laplacian kernel.
6. The underwater robot control system for detection according to claim 4, characterized in that: The backbone network consists of pulse unit modules, which are constructed by introducing the CSPNet residual structure and combining feature channel splitting, local convolution, and cross-path activation preservation mechanisms.
7. The underwater robot control system for detection according to claim 3, characterized in that: The process of generating robot control commands using an extended PPO model to perform obstacle avoidance path planning and target tracking tasks includes the following steps: Step B1: Collect historical trajectory data and combine it with the real-time environment map as the basis for optimizing the extended PPO model strategy. Use generalized advantage estimation to obtain the robot's original advantage value in the current state. Calculate statistical indicators based on the original advantage value and perform normalization processing to generate modulated advantage values. The statistical indicators include L2 norm and standard deviation. Step B2: Construct a policy loss function using the modulated advantage value and iteratively update the policy network parameters; introduce an extended policy update lower bound to control the update magnitude of the policy network; Step B3: Construct a training objective based on the modulated advantage value, optimize the value function network, ensure that the training signals of the policy network and the value function network are consistent, and generate an updated policy distribution; Step B4: Sample actions based on the updated policy distribution to generate robot control commands; Step B5: Execute robot control commands to perform obstacle avoidance path planning and target tracking tasks.
Citation Information
Patent Citations
AUV dynamic obstacle avoidance method based on near-end strategy optimization algorithm
CN115291616A
Underwater net cage inspection robot based on side-scan sonar and AUV platform
CN119460031A
Underwater target detection method based on spiking neural network
CN120451760A
Cited By
Large shield slurry pipeline detection robot obstacle crossing algorithm system based on AI calculation
CN121433312A
Visual servo scanning track control system of deepwater ROV intelligent electromagnetic detection robot
CN121857499A