A method, system and medium for controlling a biomimetic underwater robot

By performing beam pointing compensation and motion distortion correction on the acoustic scanning data of the biomimetic underwater robot, and constructing a simulation environment in conjunction with a dynamic model, the problems of acoustic scanning data mismatch and pose estimation of the biomimetic underwater robot in complex seabed terrain were solved. Stable autonomous detection and control closed loop were achieved, improving the stability of terrain detection and the training efficiency of control strategies in the marine environment.

CN122284631APending Publication Date: 2026-06-26HOHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HOHAI UNIV
Filing Date
2026-05-14
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing biomimetic underwater robots are prone to beam pointing mismatch and point cloud distortion accumulation in acoustic scanning data under complex seabed topography, strong near-bottom disturbances, and non-uniform oscillating propulsion conditions. This leads to inaccurate robot pose estimation, reduced consistency of seabed topography maps, and difficulty in forming a stable closed loop for detection, perception, localization, mapping, and autonomous control.

Method used

By introducing the motion state data of the robot body, beam pointing compensation, pose registration and motion distortion correction are performed on the acoustic scanning data acquired by multi-beam mapping sonar. Combined with the dynamic model, a simulation environment is constructed to realize stable correction of acoustic scanning data and autonomous detection control closed loop.

Benefits of technology

Under conditions of non-uniform oscillating propulsion and strong near-bottom disturbance, the spatial consistency and mapping accuracy of acoustic scanning data are maintained, the robustness of pose estimation and the global consistency of seabed topographic maps are enhanced, an autonomous closed-loop mechanism of detection and perception, positioning and mapping and control is formed, and the training efficiency of control strategies and the adaptability to real platforms are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122284631A_ABST
    Figure CN122284631A_ABST
Patent Text Reader

Abstract

This application discloses a biomimetic underwater robot control method, system, and medium: controlling the robot body to perform exploration tasks in a target sea area and acquiring acoustic scanning data of the seabed topography, motion state data of the robot body, and seabed environment perception data; based on the motion state data, performing beam pointing compensation, pose registration, and motion distortion correction on the acoustic scanning data; based on the corrected acoustic scanning data, motion state data, and seabed environment perception data, performing underwater positioning and map building of the robot body; generating motion control commands for the robot body based on the seabed topography map, pose estimation results, and the current exploration task; and adjusting the propulsion state and / or attitude state of the robot body. This application solves the problem that under non-uniform oscillating propulsion conditions, the acoustic scanning data of multi-beam mapping sonar is prone to beam pointing mismatch, leading to inaccurate pose estimation of the robot body and decreased consistency of the seabed topography map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control, and in particular to a biomimetic underwater robot control method, system and medium. Background Technology

[0002] The ocean contains abundant topographic, resource, and environmental information, and seabed topographic surveying is a fundamental link in marine environmental understanding, resource development, underwater engineering construction, and marine safety assurance. Due to the characteristics of the seabed environment, such as high pressure, low light, strong disturbance, and complex topography, underwater robots have become an important vehicle for performing seabed exploration missions. Existing underwater exploration platforms mainly include remotely operated vehicles (ROVs) and autonomous underwater vehicles (AUVs), but they generally suffer from insufficient maneuverability, limited environmental adaptability, and poor stability during continuous operation in complex seabed topographic environments. Especially in seabed canyons, reef areas, hydrothermal vent areas, and areas with significant near-bottom undulations, traditional propeller-driven underwater robots struggle to balance stable navigation with precise detection, easily leading to increased blind spots, track deviations, and decreased data quality.

[0003] Biomimetic underwater robots, by mimicking the morphology and propulsion mechanisms of marine organisms such as fish and dolphins, possess advantages in propulsion efficiency, maneuverability, environmental concealment, and adaptability to complex flow fields, and have become an important research direction in recent years. However, existing research on biomimetic underwater robots mainly focuses on propulsion efficiency, structural design, or optimization of single motion control performance, and is still insufficient for specialized integration for seabed topographic exploration tasks. In particular, there is a lack of systematic technical solutions suitable for biomimetic propulsion platforms in areas such as multi-source sensor integration, real-time processing of detection data, underwater positioning and mapping accuracy, and detection and control coordination.

[0004] Seabed topography surveying typically employs techniques such as acoustic, optical, and magnetic detection. Among these, multibeam sonar, due to its wide coverage and high depth-sounding efficiency, has become a crucial method for seabed topography mapping. Because seawater significantly attenuates electromagnetic waves, sound waves remain the most effective information carrier for long-distance underwater topographic sensing. Therefore, topographic reconstruction, underwater positioning, and map building based on acoustic scanning data constitute the core technological chain of marine topography surveying.

[0005] However, most existing multibeam mapping sonar application methods are based on the assumption of near-uniform motion and minimal attitude disturbance of traditional underwater platforms, and their data processing workflows are typically designed for conventional rigid-body propulsion platforms. When these methods are applied to biomimetic underwater robots, the robot body experiences periodic attitude changes, velocity fluctuations, and spatial position shifts during navigation, especially under non-uniform oscillating propulsion conditions, due to the oscillating propulsion method used by the biomimetic platform. Simultaneously, in complex near-bottom sea areas, the superposition of topographic undulations, current disturbances, and environmental echo interference easily leads to problems such as beam pointing mismatch, sampling point registration deviation, and accumulated point cloud distortion in the acoustic scanning data acquired by multibeam mapping sonar. This results in inaccurate robot pose estimation and inaccurate seabed topographic data. Figure 1 This leads to a decrease in consistency and makes it difficult to form a stable closed loop between detection and sensing, localization and mapping, and autonomous control.

[0006] Furthermore, the autonomous control of biomimetic underwater robots faces the challenge of high-dimensional, strongly coupled, nonlinear mapping between control parameters and navigation actions. Training the policy network directly in a real underwater environment presents problems such as high training costs, difficulty in obtaining samples, and significant environmental uncertainties. If the simulation environment cannot accurately reflect the dynamic characteristics of the biomimetic propulsion process, the control strategy obtained from the simulation is difficult to effectively transfer to the real platform, further affecting the application effect of autonomous navigation control in actual marine exploration missions.

[0007] Therefore, how to achieve stable correction of acoustic scanning data, reliable underwater positioning and mapping, and autonomous detection and control closed loop for biomimetic propulsion platforms under complex undulating seabed topography, strong near-bottom disturbances, and non-uniform oscillating propulsion conditions has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0008] In existing technologies, under conditions of complex undulating seabed topography, strong near-bottom disturbances, and non-uniform oscillating propulsion, the acoustic scanning data acquired by multibeam mapping sonar is prone to beam pointing mismatch and point cloud distortion accumulation, which in turn leads to inaccurate robot pose estimation and seabed topography. Figure 1 To address the issues of decreased consistency and the difficulty in achieving stable closed-loop control for detection, localization, mapping, and autonomous control, this application provides a biomimetic underwater robot control method, system, and medium. This method enables stable correction of acoustic scanning data, reliable underwater localization and mapping, and autonomous detection and control closed-loop control for biomimetic propulsion platforms.

[0009] One aspect of this application provides a biomimetic underwater robot system, comprising: a robot body; a watertight electronic compartment disposed inside the robot body; a biomimetic actuator connected to the robot body; a multibeam sonar disposed on the robot body; and an inertial measurement unit, a depth gauge, an optical imaging module, and an industrial control computing module disposed on the robot body. Among them, the bionic actuator is used to drive the robot body to perform forward movement, turning, hovering attitude adjustment and anti-current attitude adjustment; the multibeam mapping sonar is used to acquire acoustic scanning data of seabed topography; the inertial measurement unit, depth gauge and optical imaging module are used to collect motion state data of the robot body and environmental perception data; The industrial control computing module is connected to the bionic actuator, multibeam mapping sonar, inertial measurement unit, depth gauge, and optical imaging module, respectively, and is used for: performing beam pointing compensation and motion distortion correction on acoustic scanning data based on the robot's motion state data to obtain corrected terrain data; performing simultaneous localization and map building based on the corrected terrain data, the robot's motion state data, and environmental perception data to obtain the robot's pose estimation result and seabed topographic map; and generating drive control commands based on the seabed topographic map, pose estimation result, and current exploration task to control the robot to perform ocean topographic exploration tasks. Another aspect of this application provides a biomimetic underwater robot control method, applied to the biomimetic underwater robot system of this application: The robot body is controlled to perform exploration tasks in the target sea area and acquire acoustic scanning data of the seabed topography, motion state data of the robot body, and seabed environment perception data. Among them, the acoustic scanning data is acquired through multibeam sonar, the motion state data is acquired through inertial measurement unit and depth gauge, and the seabed environment perception data is acquired through optical imaging module. Based on motion state data, beam pointing compensation, pose registration and motion distortion correction are performed on acoustic scanning data to obtain corrected acoustic scanning data. Based on the corrected acoustic scanning data, motion state data, and seabed environment perception data, the robot performs underwater localization and map building, and obtains the robot's pose estimation results and seabed topography map. Based on the seabed topographic map, pose estimation results, and current exploration mission, generate motion control commands for the robot body; Based on motion control commands, the propulsion state and / or attitude state of the robot body are adjusted to perform subsequent marine topography exploration tasks; wherein, the adjustment of the propulsion state and / or attitude state is achieved through biomimetic actuators; Compared to existing technologies, the advantages of this application are: By incorporating the robot's motion state data, beam pointing compensation, pose registration, and motion distortion correction are sequentially applied to the acoustic scanning data acquired by multibeam mapping sonar. This ensures that the acoustic scanning data maintains high spatial consistency and mapping accuracy even under non-uniform oscillating propulsion and strong near-bottom disturbance conditions, effectively mitigating point cloud distortion caused by attitude fluctuations, propulsion disturbances, and environmental echo interference. Furthermore, by combining the corrected acoustic scanning data with seabed environmental perception data for underwater positioning and map building, the robustness of pose estimation results and the global consistency of the seabed topographic map are enhanced. Finally, by further using the seabed topographic map and pose estimation results to generate motion control commands for the robot, the robot can adjust its propulsion and / or attitude states in real time according to the current detection task, thus forming an autonomous closed-loop mechanism of mutual feedback between detection perception, positioning mapping, and motion control.

[0010] Furthermore, this application establishes a dynamic model based on the robot's body dynamic parameters and constructs a simulation environment that matches the motion characteristics of the real platform. This allows the control strategy training process to reflect the dynamic characteristics of the biomimetic underwater robot during its oscillating propulsion process, such as propulsion force, fluid resistance, lift, and torsional torque, thereby improving the simulation environment's ability to represent the real motion process. Simultaneously, the action space of the control strategy model is set as high-level control parameters for adjusting the robot's oscillating propulsion. This transforms the strategy network output from directly facing the low-level drive execution to facing the adjustment of high-level oscillating behavior, which helps reduce the learning difficulty caused by the high-dimensional nonlinear mapping between terrain detection and navigation actions and biomimetic propulsion control, improving the efficiency of control strategy training and adaptability to the real platform. By setting the time ratio parameter, the asymmetric oscillating behavior during the robot's turning process can be characterized, thereby improving the turning control capability and trajectory tracking capability in complex terrain environments. Therefore, this application can significantly improve the stability of terrain detection, the reliability of localization and mapping, the efficiency of control strategy training, and the ability to transfer simulation strategies to real platforms for biomimetic underwater robots in complex marine environments. Attached Figure Description

[0011] This application will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is a schematic diagram illustrating a real-world application scenario of a biomimetic underwater robot platform used for ocean topography exploration. Figure 2 A schematic diagram of the system architecture and information flow of a biomimetic underwater robot platform used for ocean topography exploration; Figure 3Flowchart of an underwater synchronous localization and mapping system for a biomimetic underwater robot used for ocean topography exploration; Figure 4 This is an exemplary flowchart of a biomimetic underwater robot control method according to this application; Figure 5 The output of the CPG model is shown in the figure with different parameter inputs for the biomimetic underwater robot platform used for ocean topography exploration. Figure 6 The graph shows the hydrodynamic forces in the x and y directions of a biomimetic underwater robot used for ocean topography exploration as a function of time. Figure 7 This is a diagram of the core intelligent control system architecture for a biomimetic underwater robot used for ocean topography exploration.

[0012] Explanation of the labels in the diagram: 1. Robot body; 2. Watertight electronic cabin; 3. Bionic actuator; 4. Multibeam mapping sonar. Detailed Implementation

[0013] The methods and systems provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a more comprehensive understanding of the biomimetic underwater robot platform, marine topography detection method, and related control and positioning mapping scheme provided in this application. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of protection of this application. For those skilled in the art, various substitutions, modifications, or improvements can be made to this application without departing from its conceptual framework, and all such modifications and modifications should fall within the scope of protection of this application.

[0015] Example 1 This application aims to address the shortcomings of existing seabed topographic detection technologies under conditions of complex undulating seabed topography, strong near-bottom disturbances, and non-uniform oscillating propulsion. These shortcomings include acoustic scanning data beam pointing mismatch, point cloud distortion accumulation, inaccurate pose estimation, and other issues. Figure 1This application addresses technical challenges such as decreased consistency and the difficulty in forming a stable closed loop for detection, localization, mapping, and autonomous control. Furthermore, it proposes an integrated perception, localization, mapping, and control scheme for biomimetic propulsion platforms to address the issue of high-dimensional, strongly coupled nonlinear mapping between control parameters and navigation actions in biomimetic underwater robots, which leads to low strategy training efficiency and difficulty in effectively transferring simulation strategies to real platforms. This scheme integrates platform-level sensors, actuators, power supply modules, and control computing modules. It combines multi-source data preprocessing (such as sonar and inertial data), motion state estimation based on dynamic models, loop closure detection and pose graph optimization, map construction, and reinforcement learning-based motion control strategy generation to build an autonomous closed-loop technology system encompassing perception, preprocessing, motion estimation, localization and mapping, decision control, execution feedback, and re-perception. This enables stable detection and reliable reconstruction of seabed topography and autonomous operation of the robot itself.

[0016] In one embodiment, such as Figure 1 As shown, the biomimetic underwater robot platform provided in this application includes a robot body 1, a watertight electronic cabin 2, a biomimetic actuator 3, a multibeam sonar 4, an inertial measurement unit (IMU), a depth sensor, an underwater camera, a transducer, a current sensor, an industrial computer, a data acquisition card, a transceiver amplifier board, a power supply battery, and a DC conversion module. See attached document for details. Figure 2 In this embodiment, the biomimetic underwater robot platform can be divided into a decision-making layer, a sensing layer, a driving layer, and a power supply layer according to its functions. The decision-making layer, sensing layer, driving layer, and power supply layer are located inside the robot fish, while the multi-beam mapping sonar 4 is located outside the robot fish and is connected to the internal control system through a transceiver amplifier board and a data acquisition card.

[0017] Specifically, the decision-making layer may include an industrial control computer (ICC) and a data acquisition card. The ICC is used for task scheduling, sensor data processing, motion state estimation, simultaneous localization and mapping (SLAM) construction, control strategy reasoning, and motion control command generation; the data acquisition card is used to manage the data acquisition and transmission interface between the ICC and peripheral devices. (Appendix) Figure 2 The industrial control computer and the data acquisition card shown can communicate bidirectionally. The industrial control computer sends acquisition control commands to the data acquisition card and receives sonar data and other peripheral data uploaded by the data acquisition card. The data acquisition card is also connected to a transceiver amplifier board, used to send the sonar control signals output by the industrial control computer to the transceiver amplifier board and receive the acoustic measurement data returned by the transceiver amplifier board.

[0018] In one embodiment, the transceiver amplifier board is used for electrical connection and signal interaction with the multibeam mapping sonar 4. Specifically, the transceiver amplifier board can amplify the power of the transmit drive signal and receive and front-end process the echo signal returned by the multibeam mapping sonar 4, thereby realizing the integration of sonar signal transmission and reception. (Appendix) Figure 2 The multibeam mapping sonar 4 shown is located outside the robotic fish and is used to map and scan the seabed. The acoustic scanning data it acquires is transmitted to the industrial control computer via a transceiver amplifier board and a data acquisition card for subsequent data preprocessing, beam compensation, motion distortion correction, point cloud generation, and map construction.

[0019] In one embodiment, the sensing layer may include an underwater camera, a current sensor, an inertial measurement unit (IMU), a depth sensor, and a transducer. The underwater camera is used to acquire images of the seabed environment; the IMU is used to acquire the robot's attitude angles, angular velocities, and acceleration information; the depth sensor is used to acquire the robot's depth information; the current sensor is used to detect current changes during platform operation to reflect the drive load status, energy consumption status, or equipment operating status; and the transducer can be used for acoustic signal conversion or other underwater sensing / communication functions. All multi-source data acquired by the above sensors can be transmitted to an industrial control computer as input information for environmental perception, state estimation, energy consumption monitoring, and task control.

[0020] In one embodiment, the drive layer may include multiple pectoral fin servos, a tail servo, and a lead screw motor. (See appendix) Figure 2 The illustration schematically depicts multiple pectoral fin servos, a tail servo, and a lead screw motor. The pectoral fin servos drive the left and right pectoral fins to perform flapping, spreading, retracting, or attitude trimming actions, assisting in heave control, attitude stabilization control, low-speed maneuvering control, or anti-current adjustment. The tail servo drives the flexible tail section, tail fin, or tailstock structure to perform periodic oscillations, generating primary propulsion and steering torque. The lead screw motor can drive internal mechanisms to perform center of gravity adjustment, buoyancy adjustment, sensor attitude adjustment, or other linear actuation actions. This application does not limit these aspects. Each actuator in the drive layer is connected to an industrial control computer or its corresponding drive control interface to execute corresponding actions according to motion control commands.

[0021] In one embodiment, the power supply layer may include a 24V 15AH lithium battery and a DC conversion module. The 24V 15AH lithium battery provides the main power input for the entire platform; the DC conversion module converts the battery output voltage to a target voltage suitable for the operation of the decision layer, sensing layer, and driving layer modules. (See appendix) Figure 2 The power supply layer shown is connected to both the sensing layer and the drive layer to provide stable power to the relevant sensors, servos, motors, data acquisition cards, power amplifier boards, and industrial control computers. By setting up an independent power supply layer and voltage conversion link, the power supply stability and system reliability of the platform during long-term underwater operations can be improved.

[0022] In one embodiment, the robot body 1 may adopt a streamlined fish-like shape, with its outline designed based on the morphology of fish, dolphins, or other marine organisms suitable for near-bottom navigation, to reduce hydrodynamic drag and improve maneuverability in complex flow fields. A watertight electronics compartment 2, located inside the robot body 1, houses the industrial control computer, data acquisition card, power supply battery, DC conversion module, interface circuits, and other electronic components, ensuring the airtightness and reliability of the electronic system in a high-pressure underwater environment. A biomimetic actuator 3 may include a tail servo, pectoral fin servos, and an actuator connected to a flexible drive mechanism, used to drive the robot body 1 to perform forward swimming, steering, hovering attitude adjustment, and current-resistant attitude adjustment. A multibeam mapping sonar 4 is located on the exterior or underside of the robot body 1, used to transmit and receive acoustic signals towards the seabed to acquire acoustic scanning data of the seabed topography. The industrial control computer can be connected to the bionic actuator 3, multibeam mapping sonar 4, inertial measurement unit (IMU), depth sensor, underwater camera, current sensor, and other functional modules to perform data acquisition, data preprocessing, initialization, motion state estimation, synchronous positioning and map building, motion control command generation, and task execution management.

[0023] In this embodiment, the appendix Figure 1 The illustration depicts the actual operational scenario of the biomimetic underwater robot platform of this application. The background area is a seabed canyon or slope terrain with significant undulations. The robot body 1 navigates within the target sea area according to a preset detection route or an autonomously planned route. The watertight electronics compartment 2 is located within the main body of the robot body 1 and integrates computing and power supply units. The biomimetic actuator 3 is located inside the robot body 1 and connected to a flexible torso or tail fin drive mechanism to generate fish-like undulating propulsion. The multibeam sonar 4 is located in the ventral region of the robot body 1 and performs strip scanning towards the seabed to continuously acquire acoustic scanning data of the seabed topography. Simultaneously, the inertial measurement unit (IMU) outputs attitude angle, angular velocity, and acceleration data, the depth sensor outputs depth data, and the underwater camera acquires images of the seabed environment. After the above multi-source sensor data is input into the industrial control computer within the watertight electronics compartment 2, preprocessing, correction, motion state estimation, and local map construction can be performed in real time on the robot. Optionally, the processed raw data, intermediate results, and map results can be stored in a solid-state storage device and exported through a physical interface after the task is completed, so that post-processing software can perform data fusion, 3D modeling, and visualization. This application embodiment does not limit this.

[0024] Please further details are as follows: Figure 3In this embodiment, the perception and localization mapping process of this application may include a sensing data layer, a data preprocessing layer, an initialization module, a motion estimation module, a loop closure detection and relocation module, a pose graph optimization engine, and a map building layer. Specifically, the sensing data layer may include a multi-beam sonar, an inertial measurement unit (IMU), an optical camera, and a depth gauge, used to collect raw multi-source observation data during seabed exploration. The data preprocessing layer is used to uniformly process the raw multi-source observation data. Its processing content may include: preprocessing sonar data, beam compensation and motion distortion correction, pre-integration and relative pose increment calculation of inertial data, covariance estimation of observation noise, and performing point cloud generation and optimization, coordinate transformation, filtering, and feature extraction. Through the above preprocessing, high-quality input data under a unified spatiotemporal reference can be provided for subsequent pose solving and map building.

[0025] In one embodiment, the initialization module is used to perform extrinsic parameter calibration and initial pose determination. Extrinsic parameter calibration is used to determine the mounting pose relationships of the multibeam sonar 4, inertial measurement unit (IMU), depth sensor, and underwater camera relative to the robot's body coordinate system or a unified reference coordinate system. Initial pose determination is used to provide the initial position and orientation of the robot body at the beginning of the task, serving as initial values ​​for subsequent motion state recursion and map construction.

[0026] In one embodiment, the motion estimation module can estimate the motion state changes of the robot body based on a dynamic model. Specifically, the motion estimation module can integrate inertial pre-integration results, depth information, sonar point cloud matching results, underwater image features, and robot body dynamic constraints to continuously estimate the position, attitude, velocity, and their changes of the robot body. Since biomimetic oscillating propulsion involves non-uniformity, periodicity, and enhanced near-bottom disturbances, relying solely on geometric matching can easily lead to estimation drift. Therefore, this application introduces state change estimation based on a dynamic model, enabling the pose solution process to simultaneously consider dynamic factors such as propulsion force, fluid resistance, lift, and torsional torque, thereby improving the stability and reliability of motion estimation results in complex seabed environments.

[0027] In one embodiment, the loop closure detection and relocalization module is used to match and identify historical keyframes or historical local maps to determine whether the robot body has returned to the detected area. (Appendix) Figure 3 The location identification / geometric verification shown can be used to determine geometric consistency between candidate loop closures to reduce the probability of false matches; when a valid loop closure is detected, relocation can be performed to correct pose drift errors accumulated during long-term navigation.

[0028] In one embodiment, the pose graph optimization engine is used to perform global optimization based on motion estimation constraints, loop closure constraints, and other observation constraints. Specifically, see Appendix Figure 3The "back-end: nonlinear optimization" can be understood as jointly solving for each key node and edge constraint in the pose graph to obtain a globally consistent pose estimation result. The optimized pose result can be further output to the map building layer to update the seabed topography map.

[0029] In one embodiment, the map construction layer may include a topological map, a scale map, a semantic map, and a hybrid map management module. The topological map expresses the connectivity between different exploration areas; the scale map may include a 3D point cloud map and an occupied grid map to express the spatial geometry of the seabed topography; the semantic map may include terrain classification results and feature annotation information to express seabed targets, special landforms, or environmental categories; and the hybrid map management module is used to perform multi-map fusion, hierarchical management, and task invocation. Through this map construction method, multiple task requirements such as seabed topographic mapping, environmental understanding, path planning, and detailed exploration can be simultaneously met.

[0030] In this embodiment, the robot body 1 preferably adopts a biomimetic tail fin oscillation propulsion method, that is, the flexible torso or tail fin is driven by the biomimetic actuator 3 to perform fish-like undulating motion, thereby generating propulsion force and steering torque. (See attached diagram) Figure 2 The tail servo motor serves as the primary actuator for tail fin oscillation propulsion, while the pectoral fin servo motors provide attitude stabilization and auxiliary maneuvering. The lead screw motor can be used for center-of-gravity adjustment or attitude balancing, collectively enabling the robot body 1 to achieve coordinated propulsion and attitude control in complex seabed environments. Compared to traditional propeller propulsion, biomimetic oscillating propulsion allows for more precise and maneuverable movements in complex seabed canyons, reef areas, hydrothermal vent areas, and areas with significant near-bottom undulations, while reducing mechanical noise and seabed sediment disturbance. Especially in the near... Figure 1 In the undulating seabed scenario shown, this propulsion method is more conducive to the robot body 1 performing bottom-hugging navigation, local turning, and low-speed stable observation in a small space near the bottom. Optionally, the shell shape of the robot body 1 can also incorporate fluid optimization design to further reduce underwater navigation resistance and improve current resistance stability.

[0031] Furthermore, based on the attached Figure 2 The platform's layered structure and appendix shown Figure 3The sensing, preprocessing, motion estimation, loop closure optimization, and map building processes illustrated in this application enable the unified fusion of multi-source data acquired from multibeam sonar, inertial measurement unit (IMU), underwater camera, and depth sensor. The resulting pose estimation and map results are then fed back to the subsequent control module, providing a state foundation for the robot's autonomous navigation and exploration tasks. Simultaneously, the decision layer generates motion control commands based on this state information, which are then executed by the drive layer for propulsion and attitude adjustment. The power supply layer provides stable power to each module, and the sensing layer continuously collects updated state information and returns it to the decision layer. This forms a closed-loop collaborative mechanism from perception, estimation, mapping to control execution and state feedback. Therefore, this application not only achieves stable acquisition of seabed topographic data and map building but also provides reliable state input and environmental representation for subsequent reinforcement learning-based motion control strategy generation, realizing integrated closed-loop collaboration of perception, localization, mapping, and control.

[0032] Example 2 In one embodiment, this ocean topography detection method is applied to the aforementioned biomimetic underwater robot platform and can be executed by the industrial control computing module calling the corresponding software program. For example... Figure 4 As shown, for ease of description, the method can be divided into steps S100 to S500.

[0033] In step S100, the robot body is controlled to perform a detection task in the target sea area and acquire acoustic scanning data of the seabed topography, motion state data of the robot body, and seabed environment perception data.

[0034] Specifically, the industrial control computing module can control the robot body 1 to enter the target sea area and navigate along a designated path according to preset task instructions, detection area information issued by the control terminal, or system autonomous planning results. The detection task can be area-wide mapping, route-based mapping, local fine reconstruction, or fixed-point terrain imaging, etc., and this embodiment does not limit this. During the robot's detection task, the multibeam sonar 4 continuously acquires acoustic scanning data of the seabed topography, the inertial measurement unit and depth gauge continuously output motion state data, and the optical imaging module outputs seabed environment perception data. Optionally, to improve the time alignment accuracy of multi-source data, the industrial control computing module can configure a unified timestamp or synchronization trigger mechanism for the multibeam sonar 4, inertial measurement unit, depth gauge, and optical imaging module to achieve time-synchronized data acquisition.

[0035] Step S200: Based on the motion state data, beam pointing compensation, pose registration and motion distortion correction are performed on the acoustic scanning data to obtain the corrected acoustic scanning data.

[0036] Among them, acoustic scanning data refers to seabed echo measurement data acquired by multibeam mapping sonar in one or more scanning cycles, including at least one or more of the following: transmission time, echo time, beam angle, slant range, and echo intensity for each beam; motion state data refers to sensor data characterizing the kinematic state of the robot body at each sampling time, including at least attitude angle data, angular velocity data, linear acceleration data, and depth data; beam pointing compensation refers to the process of correcting the spatial pointing of each transmitting / receiving beam of the multibeam sonar in the carrier coordinate system or navigation coordinate system based on the actual attitude and dynamic changes of the robot body at the scanning time; pose registration refers to the process of converting acoustic scanning points acquired at different sampling times to the same reference coordinate system according to a unified coordinate reference relationship; motion distortion correction refers to the process of compensating for and correcting the spatial position distortion of sampling points caused by non-uniform propulsion, periodic oscillation, attitude disturbance, and near-bottom flow field disturbance within the scanning time window.

[0037] Because this application employs a biomimetic oscillating propulsion method, the robot body generates periodic yaw, micro-vibrations, and velocity fluctuations during propulsion, which differ from those of traditional propeller-driven uniform propulsion platforms. This means that the sampling points corresponding to the multibeam sonar within a single scanning cycle are not acquired in a static or quasi-static pose. If this dynamic effect is not compensated for, the seabed measurement points acquired at different sampling times will exhibit strip stretching, lateral tearing, local distortion, and edge misalignment in a unified coordinate system, thereby reducing the accuracy of subsequent terrain reconstruction and localization mapping.

[0038] In one embodiment, step S200 can be further subdivided into steps S210 to S250.

[0039] Step S210: Obtain the sampling time sequence corresponding to the acoustic scanning data and the motion state data corresponding to the sampling time sequence, wherein the motion state data includes attitude angle, angular velocity, acceleration and depth data.

[0040] The sampling time sequence refers to a discrete time set formed according to the chronological order of the transmission, reception, or scanning frame generation of each beam during the detection process of a multibeam mapping sonar, denoted as: In the formula, Represents a sequence of sampling times; This represents the nth sampling time; n represents the total number of sampling times within the current processing window, and n = 1, 2, ..., N.

[0041] Specifically, the industrial control computing module can establish a sampling time sequence for acoustic scanning data based on the transmission time, echo acquisition time, or scan frame timestamp of the multibeam sonar; and extract motion state data corresponding to each sampling time from the inertial measurement unit and depth gauge based on a unified time axis. For the nth sampling time... The corresponding motion state data can be represented as: In the formula, Indicates the sampling time The corresponding set of motion state data; Indicates the roll angle; Indicates the pitch angle; Indicates the heading angle; , , These represent the angular velocity components of the robot body around its own x-axis, y-axis, and z-axis, respectively. These represent the linear acceleration components of the robot body along its own x-axis, y-axis, and z-axis, respectively. Indicates the robot body at the sampling time The depth value.

[0042] Step S220: Based on the attitude angle, angular velocity, and depth data, determine the beam pointing parameters and initial pose parameters corresponding to each sampling moment. The beam pointing parameters refer to the set of parameters characterizing the actual spatial pointing of each beam of the multi-beam mapping sonar in the reference coordinate system at a given sampling moment; the initial pose parameters refer to the set of parameters characterizing the initial spatial position and attitude of the robot body at a given sampling moment.

[0043] For the nth sampling time The robot's posture rotation matrix can be represented as: In the formula, Indicates the sampling time The attitude rotation matrix of the robot body relative to the navigation coordinate system; Indicates the roll angle The generated rotation matrix around the x-axis; Indicates the pitch angle The generated rotation matrix around the y-axis; Indicates the heading angle The generated rotation matrix around the z-axis.

[0044] For the j-th beam of a multibeam mapping sonar, its ideal direction vector in the sonar body coordinate system is denoted as . Then its actual beam direction vector in the navigation coordinate system can be expressed as: In the formula, This indicates that the j-th beam is at the sampling time. The actual direction vector in the corresponding navigation coordinate system; This represents the nominal direction vector of the j-th beam in the sonar coordinate system; represents the extrinsic rotation matrix of the sonar coordinate system relative to the robot's body coordinate system; j represents the beam number.

[0045] Therefore, the beam pointing parameters corresponding to each sampling time can be defined as: In the formula, Indicates the sampling time The set of beam pointing parameters; J represents the total number of beams in the current scan frame. Simultaneously, the robot body at the sampling time... The initial pose parameters can be expressed as: In the formula, Indicates the sampling time The initial pose parameters; This represents the initial position vector of the robot body in the navigation coordinate system; This represents the attitude rotation matrix at that moment.

[0046] In actual implementation, Can be determined by depth value The depth value is determined by combining the initial trajectory calculation results. For example, in near-bottom detection scenarios, the depth value is used to constrain the vertical position component, thereby improving the stability of pose initialization.

[0047] Step S230: Based on acceleration and angular velocity, pre-integration processing is performed on the motion state data between adjacent sampling times to obtain the pose increment between adjacent sampling times. Here, pose increment refers to the combined amount of attitude change, velocity change, and position change of the robot body between adjacent sampling times; pre-integration processing refers to the method of integrating and accumulating the linear acceleration and angular velocity output by the inertial measurement unit within a discrete time interval to obtain the relative motion quantity related to the initial reference time.

[0048] For adjacent sampling times and The time interval is defined as: In the formula, This represents the sampling time interval for the nth time interval. In one embodiment, the attitude increment between adjacent sampling moments can be expressed as: In the formula, Indicates from the sampling time up to the sampling time The attitude increment; Indicates time angular velocity vector; This represents the zero bias vector of the gyroscope; This represents the time interval of the corresponding infinitesimal integral.

[0049] The velocity increment can be expressed as: In the formula, Indicates from the sampling time up to the sampling time The speed increment; Indicates time The linear acceleration vector; represents the zero bias vector of the accelerometer; g represents the gravitational acceleration vector.

[0050] The position increment can be expressed as: In the formula, Indicates from the sampling time up to the sampling time Position increment; Let represent the velocity change term corresponding to the k-th discrete time step during the integration process. Therefore, the pose increment between adjacent sampling times can be defined as: In the formula, Indicates from the sampling time up to the sampling time The set of pose increments.

[0051] In this embodiment, the pre-integration processing can also be combined with depth constraints, sensor bias compensation, and filtering models to suppress integration drift. For example, the depth data measured by the depth gauge can be... As a vertical position constraint, it reduces the cumulative error of inertial integrals in the vertical channel.

[0052] Step S240: Based on the beam pointing parameters, initial pose parameters, and pose increments corresponding to each sampling time, beam pointing compensation and pose registration are performed on the acoustic scanning data to obtain registered acoustic scanning data. The registered acoustic scanning data refers to the set of acoustic sampling points that have been corrected in beam direction and uniformly transformed to a preset reference coordinate system.

[0053] For sampling time The slant range measurement value corresponding to the j-th beam The coordinates of its seabed measuring point in the navigation coordinate system can be expressed as: In the formula, Indicates the sampling time The coordinates of the seabed measuring point corresponding to the j-th beam; Indicates the robot body at the sampling time The position vector; This represents the ranging value of the j-th beam; This represents the actual direction vector of the beam in the navigation coordinate system.

[0054] Furthermore, the robot body at the sampling time position vector It can be obtained recursively from the initial pose parameters and pose increment: In the formula, Represents the initial position vector at the start of sampling; Indicates from time At the time The position increment.

[0055] Therefore, the set of all sampled points in the unified reference coordinate system can be represented as: In the formula, Q represents the acoustic point set after beam pointing compensation and pose registration. In the specific implementation, the industrial control computing module first compensates the acoustic beam direction of each sampling point according to the beam pointing parameters, and then, combined with the initial pose parameters and pose increments, transforms the acoustic scanning points acquired at different sampling times point by point into a unified reference coordinate system. Through this method, seabed measuring points acquired at different sampling times can form a continuous and consistent geometric structure in a unified space.

[0056] Step S250: Based on the registered acoustic scanning data and the pose increment between adjacent sampling times, motion distortion correction is performed on the registered acoustic scanning data to obtain the corrected acoustic scanning data.

[0057] Among them, the propulsion disturbance parameter refers to the set of parameters used to characterize the periodic attitude deviation, velocity pulsation and spatial position disturbance caused by fish body undulation, tail fin oscillation, asymmetric steering oscillation and near-bottom flow field coupling during the biomimetic oscillating propulsion process of the robot body; the position compensation amount refers to the correction amount applied to the spatial translation error of each sampling point; the direction compensation amount refers to the correction amount applied to the corresponding observation direction or local normal deviation of each sampling point.

[0058] Furthermore, in one embodiment, step S250 may further include: Step S251: Based on the pose increment and depth between adjacent sampling times, determine the propulsion perturbation parameters of the robot body during the acoustic scanning process. In this application, the propulsion perturbation parameters are preferably defined as: In the formula, Indicates the sampling time The propulsion disturbance parameters are as follows; Indicates the positional disturbance component; Indicates the velocity disturbance component; Indicates the attitude disturbance component; The propulsion periodic characteristic components are represented. The aforementioned propulsion perturbation parameters are used to describe the structured motion perturbation characteristics of the robot body during acoustic scanning under biomimetic propulsion conditions, caused by oscillating propulsion, near-bottom flow field disturbance, and attitude coupling, so as to compensate and constrain it in subsequent sonar beam compensation, motion distortion correction, and pose optimization processes.

[0059] In a preferred embodiment, the propulsion perturbation parameters can be constructed using a biomimetic propulsion mechanism. According to Lighthill's slender body theory, a slender fish body generates propulsion through body wave propagation, and this thrust satisfies the following relationship: In the formula, T represents the propulsion thrust; The fluid density is represented by U; the traveling wave velocity along the fish's body is represented by U. Indicates the amplitude of the oscillation; This represents the traveling wave wavelength. Because the thrust generated by biomimetic propulsion has periodic variation characteristics, it causes velocity pulsations, yaw disturbances, and micro-vibrations in the robot body during scanning. These disturbances are not completely random but are correlated with the oscillating propulsion parameters, propulsion period, and attitude response.

[0060] To further describe the midline undulation of the fish, the modified Barrett wave equation can be used: In the formula, The value represents the displacement of the fish's midline in the lateral direction; x represents the arc length coordinate or axial position coordinate along the head-to-tail direction of the fish. and represents the fish body wave envelope coefficient; k represents the traveling wave number, and ; t represents the volume wave angular frequency; t represents time. From the above relationship, it can be seen that the fish's body oscillation output is affected by the amplitude envelope, wavelength, and frequency, which in turn affects the robot's instantaneous propulsion state and attitude disturbance at each sampling moment.

[0061] In discrete control implementation, a discrete form can be used: In the formula, i represents the discrete time index within a single oscillation period, and M represents the number of discrete samples within one oscillation period; and This represents the discrete waveform coefficients. The robot's propulsion speed, yaw amplitude, and attitude disturbances are coupled with the yaw amplitude, wavelength, yaw frequency, and discrete control rhythm. Therefore, this application does not treat propulsion disturbances as random noise, but rather models them as structured disturbances related to the biomimetic yaw parameters. Based on this mechanism, the pose increment sequence at adjacent sampling times can be combined. Depth changes The impact of attitude change frequency and propulsion period on propulsion disturbance parameters Make an estimate.

[0062] Furthermore, for biomimetic thrusters employing CPG control, CPG state variables can be used to aid in describing the propulsion disturbance characteristics. In one embodiment, the aforementioned single-joint asymmetric oscillation control can be represented by the following CPG dynamics model: ; ; ; ; Where: b represents the swing bias state quantity, which is used to describe the DC offset of the current swing trajectory relative to the zero-position center line. This represents the first derivative of the oscillation bias state quantity b with respect to time, i.e., the rate of change of the bias. This represents the second derivative of the oscillating bias state quantity b with respect to time, which is the second-order dynamic response of the bias state. This represents the bias convergence coefficient, used to characterize how quickly the bias state variable b converges to the target bias variable b. The larger the value, the faster the dynamic response of the bias channel. B represents the target bias control quantity, which is the target value of the oscillation bias state quantity b or a high-level control input, used to adjust the degree of offset of the output oscillation center relative to zero. m represents the oscillation amplitude state quantity, used to describe the amplitude of the current oscillation output relative to the oscillation center. It represents the first derivative of the oscillation amplitude state quantity M with respect to time, i.e., the rate of change of amplitude. It represents the second derivative of the oscillation amplitude state quantity M with respect to time, that is, the second-order dynamic response quantity of the amplitude state. It represents the amplitude convergence coefficient, which is used to characterize how quickly the oscillation amplitude state variable M converges to the target amplitude M; The larger the value, the faster the dynamic response of the amplitude channel. M represents the target amplitude control quantity, which is the target value of the oscillation amplitude state quantity M or a high-level control input used to adjust the oscillation amplitude of the oscillation output. It represents the phase state quantity of the oscillation, used to characterize the phase position of the current periodic oscillation within one oscillation period. Represents the oscillating phase state quantity The first derivative with respect to time is the instantaneous rate of phase change. ω represents the phase reference angular frequency, used to determine the basic oscillation speed of the swing process, and its unit can be radians per second. R represents the asymmetric time ratio parameter, used to characterize the duration distribution relationship between the positive and negative half-cycles of the swing period, or between the left and right swing phases; in this application, R is used to adjust the time asymmetry of the swing to generate the asymmetric propulsion mode required for steering. Represents a symbolic function. Represents phase state quantity The value of the sine function. Represents phase state quantity The value of the cosine function. This represents the drive control output quantity, used to characterize the target output angle ultimately applied to the tail drive mechanism, servo, or oscillating joint. In a preferred embodiment of this application, It can be used as the target joint angle output by the controller. This represents the auxiliary state output quantity, which is related to... Orthogonal intermediate variables are used to characterize another output component of the oscillator in phase space; in some implementations, It can be used for phase tracking, state observation, waveform reconstruction, or control smoothing.

[0063] The sign function satisfies: In the formula, These are the input variables for the symbolic function. In the formula, the input variables are specifically... ,therefore This represents the result of the operation of taking the sign of the phase sine value.

[0064] This indicates the duration of the second phase within one oscillation cycle. Duration of Phase 1 The ratio of the two phases is used to describe the propulsion parameters. Since the oscillating propulsion during the turning process typically no longer satisfies the time symmetry of the first and second halves of the cycle, the lateral forces and yaw moments exerted on the aircraft by the two phases differ in duration and cumulative effect. Consequently, the perturbation contributions to the imaging carrier's attitude and scanning axis are no longer balanced. The time ratio R directly characterizes the duration distribution relationship between these two phases, thus reflecting the temporal structural asymmetry within the oscillation cycle and further characterizing the resulting uneven distribution of yaw disturbance strength. Based on this, the propulsion disturbance parameters of this application can further include the oscillation angle output by the CPG. Phase The periodic motion characteristics represented by the time ratio R and frequency ω are used in this application to improve the interpretability and compensation specificity of the disturbance model.

[0065] Appendix Figure 5 The model output results are shown under different CPG advanced control parameter inputs. The horizontal axis represents time, and the vertical axis represents the output angle. (See attached diagram.) Figure 5 It can be seen that when the control parameters When the output reaches a stable periodic oscillation within a short time, and the oscillation center is basically symmetrical about the zero position, the biomimetic robotic fish executes symmetrical servo oscillations, enabling relatively fast linear swimming; when the control parameters When the output swing angle deviates from the zero center and the periodic time structure shows obvious asymmetry, it indicates that the thruster output has bias and unbalanced time distribution characteristics. At this time, the bionic robotic fish can perform steering motion; when the control parameters When the output amplitude decreases and the frequency decreases, it indicates that the robotic fish can complete a relatively slow forward swimming motion under the corresponding propulsion state.

[0066] From the appendix Figure 5 Furthermore, it is known that different advanced control parameters directly affect the amplitude, frequency, center offset, and periodic symmetry of the output swing angle. These factors correspond to the robot's propulsion speed, yaw trend, attitude disturbance, and scanning stability changes during acoustic scanning. Therefore, by introducing CPG state variables and control parameters into the propulsion disturbance parameter estimation process, this application can more accurately distinguish the sources of disturbance in different motion modes such as straight-line movement, turning, and slow propulsion, and accordingly make targeted corrections to the sonar beam pointing compensation, point cloud distortion correction, and pose calculation processes.

[0067] Step S252: Based on the propulsion disturbance parameters, determine the position compensation amount and orientation compensation amount corresponding to each sampling point in the registered acoustic scanning data.

[0068] In one embodiment, the j-th sampling point is at time... The position compensation amount can be expressed as: In the formula, This indicates the location compensation amount of the sampling point; This represents the position compensation mapping function, whose inputs are propulsion disturbance parameters, beam direction, and ranging value. Correspondingly, the direction compensation amount can be expressed as: In the formula, This represents the directional compensation amount corresponding to the observation direction of the sampling point; This represents the orientation compensation mapping function. Those skilled in the art will understand that the position compensation primarily corrects spatial deviations at sampling points caused by platform translational disturbances and velocity fluctuations, while the orientation compensation primarily corrects observation orientation deviations at sampling points caused by tail yaw, yaw micro-vibrations, and roll / pitch disturbances.

[0069] Step S253: Based on the position compensation amount and direction compensation amount corresponding to each sampling point, the sampling points in the registered acoustic scanning data are compensated and corrected to obtain the corrected acoustic scanning data.

[0070] In one embodiment, the corrected sampling point coordinates can be expressed as: In the formula, Indicates the corrected coordinates of the sampling points; Indicates the coordinates of the original sampling points after registration; Indicates the position compensation amount; Indicates the directional compensation amount; This represents the corresponding beam ranging value. Finally, the corrected acoustic scan data can be expressed as: In the formula, This represents the acoustic scan point set after motion distortion correction.

[0071] In step S300, based on the corrected acoustic scanning data, motion state data, and seabed environment perception data, underwater localization and map building of the robot body are performed to obtain the pose estimation results of the robot body and the seabed topography map.

[0072] This step is used to fuse multi-source information such as acoustic, inertial, and optical data to obtain stable robot pose and terrain map results in complex seabed environments. In one embodiment, step S300 can be further subdivided into steps S310 to S350.

[0073] Step S310: Generate terrain point cloud data corresponding to the current moment based on the corrected acoustic scan data, and extract seabed terrain features from the terrain point cloud data.

[0074] Specifically, the industrial control computing module can convert the corrected multibeam sonar data into topographic point cloud data corresponding to the current moment. The extracted seabed topographic features can be local elevation change features, slope features, edge features, corner features, surface curvature features, or other geometric features suitable for underwater topographic matching, and this application embodiment does not limit this.

[0075] Step S320: Extract seabed environment graphic features based on seabed environment perception data.

[0076] In one embodiment, seabed environment perception data is output by an optical imaging module, from which an industrial control computing module can extract seabed environment graphic features. These features may include texture features, edge contour features, regional distribution features, target appearance features, or semantically relevant features. Optionally, image feature extraction networks, local descriptors, or traditional image processing algorithms can be used for extraction; this embodiment does not limit the specific methods used. By introducing seabed environment graphic features, an environmental perception dimension can be added beyond purely acoustic terrain features, thereby improving the robustness of subsequent pose estimation.

[0077] Step S330: Based on motion state data, seabed topographic features and seabed environment graphic features, the pose of the robot body at the current moment is estimated to obtain the initial pose estimation result of the robot body.

[0078] In one embodiment, step S330 may further include: based on motion state data, using a pre-established robot body dynamics model to perform state propagation on the robot body's pose at the current moment to obtain a prior pose result; constructing observation constraints based on seabed topographic features and seabed environment image features; and matching and correcting the prior pose result based on the observation constraints to obtain an initial pose estimation result. Here, the prior pose result is used to characterize the predicted pose of the robot body at the current moment, and the observation constraints are used to characterize the correction relationship between seabed topographic observation data and seabed environment image observation data and the predicted pose.

[0079] Specifically, prior pose results can be provided jointly by the inertial measurement unit, depth gauge, and dynamic model, while observation constraints can be constructed from terrain point cloud matching relationships, terrain feature correlation relationships, and image feature matching relationships. By combining dynamic model predictions with multi-source observation constraints, more stable initial pose estimation results can be obtained in complex underwater environments.

[0080] Step S340: Based on the initial pose estimation result, construct the pose graph of the robot body at different times, and optimize the pose graph by combining the pose constraints of the historical time to obtain the pose estimation result of the robot body.

[0081] In one embodiment, the pose states of the robot body at different times can be used as graph nodes, and the pose relationships between adjacent times, terrain matching constraints, image observation constraints, and optional loop closure constraints can be used as graph edges to construct a pose graph. Subsequently, the pose graph is optimized using graph optimization methods to eliminate accumulated errors and improve global consistency. Optionally, loop closure detection and relocalization mechanisms can be included as part of the backend optimization process to further correct accumulated drift using historical map information; this embodiment does not limit this approach.

[0082] Step S350: Based on the pose estimation results, the terrain point cloud data is stitched together to obtain the seabed terrain map.

[0083] Specifically, the industrial control computing module can transform topographic point cloud data acquired at different times into a unified global coordinate system and stitch them together to form a seabed topographic map. The seabed topographic map can be a dense point cloud map, a raster elevation map, a hybrid topological-geometric map, or a composite map with environmental semantic information; this application embodiment does not limit this. By utilizing optimized pose results for map construction, the global consistency and representational accuracy of the seabed topographic map can be improved.

[0084] Step S400 involves generating motion control commands for the robot body based on the seabed topography map, pose estimation results, and the current exploration task. This step generates motion control commands suitable for the biomimetic swing propulsion platform to execute, based on the perceived environment and current task requirements, to achieve autonomous path planning and motion control in the terrain exploration task. In one embodiment, step S400 can be further subdivided into steps S410 to S460.

[0085] Step S410: Obtain the dynamic parameters of the robot body and establish a dynamic model of the robot body based on the dynamic parameters.

[0086] In one embodiment, the dynamic parameters include the robot body's mass parameters, additional mass parameters, flexible torso motion parameters, and hydrodynamic parameters. The process of establishing a dynamic model of the robot body based on these dynamic parameters may include: determining the traveling wave motion parameters of the robot body's flexible torso based on Barrett's traveling wave equations, and establishing the centerline waveform of the robot body's flexible torso based on these parameters; determining the propulsion force of the robot body during its oscillating propulsion process based on the centerline waveform and Lighthill's slender body theory; determining the corresponding fluid resistance, lift, and torsional torque of the robot body based on its velocity, attitude, and hydrodynamic parameters during underwater motion; and establishing a dynamic model characterizing the relationship between the robot body's acceleration and forces using the Newton-Euler equations based on the propulsion force, fluid resistance, lift, torsional torque, mass parameters, and additional mass parameters.

[0087] In this embodiment, the kinematics and dynamics of the robotic fish's flexible torso can be described based on Lighthill's slender body theory. Specifically, the midline waveform of the fish can be described in the following form: In the formula, The value represents the lateral displacement of the robot's flexible torso centerline at position x and time T; x represents the longitudinal arc length coordinate or body axis coordinate along the head-to-tail direction of the robot body. This represents the amplitude coefficient of the first term of the midline waveform; The quadratic amplitude coefficient of the midline waveform is represented by ; k represents the volume wave number, which is used to characterize the density of spatial fluctuations. The value represents the body wave angular frequency, used to characterize the rate of change of the centerline oscillation over time; T represents time. The subscript "body" indicates that the displacement corresponds to the centerline of the robot's torso, and not to the displacement of the external environment or other components.

[0088] The relative velocities of the torso elements can be further derived from the midline waveform: In the formula, Indicates the positional parameters of the torso curve The velocity vector of the corresponding infinitesimal element relative to the surrounding fluid; This indicates that the velocity corresponds to the parameter on the torso. The identified micro-element position; The vector represents the translational velocity of the robot's center of mass C; C indicates that this velocity corresponds to the center of mass. This represents the angular velocity of the robot body around the z-axis, that is, the yaw angular velocity in a two-dimensional plane; Let C represent the position vector component from the centroid to the infinitesimal element in the horizontal direction, where C is the scalar coefficient of this horizontal distance. This is the horizontal unit vector of the robot's body coordinate system; This indicates the angle of inclination of the local tangent direction of the flexible torso relative to the body axis reference direction; Indicates the swing angle The first derivative with respect to time; The parameter representing the arc length of the integral along the midline of the fish's body; This represents the unit normal vector of the infinitesimal element. Therefore, it can be seen that the instantaneous relative velocity of each infinitesimal element of the torso is simultaneously affected by the translation of the center of mass, the yaw rotation of the body, and the changes in local yaw angles.

[0089] Furthermore, the force exerted by the fluid on the torso element can be expressed as: In the formula, Indicates the effect of fluid on position parameters The fluid force vector per unit length corresponding to the torso element; m represents the equivalent additional mass coefficient at the corresponding element; Represents the relative velocity of the infinitesimal element The component in the normal direction; This represents the derivative with respect to time. This expression shows that during the oscillation of a slender body, its normal acceleration will induce an additional inertial reaction in the surrounding fluid, thereby forming an instantaneous fluid force acting on the body's infinitesimal elements.

[0090] By dividing the entire fish volume, the overall hydrodynamic forces acting on the robot body can be obtained: In the formula, The total additional hydrodynamic force acting on the robot body is represented by L; L represents the effective length of the robot body's flexible torso. Represents the overall hydrodynamic force in the tangential unit vector. Components in direction; Represents the overall hydrodynamic force in the normal unit vector. The component in the direction. Therefore, the distributed fluid load induced by the torso swing can be equivalently summarized as the total resultant force acting on the robot body, providing a dynamic basis for subsequent propulsion performance analysis and control compensation.

[0091] At the same time, the drag, lift, and torsional torque experienced by the robot body during motion can be calculated by combining the fluid resistance equation: ; ; In the formula, This indicates the fluid resistance experienced by the robot body; This represents the fluid lift force acting on the robot body; This represents the fluid damping torque acting on the robot body; Indicates fluid density; Represents the velocity vector of the center of mass. The square of the modulus; S represents the reference force-bearing area of ​​the robot body in the fluid; Indicates the drag coefficient; Indicates the lift coefficient; Indicates angle of attack or sideslip angle, used to characterize the angular relationship between the robot's velocity direction and body axis direction; Indicates the yaw damping coefficient; Represents the yaw rate The sign function that takes the sign.

[0092] like Figure 6 As shown, based on Lighthill's slender body theory and pseudo-rigid body theory, the hydrodynamic force in the x-direction generated during the tail swing of the biomimetic underwater robot is analyzed. and hydrodynamics in the y direction Matlab simulation calculations were performed. Figure 6 The horizontal axis is unified to time t (s), and the vertical axis is respectively... and The upper figure shows the hydrodynamic force in the x-direction as a function of time, and the lower figure shows the hydrodynamic force in the y-direction as a function of time.

[0093] from Figure 6 It can be seen that, During the simulation time, the hydrodynamic force in the x-direction It exhibits obvious periodic oscillation characteristics, with its waveform generally varying around a small negative bias, completing approximately 9–10 oscillation cycles within the simulation range, corresponding to an oscillation frequency of approximately 4.5–5 Hz. In contrast, the hydrodynamics in the y-direction… It also exhibits stable periodic changes, but completes approximately 4 to 5 oscillation cycles within the same time frame, corresponding to an oscillation frequency of approximately 2 Hz to 2.5 Hz. As can be seen from the attached figure, The main oscillation frequency is approximately The result, which is twice the main oscillation frequency, is consistent with the propulsion mechanism of a slender body. That is, when the tail fin or tail flexible structure completes a left-right swing cycle, the hydrodynamic pulsation in the forward propulsion direction usually exhibits two characteristic fluctuations. Therefore, the pulsation frequency of hydrodynamics in the x-direction is higher than the lateral hydrodynamic frequency.

[0094] In addition, by Figure 6 It can also be seen that the hydrodynamic force in the x-direction The order of magnitude is approximately hydrodynamics in the y direction The order of magnitude is approximately The difference between the two is approximately two orders of magnitude. This indicates that, under the set parameters, the lateral hydrodynamic force generated during the tail sway of the bionic robot is significantly greater than the forward hydrodynamic force. This phenomenon suggests that while the flexible tail sway provides propulsion, it also inevitably generates a strong lateral force, which in turn affects the robot's directional stability, attitude maintenance capability, and the stable pointing of sensor payloads.

[0095] also, The waveform is generally smooth and approximates a low-frequency sinusoidal change, while This results in higher frequency pulsation characteristics and more pronounced asymmetric waveform details. This indicates that during biomimetic propulsion, the lateral force caused by the lateral oscillation more directly reflects the periodic oscillation of the tail itself, while the forward propulsion force is more of a net pulsating effect formed by the coupling of the periodic oscillation through nonlinear fluid action. Therefore, when designing a robot propulsion system, it is necessary not only to focus on the average forward propulsion capability but also to consider lateral force suppression, yaw compensation, and attitude stability control to avoid the adverse effects of large lateral periodic hydrodynamic forces on linear swimming and load stability.

[0096] Finally, a dynamic model of the robot body in two-dimensional planar motion is established using the Newton-Euler equations: ; ; ; In the formula, Indicates the mass of the robot itself; the subscript "b" uniquely represents the robot itself. Let x represent the additional mass term along the x-direction; where, This represents the component of the center of mass velocity along the body axis x-direction; This represents the additional mass term along the y-direction; where This represents the component of the center of mass velocity along the body axis in the y-direction; This represents the x-component of the center of mass velocity. The time derivative; This represents the y-component of the center of mass velocity. The time derivative; This represents the velocity component of the robot's center of mass along the longitudinal axis. This represents the velocity component of the robot's center of mass along the body axis in the transverse direction; This represents the moment of inertia of the robot body about the z-axis; This represents the additional moment of inertia term about the z-axis; This represents the resultant force on the robot body in the x-direction; This represents the resultant force on the robot body in the y-direction; This represents the resultant torque of the robot body about the z-axis. It should be understood that the above equation is only an exemplary implementation. The embodiments of this application do not limit the specific parameter form, solution method, and dimension selection of the dynamic model, as long as it can characterize the main dynamic behavior of the biomimetic underwater robot during the oscillating propulsion process.

[0097] Furthermore, in this embodiment, for ease of engineering implementation, the modified wave equation proposed by Barrett can be used to parameterize the fish body undulations, and the waveform coefficients can be adjusted according to the robot's body length, flexible torso segmentation method, actuator arrangement, and target propulsion mode. By introducing the above dynamic model, the subsequent control strategy training process can more closely resemble the motion characteristics of a real platform, improving the consistency between the simulation environment and the real operating environment.

[0098] Step S420: Based on the seabed topographic map, pose estimation results, and the current exploration mission, determine the state input of the robot body, the local navigation sub-target, and the control target.

[0099] Specifically, the state input can include two parts: robot body state and environment state. The robot body state can include position, attitude, linear velocity, angular velocity, depth, and relative position and orientation with respect to the current local sub-target. The environment state can include terrain gradients of the area ahead, obstacle distances, local terrain features, or other processed environmental perception information. Local navigation sub-targets can be obtained by decomposing the global detection task. For example, a region coverage scanning task can be decomposed into a series of continuous coverage waypoints, and a local fine detection task can be decomposed into sub-targets such as fixed-point approach, surrounding scanning, or directional docking. This application embodiment does not limit this. The control target can be comprehensively determined by indicators such as reachability, safety, energy consumption constraints, attitude stability, and detection coverage.

[0100] Step S430: Based on the dynamic model and the seabed topography map, construct a simulation environment corresponding to the target sea area.

[0101] In one embodiment, the simulation environment can be constructed based on the seabed topography map, obstacle distribution, terrain undulation characteristics, near-bottom current field characteristics, and robot dynamics of the target sea area. This environment is used to simulate the motion response of the robot body when performing exploration tasks in real sea areas. By embedding the dynamics model into the simulation environment, the simulated state transition process can more realistically reflect the propulsion behavior of the robot body under different control parameters, thereby providing a high-fidelity environment foundation for training reinforcement learning control strategies.

[0102] Step S440: Construct a control strategy model, which includes a policy network and a value assessment network.

[0103] In one embodiment, the control strategy model is a flexible actor-critic control strategy model. The action space of the control strategy model includes control parameters for adjusting the robot's oscillating propulsion. These control parameters include oscillation bias parameters, oscillation amplitude parameters, oscillation frequency parameters, and time ratio parameters. These control parameters can be understood as a set of parameters for high-level oscillation behavior adjustment, rather than high-dimensional execution signals directly output to the underlying servos or actuators. By adopting this action space design, complex navigation control requirements can be mapped to a set of oscillation parameters with explicit physical meaning, thereby reducing the difficulty of strategy learning and enhancing the adaptability of the control output to the biomimetic actuator.

[0104] In one embodiment, the yaw bias parameter can be used to adjust the center offset of the robot's left and right yaws to achieve yaw trend control; the yaw amplitude parameter can be used to adjust the yaw amplitude of the tail fin or flexible torso to affect the propulsion force; the yaw frequency parameter can be used to adjust the yaw period to affect the propulsion speed; and the time ratio parameter can be used to characterize the asymmetric yaw behavior during steering to enhance steering control and trajectory adjustment capabilities. It should be noted that the time ratio parameter can be defined as... ,in, This indicates the duration of the first phase within one oscillation cycle. R represents the duration of the second stage within the same oscillation cycle. When R changes, it indicates a change in the time allocation between the two oscillation stages, thus reflecting the asymmetry in the time structure within the oscillation cycle. Since the lateral forces and yaw moments exerted on the aircraft by the two stages during the turn differ in duration and cumulative effect, the time ratio parameter R can further characterize the uneven contribution of yaw disturbances caused by asymmetric oscillation.

[0105] Step S450: Based on the state input, local navigation sub-objective, and control objective, train the control strategy model in the simulation environment.

[0106] See appendix for details Figure 7In this embodiment, the control strategy model can be trained using a flexible actor-critic algorithm. The strategy network acts as the actor, sampling control actions based on the probability distribution of input and output actions in the current state; the value evaluation network acts as the critic, evaluating the long-term benefits of executing the corresponding control actions in the current state. Optionally, the reward function can comprehensively consider multiple factors such as rewards for moving towards the target, rewards for reaching the target, collision penalties, energy consumption penalties, and stability rewards, to guide the strategy network in learning a motion control strategy that meets the requirements of the marine topography exploration mission.

[0107] In one embodiment, the policy network outputs the current environment state. Action probability distribution Its optimization objective can be expressed as: In the formula, Representation strategy The objective function value during the training time domain; This represents the policy function to be optimized; T represents the termination time of the training rounds or decision sequence. Indicates the policy Induced state-action distribution Find the expected value; This represents the state-action access distribution determined by policy π; Indicates the state input at time T; Indicates the control action output at time T; Indicates the state Next action The instant rewards received; This represents the entropy temperature parameter, used to adjust the importance of the strategy entropy term; Indicates the state The entropy value of the policy distribution; This represents the set of parameters for the policy network.

[0108] ;in, Rewards are given for the progress of the biomimetic underwater robot toward its current local navigation target. The target achievement reward obtained by the biomimetic underwater robot when it reaches the current local navigation target; The collision penalty for a biomimetic underwater robot when it collides with another object; To control actions Amplitude-dependent energy consumption penalty; Stability rewards related to the smoothness of motion and posture of biomimetic underwater robots; These are the weighted coefficients for location progress rewards, target achievement rewards, collision penalties, energy consumption penalties, and stability rewards, respectively.

[0109] Furthermore, the objective of the valuation network can be expressed based on the soft Bellman equation as follows: ; in, , The discount factor is 0 < <1, used to measure the weight of future returns at the current moment; This represents the state transition probability distribution of the environment, that is, given the current state. and actions Next state under the condition The generation probability of ; where, Represents the state at time t Take action This is used to update the target Q value of the soft state-action value function; Indicates the agent's state Next action The instant rewards received; Indicates the execution of an action The state obtained after the transition; E represents the mathematical expectation; Indicates by parameters The target soft-state value function is characterized in the state The value at; The parameters represent the target soft-state value network; Indicates the state The next action is obtained by sampling the strategy. Indicates by parameters The stochastic policy represented in the state The probability of the following action; Indicates the policy network parameters; Indicates by parameters The target soft state-action value function is represented in the state-action pair. The value at; These represent the parameters of the target soft Q-network; This represents the temperature parameter, used to adjust the weight of the strategy entropy term in value assessment; In one embodiment, the offline training phase can be conducted on a large scale in a numerical simulation environment built based on a dynamic model, reducing the cost and risks associated with direct training in a real underwater environment. After training, the trained policy network can be deployed to the robot's industrial control computing module to achieve rapid forward inference output during actual operations. Optionally, after deployment, online fine-tuning, transfer learning, or adaptive parameter update mechanisms can be combined to mitigate the differences between the simulation environment and the real sea area, improving the practical applicability of the policy.

[0110] Step S460: Based on the state input, output motion control commands for the robot body using the trained control strategy model.

[0111] In one embodiment, the step includes: generating an action probability distribution using a policy network based on the state input, and sampling control actions based on the action probability distribution; evaluating the state input and control action input value evaluation network, and determining the control action output based on the evaluation result; determining adjustment parameters corresponding to the robot's propulsion state and / or posture based on the control parameters in the output control action; and generating motion control commands for the robot based on the adjustment parameters to control the robot to perform at least one of forward movement, turning, hovering posture adjustment, and anti-current posture adjustment.

[0112] Specifically, the adjustment parameters may include the servo target angle, swing phase, drive frequency, PWM control quantity, or other execution parameters suitable for driving the bionic actuator, which are not limited in this embodiment. By further mapping the high-level control parameters output by the control strategy model to low-level executable adjustment parameters, the high-level decision results obtained from reinforcement learning can be effectively converted into actual driving actions, realizing the connection between navigation decisions and bionic propulsion execution.

[0113] A time ratio parameter can be introduced into the traditional CPG model to characterize the asymmetric oscillation behavior of the robot body during turning. For a biomimetic robotic fish driven by a single joint oscillation, different motion modes such as rapid linear swimming, smooth turning, and slow forward swimming can be achieved by adjusting the oscillation bias parameter, oscillation amplitude parameter, oscillation frequency parameter, and time ratio parameter. Furthermore, the time ratio parameter R, together with the oscillation bias parameter, oscillation amplitude parameter, and oscillation frequency parameter, constitutes a set of periodic motion characteristic quantities for oscillation propulsion. Among them, the oscillation bias parameter is used to characterize the oscillation center position, the oscillation amplitude parameter is used to characterize the oscillation output intensity, the oscillation frequency parameter is used to characterize the oscillation rhythm, and the time ratio parameter R is used to characterize the distribution relationship of the duration of the two stages within the same oscillation cycle. Since the disturbance effects of the two oscillation stages on the body attitude and scanning view axis under the turning condition are not the same in terms of duration and cumulative effect, the introduction of R can further reflect the unbalanced yaw disturbance phenomenon caused by asymmetric oscillation. Therefore, the compensation for scanning distortion caused by biomimetic oscillating propulsion can be upgraded from a posteriori correction based on image results to mechanism-aware compensation that incorporates propulsion mechanism characteristics, thereby improving the interpretability of disturbance modeling and the specificity of compensation. Simulation verification shows that this parameterization method can achieve stable output in a short time and has good smoothness and controllability during parameter switching, which is beneficial to improving trajectory tracking capability and motion stability in complex seabed topographic environments.

[0114] In step S500, based on motion control commands, the propulsion state and / or attitude state of the robot body are adjusted to perform subsequent ocean topography exploration tasks.

[0115] Specifically, see attached Figure 7 In this embodiment, the motion control process can be completed collaboratively by the control layer and the execution layer. The control layer includes a SAC policy network, a SAC value network, and a target trajectory generator, while the execution layer includes a dynamic model, a motion attitude estimation module, and a motion attitude acquisition module. Two-way information exchange occurs between the external aquatic environment and the execution layer: on one hand, the robot moves in the external aquatic environment and is affected by fluid, terrain, and disturbances; on the other hand, the execution layer obtains the actual motion state information of the robot in the external aquatic environment through the motion attitude acquisition module and estimates the robot's current pose, velocity, and attitude changes through the motion attitude estimation module.

[0116] In one embodiment, the industrial control computing module can input the motion control commands generated in step S400 into the control layer. The target trajectory generator generates target trajectory information based on the current exploration task, seabed topography map, local navigation sub-targets, and motion attitude acquisition results, and inputs the target trajectory information into the SAC value network and SAC policy network. The SAC value network is used to evaluate the long-term benefits of the robot body executing candidate control actions in the current state, and the SAC policy network is used to output control actions based on the state input and value evaluation results. The control actions can be further mapped via a dynamic model into propulsion state adjustment amounts and / or attitude state adjustment amounts suitable for execution by the bionic actuator 3, and the bionic actuator 3 drives the robot body 1 to execute corresponding propulsion state and attitude state adjustments.

[0117] Furthermore, attached Figure 7 The dynamic model shown represents the motion response of the robot body in the underwater environment under control actions, enabling the control parameters output by the control layer to be converted into actual executable driving quantities based on the robot's dynamic characteristics. The motion attitude acquisition module collects sensor data such as the robot's velocity, attitude angle, angular velocity, depth, or other characteristics of its motion state. The motion attitude estimation module performs a fusion estimation of the robot's current motion state based on the collected data and feeds the estimation results back to the control layer for use by the target trajectory generator, SAC policy network, and SAC value network to update control decisions. Thus, a closed-loop control link is formed between the control layer output, the execution layer response, and the state feedback.

[0118] For example, when a local navigation sub-target requires the robot to quickly approach the target area, the target trajectory generator can generate a corresponding forward target trajectory, and the SAC policy network outputs forward control actions accordingly. When the local map shows obstacles or significant terrain undulations ahead, the target trajectory generator can generate obstacle avoidance or detour trajectories, and the SAC policy network outputs steering control actions. When fine-grained detection or near-bottom stable observation is required, a low-speed stable hovering trajectory can be generated, and hovering attitude adjustment control actions can be output. When the robot is disturbed by near-bottom currents, an anti-current holding trajectory can be generated, and anti-current attitude adjustment control actions can be output. Through the above control processes, the robot body 1 can dynamically adjust its propulsion and attitude states based on real-time detection results and motion state feedback, and continuously perform subsequent ocean topography exploration tasks.

[0119] Furthermore, in conjunction with the appendix Figure 7 The SAC policy optimization module shown can achieve target trajectory generation and path optimization through the RL optimization module during the training phase. The SAC policy network parameters can be updated according to the policy gradient to gradually improve the control policy's adaptability to complex seabed environments. (Appendix) Figure 7 The parameter update relationship shown Indicates policy network parameters Based on the performance objective function The gradient information is used for iterative optimization, where, This indicates the parameter update step size or learning rate. This represents the gradient of the objective function with respect to the policy. Through this optimization process, the SAC policy network can gradually learn a control policy that better meets the requirements of target trajectory tracking, environmental constraints, and motion stability.

[0120] In one embodiment, append Figure 7 The target trajectory generator in this application receives current state information from the motion attitude acquisition module and trajectory optimization results from the target trajectory generation path optimization (RL) module. This generates a target trajectory suitable for the current task scenario, which is then used as the basis for control layer decisions and input into the SAC value network and SAC policy network. Therefore, the motion control in this application does not simply issue commands based on fixed rules, but rather performs closed-loop adjustments based on a collaborative mechanism of target trajectory generation, value assessment, policy output, dynamic response, and attitude feedback.

[0121] In this embodiment, steps S100 to S500 are not absolutely isolated linear processes, but can be executed iteratively during actual operation. Specifically, the new propulsion and attitude states of the robot body after executing step S500 will affect the actual motion process in the external aquatic environment. These states are collected in real time by the motion attitude acquisition module, processed by the motion attitude estimation module, and fed back to the control layer and the aforementioned detection processing flow. This affects the acoustic scanning data, environmental perception data, pose estimation results, and map update results at the next moment, and further affects the subsequent target trajectory generation and control strategy output. Thus, this application constructs a dynamic closed-loop detection framework for biomimetic propulsion platforms, enabling the robot to continuously execute an autonomous operation process of detection—correction—localization and mapping—trajectory generation—control—re-detection in complex seabed environments.

[0122] In practical applications, the biomimetic underwater robot platform and its marine topography detection method provided in this application can be widely used in marine resource exploration, seabed engineering construction, marine scientific research, environmental monitoring of sensitive sea areas, and other scenarios requiring high-precision seabed topography perception and autonomous near-bottom detection. Compared with existing technologies, this application introduces the motion state data of the robot body to perform beam pointing compensation, pose registration, and motion distortion correction on the acoustic scanning data acquired by multibeam mapping sonar. This enables the acoustic scanning data to maintain high spatial consistency and mapping accuracy even under non-uniform oscillating propulsion and strong near-bottom disturbance conditions, thereby effectively mitigating the point cloud distortion problem caused by attitude fluctuations, propulsion disturbances, and environmental echo interference. By combining the corrected acoustic scanning data with seabed environmental perception data for underwater positioning and map construction of the robot body, it helps to enhance the robustness of pose estimation results and the global consistency of seabed topography maps. Furthermore, by using the seabed topography map and pose estimation results to generate motion control commands for the robot body, and combining them with the attached... Figure 7The target trajectory generator, SAC value network, SAC policy network, dynamic model, motion attitude acquisition module, and motion attitude estimation module shown enable the robot to dynamically adjust its propulsion state and / or attitude state based on the current detection task and real-time feedback results, thereby forming an autonomous closed-loop mechanism in which detection and perception, localization and mapping, trajectory planning, and motion control mutually feedback each other. Furthermore, this application establishes a dynamic model based on the robot's body dynamic parameters and constructs a simulation environment that matches the motion characteristics of the real platform. This allows the control strategy training process to reflect the dynamic characteristics of the biomimetic underwater robot during its oscillating propulsion, such as propulsion force, fluid resistance, lift, and torsional torque, thereby improving the simulation environment's ability to represent the real motion process. Simultaneously, the action space of the control strategy model is set as the control parameters for adjusting the robot's oscillating propulsion. This transforms the strategy network output from directly facing the low-level drive execution to facing the adjustment of high-level oscillating behavior. This helps reduce the learning difficulty caused by the high-dimensional nonlinear mapping between terrain detection and navigation actions and biomimetic propulsion control, improving the efficiency of control strategy training and its adaptability to the real platform. In particular, by setting the time ratio parameter, the asymmetric oscillating behavior during the robot's turning process can be characterized, thereby improving the turning control capability and trajectory tracking capability in complex terrain environments.

[0123] The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit the scope of protection of this application. For those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these should all fall within the scope of protection of this application.

Claims

1. A biomimetic underwater robot system, characterized in that, include: The robot itself; A watertight electronic compartment located inside the robot body; Bionic actuators connected to the robot body; Multibeam mapping sonar mounted on the robot body; The inertial measurement unit, depth gauge, optical imaging module, and industrial control computing module are installed on the robot body. in, Bionic actuators are used to drive the robot body to perform forward movement, turning, hovering posture adjustment, and anti-current posture adjustment. Multibeam mapping sonar is used to acquire acoustic scanning data of seabed topography; The inertial measurement unit, depth gauge, and optical imaging module are used to collect motion state data of the robot body and environmental perception data; The industrial control computing module is connected to the bionic actuator, multi-beam mapping sonar, inertial measurement unit, depth gauge, and optical imaging module, respectively, for the following purposes: Beam pointing compensation and motion distortion correction are performed on acoustic scanning data based on the motion state data of the robot body to obtain corrected terrain data. Based on the corrected terrain data, the robot's motion state data, and environmental perception data, simultaneous localization and map building are performed to obtain the robot's pose estimation results and the seabed terrain map. Based on the seabed topographic map, pose estimation results, and current exploration mission, drive control commands are generated to control the robot body to perform the ocean topographic exploration mission.

2. A biomimetic underwater robot control method, applied to the biomimetic underwater robot system of claim 1, characterized in that, include: The robot body is controlled to perform exploration tasks in the target sea area and acquire acoustic scanning data of the seabed topography, motion state data of the robot body, and seabed environment perception data. Among them, the acoustic scanning data is acquired through multibeam sonar, the motion state data is acquired through inertial measurement unit and depth gauge, and the seabed environment perception data is acquired through optical imaging module. Based on motion state data, beam pointing compensation, pose registration and motion distortion correction are performed on acoustic scanning data to obtain corrected acoustic scanning data. Based on the corrected acoustic scanning data, motion state data, and seabed environment perception data, the robot performs underwater localization and map building, and obtains the robot's pose estimation results and seabed topography map. Based on the seabed topographic map, pose estimation results, and current exploration mission, generate motion control commands for the robot body; Based on motion control commands, the propulsion state and / or attitude state of the robot body are adjusted to perform subsequent marine topography exploration tasks; wherein, the adjustment of the propulsion state and / or attitude state is achieved through biomimetic actuators.

3. The biomimetic underwater robot control method according to claim 2, characterized in that: The corrected acoustic scan data obtained include: Acquire the sampling time sequence corresponding to the acoustic scanning data and the motion state data corresponding to the sampling time sequence. The motion state data includes attitude angle, angular velocity, acceleration and depth data. Based on attitude angle, angular velocity and depth data, determine the beam pointing parameters and initial pose parameters corresponding to each sampling time. Based on acceleration and angular velocity, the motion state data between adjacent sampling times are pre-integrated to obtain the pose increment between adjacent sampling times. Based on the beam pointing parameters, initial pose parameters, and pose increments corresponding to each sampling time, beam pointing compensation and pose registration are performed on the acoustic scanning data to obtain the registered acoustic scanning data. Based on the registered acoustic scanning data and the pose increment between adjacent sampling times, motion distortion correction is performed on the registered acoustic scanning data to obtain the corrected acoustic scanning data.

4. The biomimetic underwater robot control method according to claim 3, characterized in that: Motion distortion correction is performed on the registered acoustic scan data, including: Based on the pose increment and depth between adjacent sampling times, the propulsion perturbation parameters of the robot body during the acoustic scanning process are determined; Based on the propulsion disturbance parameters, the position compensation and orientation compensation amounts corresponding to each sampling point in the registered acoustic scanning data are determined. Based on the position compensation and orientation compensation corresponding to each sampling point, each sampling point in the registered acoustic scanning data is compensated and corrected to obtain the corrected acoustic scanning data. Among them, the propulsion perturbation parameter is used to characterize the periodic attitude shift, velocity fluctuation and spatial position shift generated by the robot body during the biomimetic swing propulsion process.

5. The biomimetic underwater robot control method according to claim 3, characterized in that: The robot's pose estimation and seabed topography map are obtained, including: Based on the corrected acoustic scan data, generate the terrain point cloud data corresponding to the current moment, and extract the seabed terrain features from the terrain point cloud data; Extracting graphic features of the seabed environment based on seabed environment perception data; Based on motion state data, seabed topographic features, and seabed environment graphic features, the current pose of the robot body is estimated to obtain the initial pose estimation result of the robot body. Based on the initial pose estimation results, pose graphs of the robot body at different times are constructed, and the pose graphs are optimized by combining the pose constraints of historical times to obtain the pose estimation results of the robot body. The terrain point cloud data is stitched together based on the pose estimation results to obtain the seabed topographic map.

6. The biomimetic underwater robot control method according to claim 5, characterized in that: The initial pose estimation results of the robot body are obtained, including: Based on motion state data, the robot body's pose at the current moment is propagated using a pre-established robot body dynamics model to obtain the prior pose result. Observation constraints are constructed based on seabed topographic features and seabed environment image features; Based on observation constraints, the prior pose results are matched and corrected to obtain the initial pose estimation results. Among them, the prior pose result is used to characterize the predicted pose of the robot body at the current moment, and the observation constraints are used to characterize the correction relationship between the seabed topographic observation data and the seabed environment graphic observation data and the predicted pose.

7. The biomimetic underwater robot control method according to any one of claims 2 to 6, characterized in that: Generate motion control commands for the robot body, including: Obtain the dynamic parameters of the robot body and establish a dynamic model of the robot body based on the dynamic parameters; Based on the seabed topographic map, pose estimation results, and the current exploration mission, the state input, local navigation sub-targets, and control targets of the robot body are determined. Based on the dynamic model and the seabed topography map, a simulation environment corresponding to the target sea area is constructed; Construct a control strategy model, which includes a strategy network and a value assessment network; Based on state input, local navigation sub-objectives, and control objectives, a control strategy model is trained in a simulation environment. Based on the state input, the trained control strategy model is used to output motion control commands for the robot body. Among them, the control strategy model is the flexible actor-critic control strategy model; The action space of the control strategy model includes control parameters for adjusting the robot's oscillating propulsion. These control parameters include oscillation bias parameters, oscillation amplitude parameters, oscillation frequency parameters, and time ratio parameters.

8. The biomimetic underwater robot control method according to claim 7, characterized in that: The dynamic parameters include the robot's body mass parameters, additional mass parameters, flexible torso motion parameters, and fluid dynamic parameters; A dynamic model of the robot body is established based on dynamic parameters, including: Based on Barrett's traveling wave equation, the traveling wave motion parameters of the robot's flexible torso are determined, and the midline waveform of the robot's flexible torso is established based on the traveling wave motion parameters. Based on the midline waveform, the propulsion force of the robot body during the swing propulsion process is determined according to Lighthill's slender body theory; Based on the robot's speed, attitude, and fluid dynamics parameters during underwater movement, the corresponding fluid resistance, lift, and torsional torque of the robot are determined. Based on propulsion, fluid resistance, lift, and torsional torque, as well as mass parameters and additional mass parameters, a dynamic model characterizing the relationship between the robot's body motion acceleration and the forces is established using the Newton-Euler equations. Among them, the dynamic model is used to characterize the state transition relationship of the robot body under the action of motion control commands.

9. The biomimetic underwater robot control method according to claim 8, characterized in that: Based on the state input, the trained control strategy model outputs motion control commands for the robot body, including: Based on the state input, the policy network is used to generate the action probability distribution, and the control action is obtained by sampling based on the action probability distribution; The state input and control action input are evaluated by a value assessment network, and the control action output is determined based on the evaluation results. Based on the control parameters in the output control action, determine the adjustment parameters corresponding to the robot's propulsion state and / or posture; Based on the adjustment parameters, motion control commands are generated for the robot body to control the robot body to perform at least one of forward movement, turning, hovering posture adjustment and anti-current posture adjustment.

10. A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the method as described in any one of claims 2 to 9.