AUV visual dynamic docking control method based on BP neural network and reinforcement learning
Through the method of combining BP neural network with reinforcement learning, the BP neural network position controller and reinforcement learning heading controller are built, which solves the problems of visual delay and loss in AUV visual dynamic docking, and realizes high-precision and high-speed docking control.
Patent Information
- Application Number
- CN202510506142.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is difficult to effectively deal with the lag and uncertainty caused by visual delay and visual loss during the AUV visual dynamic docking process, affecting control accuracy and speed.
Using a method of combining BP neural network with reinforcement learning, a BP neural network position controller and reinforcement learning heading controller are built, and visual delay time and loss rate are used for training to achieve high-precision and high-speed control of docking AUVs.
In the case of visual delay and loss, high-precision and high-speed visual dynamic docking control of AUV is realized, reducing the impact of visual delay.
Smart Images

Figure CN120469467A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of visual dynamic docking control, and in particular to a method for visual dynamic docking control of an AUV based on BP neural network and reinforcement learning. Background Art
[0002] Autonomous underwater vehicles (AUVs) are key equipment in ocean exploration. With the increasing demand for deeper and broader ocean exploration, improving the AUV's exploration range has become a hot topic of research. Similar to aerial refueling tankers, AUVs can utilize dynamic docking technology underwater to recharge energy and transfer information. A docking AUV equipped with visual sensors will attempt to dock with a target AUV carrying a visual marker. At the end of the docking process, the docking AUV uses the relative position (relative to the target AUV) acquired through the visual sensors as docking control input. During this process, the docking AUV's visual sensors experience visual delays, and the target AUV's visual markers may be lost. These visual delays and loss introduce hysteresis and uncertainty to the docking AUV's control.
[0003] Improvements to PID control methods, such as segmented PID control, can reduce the effects of hysteresis and uncertainty. However, due to the complexity of the visual perception process, lag time can vary due to factors such as exposure time. This increases the difficulty of segmented PID control and makes it difficult to model the control process.
[0004] Utilizing existing PID control data to improve data-based control methods, thereby enhancing control adaptability, is a viable research direction. Neural networks have been used in the construction of control methods during docking, but simple neural networks perform poorly in terms of control accuracy and are unable to meet the high accuracy requirements for heading control during docking. Complex neural networks increase computational complexity and are unable to meet the timeliness requirements for position control during docking. Summary of the Invention
[0005] In view of this, an embodiment of the present application proposes a visual dynamic docking control method for AUV based on BP neural network and reinforcement learning. By using BP neural network and reinforcement learning, the visual delay time and visual loss rate are introduced into the controller, realizing high-precision and high-speed visual dynamic docking control under the influence of hysteresis and uncertainty.
[0006] In the first aspect, the embodiment of the present application proposes a visual dynamic docking control method for AUV based on BP neural network and reinforcement learning, which is suitable for docking AUV. The method comprises the following steps: constructing a BP neural network and training the BP neural network offline using an offline data set; constructing a BP neural network position controller based on the BP neural network completed by offline training, and training the BP neural network online using the segmented PID control output to obtain a trained BP neural network position controller; wherein the input of the BP neural network position controller is visual positioning information, visual loss rate and visual delay, and the output is the propulsion voltage of the main thruster of the docking AUV; constructing a BP neural network position controller based on the action network. A reinforcement learning network consisting of a position network and an evaluation network is constructed, and the evaluation network in the reinforcement learning network is trained offline using an offline dataset. A reinforcement learning heading controller is constructed based on the offline-trained reinforcement learning network, and the action network in the reinforcement learning network is trained online using the action output to obtain a trained reinforcement learning heading controller. The input of the reinforcement learning heading controller is the target heading, visual loss rate and visual delay, and the output is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV. The trained BP neural network position controller and the trained reinforcement learning heading controller are deployed on the docking AUV to realize the control of the AUV's visual dynamic docking.
[0007] In some optional embodiments, a BP neural network is constructed, and the BP neural network is trained offline using an offline data set, including: designing a network architecture of the BP neural network, including an input layer and an output layer, and adding a hidden layer between the input layer and the output layer, thereby constructing the BP neural network; obtaining a discrete data set for training the BP neural network, the discrete data set for training the BP neural network includes a large number of labeled training samples, the training samples are input into the BP neural network, the loss function is calculated based on the labels marked on the training samples and the output of the BP neural network on the training samples, and based on the calculated loss function, discrete training of the BP neural network is achieved using gradient descent.
[0008] In some optional embodiments, the sources of training samples in the discrete data set used to train the BP neural network are divided into two parts. The first part is the output values corresponding to different inputs recorded when the docking AUV uses a segmented PID control method to achieve stage position approach. The second part is the operating experience recorded under human docking control when using remote control to achieve ROV-like operations using images transmitted back by optical fiber.
[0009] In some optional embodiments, a BP neural network position controller is constructed based on a BP neural network that has been trained offline, and the BP neural network is trained online using a segmented PID control output, thereby obtaining a trained BP neural network position controller, including: using the BP neural network that has been trained offline as the processing core of the BP neural network position controller, and designing the input of the BP neural network position controller to be visual positioning information (V x ,V y ,V z ), visual loss rate σ and visual delay f t , the output is the propulsion voltage of the docking AUV main thruster, and the feedback is the updated relative distance between the docking AUV and the target AUV (x * ,y * ,z * ); Based on whether the output of the BP neural network position controller can complete the docking position control, when the time exceeds the task time, the BP neural network in the BP neural network position controller is trained online using the segmented PID control output, thereby obtaining a trained BP neural network position controller.
[0010] In some optional embodiments, a reinforcement learning network consisting of an action network and an evaluation network is constructed, and the evaluation network in the reinforcement learning network is trained offline using an offline data set, including: constructing an action network and an evaluation network separately, and connecting and combining the action network and the evaluation network to construct a reinforcement learning network; wherein the input of the action network is the target heading V ψ , visual loss rate σ and visual delay f t , the output is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV, the input of the evaluation network is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV, and the output is the evaluation reward value; a discrete data set for training the evaluation network is obtained, and the discrete data set for training the evaluation network contains a large number of training samples annotated with labels, and the labels are evaluation reward values; the training samples are input into the evaluation network, and the evaluation network is supervised for learning, thereby realizing offline training of the evaluation network.
[0011] In some optional embodiments, the training samples in the discrete data set for training the evaluation network are obtained by the following steps: designing PID controllers according to different visual loss rates and visual delays, recording the input, output, and evaluation reward value of each PID controller corresponding to the PID action as training samples, and forming a discrete data set for training the evaluation network; wherein the evaluation reward value is determined based on the degree of change in the heading state before and after the action is performed, if the relative heading angle decreases after the action is performed, the evaluation reward value is positive, and if the relative heading angle increases after the action is performed, the evaluation reward value is negative.
[0012] In some optional embodiments, a reinforcement learning heading controller is constructed based on a reinforcement learning network that has been trained offline, and the action network in the reinforcement learning network is trained online using the action output, thereby obtaining a trained reinforcement learning heading controller, including: using the action network in the reinforcement learning network that has been trained offline as the processing core of the reinforcement learning heading controller, and designing the input of the reinforcement learning heading controller to be the target heading V ψ , visual loss rate σ and visual delay f t The output is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV, and the feedback is the updated target heading The output of the action network in the reinforcement learning heading controller is sent to the evaluation network to calculate the evaluation reward value, and the calculated evaluation reward value is used to perform online training of the action network using the gradient descent method to obtain a trained reinforcement learning heading controller.
[0013] The embodiment of the present application proposes a method for visual dynamic docking control of an AUV based on BP neural network and reinforcement learning, and designs, constructs and trains a BP neural network position controller and a reinforcement learning heading controller to perform visual dynamic docking control of the docking AUV. The processing core of the BP neural network position controller is the BP neural network. The BP neural network position controller relies on the powerful adaptive ability of the BP neural network to well combine segmented PID control and manual operation experience to achieve position control, thereby reducing the impact of visual delay on the visual dynamic docking of the docking AUV. The processing core of the reinforcement learning heading controller is the action network in the reinforcement learning network. The reinforcement learning network, with its powerful self-learning ability, can continue to complete heading control when visual loss occurs and the heading control target is missing. In summary, the present application uses BP neural network and reinforcement learning to introduce visual delay time and visual loss rate into the controller, thereby achieving high-precision and high-speed visual dynamic docking control under the influence of hysteresis and uncertainty.
[0014] On the second aspect, an embodiment of the present application proposes an AUV visual dynamic docking control system based on BP neural network and reinforcement learning, the system comprising: a BP neural network construction and offline training module, for constructing a BP neural network, and using an offline data set to perform offline training on the BP neural network; a BP neural network position controller construction and online training module, for constructing a BP neural network position controller based on the BP neural network completed offline training, and using the segmented PID control output to perform online training on the BP neural network, thereby obtaining a trained BP neural network position controller, wherein the input of the BP neural network position controller is visual positioning information, visual loss rate and visual delay, and the output is the propulsion voltage of the main thruster of the docking AUV; a reinforcement learning network construction and offline training module, for constructing a reinforcement learning network consisting of an action network and an evaluation network, and using an offline data set to perform online training on the evaluation network in the reinforcement learning network. Offline training; a reinforcement learning heading controller construction and online training module, which is used to construct a reinforcement learning heading controller based on the reinforcement learning network completed by offline training, and use the action output to perform online training on the action network in the reinforcement learning network, thereby obtaining a trained reinforcement learning heading controller, wherein the input of the reinforcement learning heading controller is the target heading, visual loss rate and visual delay, and the output is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV; a deployment module, which is used to deploy the trained BP neural network position controller and the trained reinforcement learning heading controller to the docked AUV; an execution module, which is used to input the current visual positioning information, visual loss rate and visual delay of the docked AUV into the deployed BP neural network position controller, and input the current target heading, visual loss rate and visual delay of the docked AUV into the deployed reinforcement learning heading controller, so as to realize the control of the AUV visual dynamic docking.
[0015] In a third aspect, an embodiment of the present application proposes an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for visual dynamic docking of an AUV based on BP neural network and reinforcement learning as described in the first aspect above.
[0016] In a fourth aspect, an embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement an AUV visual dynamic docking control method based on BP neural network and reinforcement learning as described in the first aspect above.
[0017] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the following is a brief introduction to the drawings required for use in the embodiments of the present application or the description of the related technologies. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 This is a flowchart of an AUV visual dynamic docking control method based on BP neural network and reinforcement learning provided by an embodiment of the present application;
[0020] Figure 2 This is a schematic diagram of the structure of a BP neural network position controller provided by one embodiment of the present application;
[0021] Figure 3 is a schematic diagram of a process for acquiring a discrete data set for training a BP neural network provided by an embodiment of the present application;
[0022] Figure 4 This is a schematic diagram of offline training of a BP neural network provided by an embodiment of the present application;
[0023] Figure 5 This is a schematic diagram of online training of a BP neural network provided by an embodiment of the present application;
[0024] Figure 6 This is a schematic diagram of the structure of a reinforcement learning heading controller provided by an embodiment of the present application;
[0025] Figure 7 This is a schematic diagram of the principle of training a reinforcement learning network provided by an embodiment of the present application;
[0026] Figure 8 is a schematic diagram of a process for obtaining a discrete data set for training an evaluation network provided by an embodiment of the present application;
[0027] Figure 9 is a schematic diagram of offline training of an evaluation network provided by an embodiment of the present application;
[0028] Figure 10 is a schematic diagram of online training of an action network provided by an embodiment of the present application;
[0029] Figure 11This is a structural diagram of an AUV visual dynamic docking control system based on BP neural network and reinforcement learning provided by another embodiment of the present application;
[0030] Figure 12 This is a structural diagram of an electronic device provided by another embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the various embodiments of the present application, many technical details are proposed to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can be implemented. The division of the following embodiments is only for the convenience of description and should not constitute any limitation on the specific implementation of the present application. The various embodiments can be combined with each other and referenced to each other under the premise of no contradiction.
[0032] An embodiment of the present application proposes an AUV visual dynamic docking control method based on BP neural network and reinforcement learning, which is suitable for docking AUV and is applied to electronic equipment, wherein the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described using the server as an example. The implementation details of the AUV visual dynamic docking control method based on BP neural network and reinforcement learning proposed in this embodiment are specifically described below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this solution.
[0033] The specific process of the AUV visual dynamic docking control method based on BP neural network and reinforcement learning proposed in this embodiment can be as follows: Figure 1 Shown, including:
[0034] Step 11: construct a BP neural network and use the offline data set to perform offline training on the BP neural network.
[0035] Step 12: construct a BP neural network position controller based on the offline trained BP neural network, and use the segmented PID control output to train the BP neural network online, thereby obtaining a trained BP neural network position controller.
[0036] In the specific implementation, the visual dynamic docking control of the docking AUV is divided into two parts. The first part is the position control based on distance, and the second part is the attitude control (heading control) based on heading. The BP neural network has high nonlinear fitting ability and data rule integration ability, and low computational complexity. It can be used for position control with high real-time requirements and low precision requirements in visual docking. Based on this, this embodiment designs a BP neural network position controller to realize the position control in the visual dynamic docking control of the docking AUV. This first requires constructing a BP neural network and using an offline data set to train the BP neural network offline. After obtaining the BP neural network that has been trained offline, a BP neural network position controller can be constructed based on the BP neural network that has been trained offline, and the BP neural network can be trained online using the segmented PID control output to obtain a trained BP neural network position controller. The input of the BP neural network position controller is visual positioning information, visual loss rate and visual delay, and the output is the propulsion voltage of the main thruster of the docking AUV.
[0037] In one example, the structure of the BP neural network position controller can be as follows: Figure 2 shown.
[0038] In one example, the training of BP neural network includes two parts: offline training and online training. Before discrete training, the server needs to obtain a discrete data set for training BP neural network. The process of obtaining a discrete data set for training BP neural network can be as follows: Figure 3 As shown in Figure 1, the sources of training samples in the discrete data set used to train the BP neural network are divided into two parts. The first part is the output values corresponding to different inputs recorded when the docking AUV uses a segmented PID control method to achieve stage position approach. The second part is the operating experience recorded under human docking control when using remote control to achieve ROV-like operations using images transmitted back by optical fiber.
[0039] In one example, when constructing a BP neural network, the server needs to design a network architecture of the BP neural network, including an input layer and an output layer, and add a hidden layer between the input layer and the output layer to construct the BP neural network. The process of discrete training of the BP neural network can be as follows: Figure 4 As shown, the server needs to obtain a discrete data set for training the BP neural network. The discrete data set for training the BP neural network contains a large number of labeled training samples. The training samples are then input into the BP neural network. The loss function is calculated based on the labels marked on the training samples and the output of the BP neural network for the training samples. Based on the calculated loss function, the discrete training of the BP neural network is implemented using gradient descent.
[0040] In one example, when the server constructs a BP neural network position controller, it needs to use the BP neural network trained offline as the processing core of the BP neural network position controller and design the input of the BP neural network position controller as the visual positioning information (V x ,V y ,V z ), visual loss rate σ and visual delay f t , the output is the propulsion voltage of the docking AUV main thruster, and the feedback is the updated relative distance between the docking AUV and the target AUV (x * ,y * ,z * ).
[0041] In one example, the process of online training of the BP neural network in the BP neural network position controller can be as follows: Figure 5 As shown in the figure, the server uses the output of the BP neural network position controller to complete the docking position control as the judgment basis. When the time exceeds the task time, the BP neural network in the BP neural network position controller is trained online using the segmented PID control output to obtain a trained BP neural network position controller.
[0042] Step 13: construct a reinforcement learning network consisting of an action network and an evaluation network, and use the offline dataset to perform offline training on the evaluation network in the reinforcement learning network.
[0043] Step 14: construct a reinforcement learning heading controller based on the reinforcement learning network that has completed offline training, and use the action output to perform online training on the action network in the reinforcement learning network, thereby obtaining a trained reinforcement learning heading controller.
[0044] In the specific implementation, unlike the BP neural network, the reinforcement learning network has a stronger learning ability and the computational complexity is also improved. It is very suitable for use in the visual dynamic docking of the AUV, and for the heading control with high precision requirements and relatively low real-time performance. Based on this, the present embodiment designs a reinforcement learning heading controller to realize the heading control in the visual dynamic docking of the AUV. This first requires the construction of a reinforcement learning network consisting of an action network and an evaluation network, and the offline data set is used to perform offline training on the evaluation network in the reinforcement learning network. Subsequently, a reinforcement learning heading controller is constructed based on the reinforcement learning network completed by offline training, and the action network in the reinforcement learning network is trained online using the action output, thereby obtaining a trained reinforcement learning heading controller. The inputs of the reinforcement learning heading controller are the target heading, visual loss rate and visual delay in order, and the output is the propulsion voltage of the thrusters of the full-drive vehicle docking the AUV.
[0045] In one example, the reinforcement learning heading controller can be structured as follows Figure 6 As shown, the principle of training the reinforcement learning network is as follows Figure 7 As shown in Figure 2, similar to the training of BP neural network, it also includes two parts: offline training and online training. However, discrete training targets the evaluation network in the reinforcement learning network, while online training targets the action network in the reinforcement learning network. Before discrete training, the server needs to obtain discrete data for training the evaluation network. The process of obtaining the discrete data set for training the evaluation network can be as follows: Figure 8 As shown in the figure, the server needs to design PID controllers based on different visual loss rates and visual delays, and record the input, output, and evaluation reward values of each PID controller as training samples. These are then used to form a discrete data set for training the evaluation network. The evaluation reward value is determined based on the degree of change in the heading state before and after the action is performed. If the relative heading angle decreases after the action is performed, the evaluation reward value is positive, and if the relative heading angle increases after the action is performed, the evaluation reward value is negative.
[0046] In one example, when the server constructs a reinforcement learning network, it needs to first construct an action network and an evaluation network separately, and then connect and combine the action network and the evaluation network to construct a reinforcement learning network. The input of the action network is the target heading V ψ , visual loss rate σ and visual delay f t , the output is the propulsion voltage of the propeller of the full-drive vehicle docking with the AUV, the input of the evaluation network is the propulsion voltage of the propeller of the full-drive vehicle docking with the AUV, and the output is the evaluation reward value.
[0047] In one example, the offline training process of the evaluation network can be as follows Figure 9 As shown, the server obtains a discrete dataset for training the evaluation network. This dataset contains a large number of labeled training samples, where the labels are evaluation reward values. The training samples are then fed into the evaluation network for supervised learning, enabling offline training of the evaluation network.
[0048] In one example, after obtaining the offline trained reinforcement learning network, the server needs to use the action network in the offline trained reinforcement learning network as the processing core of the reinforcement learning heading controller and design the input of the reinforcement learning heading controller as the target heading V ψ , visual loss rate σ and visual delay f t The output is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV, and the feedback is the updated target heading Thus, a reinforcement learning heading controller is constructed.
[0049] In one example, the process of online training of the action network can be as follows Figure 10 As shown, the server needs to send the output of the action network in the reinforcement learning heading controller into the evaluation network to calculate the evaluation reward value, and use the calculated evaluation reward value to perform online training on the action network using the gradient descent method, thereby obtaining a trained reinforcement learning heading controller.
[0050] In step 15, the trained BP neural network position controller and the trained reinforcement learning heading controller are deployed on the docking AUV to realize the control of the AUV visual dynamic docking.
[0051] In the specific implementation, after the server completes the construction and training of the BP neural network position controller and the reinforcement learning heading controller, the trained BP neural network position controller and the trained reinforcement learning heading controller can be deployed to the docked AUV. The docked AUV inputs the current visual positioning information, visual loss rate and visual delay into the deployed BP neural network position controller, and inputs the current target heading, visual loss rate and visual delay into the deployed reinforcement learning heading controller, thereby realizing the control of the visual dynamic docking of the docked AUV itself.
[0052] This embodiment designs, constructs and trains a BP neural network position controller and a reinforcement learning heading controller to perform visual dynamic docking control for docking AUVs. The processing core of the BP neural network position controller is the BP neural network. The BP neural network position controller relies on the powerful adaptive ability of the BP neural network and can well combine segmented PID control and manual operation experience to achieve position control, reducing the impact of visual delay on the visual dynamic docking of the docking AUV. The processing core of the reinforcement learning heading controller is the action network in the reinforcement learning network. The reinforcement learning network, with its powerful self-learning ability, can continue to complete heading control when visual loss occurs and the heading control target is missing. In summary, this embodiment uses BP neural network and reinforcement learning to introduce visual delay time and visual loss rate into the controller, achieving high-precision and high-speed visual dynamic docking control under the influence of hysteresis and uncertainty.
[0053] The steps of the various methods described above are divided for clarity of description only. During implementation, they can be combined into a single step, or some steps can be split and decomposed into multiple steps. As long as they contain the same logical relationships, they are all within the scope of protection of this application. Furthermore, minor modifications or design changes to the algorithms or processes that do not change the core design of the algorithms or processes are also within the scope of protection of this application.
[0054] Another embodiment of the present application proposes an AUV visual dynamic docking control system based on BP neural network and reinforcement learning, which is suitable for docking AUV. The following is a specific description of the implementation details of the AUV visual dynamic docking control system based on BP neural network and reinforcement learning proposed in this embodiment. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this example.
[0055] Figure 11 This is a structural diagram of an AUV visual dynamic docking control system based on BP neural network and reinforcement learning proposed in this embodiment. The system includes: a BP neural network construction and offline training module 21, a BP neural network position controller construction and online training module 22, a reinforcement learning network construction and offline training module 23, a reinforcement learning heading controller construction and online training module 24, a deployment module 25 and an execution module 26.
[0056] The BP neural network construction and offline training module 21 is used to construct the BP neural network and perform offline training on the BP neural network using an offline data set.
[0057] The BP neural network position controller construction and online training module 22 is used to construct a BP neural network position controller based on the BP neural network completed offline training, and use the segmented PID control output to train the BP neural network online, thereby obtaining a trained BP neural network position controller, wherein the input of the BP neural network position controller is visual positioning information, visual loss rate and visual delay, and the output is the propulsion voltage of the main thruster of the docked AUV.
[0058] The reinforcement learning network construction and offline training module 23 is used to construct a reinforcement learning network consisting of an action network and an evaluation network, and use an offline data set to perform offline training on the evaluation network in the reinforcement learning network.
[0059] The reinforcement learning heading controller construction and online training module 24 is used to construct a reinforcement learning heading controller based on the reinforcement learning network completed by offline training, and use the action output to perform online training on the action network in the reinforcement learning network, thereby obtaining a trained reinforcement learning heading controller, wherein the input of the reinforcement learning heading controller is the target heading, visual loss rate and visual delay, and the output is the propulsion voltage of the thruster of the full-drive vehicle docked with the AUV.
[0060] The deployment module 25 is used to deploy the trained BP neural network position controller and the trained reinforcement learning heading controller to the docked AUV.
[0061] Execution module 26 is used to input the current visual positioning information, visual loss rate and visual delay of the docked AUV into the deployed BP neural network position controller, and input the current target heading, visual loss rate and visual delay of the docked AUV into the deployed reinforcement learning heading controller to achieve control of the AUV's visual dynamic docking.
[0062] It is worth mentioning that all modules and modules involved in this embodiment are logical modules. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovation of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application. However, this does not mean that other units do not exist in this embodiment.
[0063] It is not difficult to find that this embodiment is a system embodiment corresponding to the above-mentioned method embodiment. This embodiment can be implemented in conjunction with the above-mentioned method embodiment. The relevant technical details and technical effects mentioned in the above-mentioned method embodiment are still valid in this embodiment. In order to reduce repetition, they will not be repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above-mentioned method embodiment.
[0064] Accordingly, another embodiment of the present application provides an electronic device, the specific structure of which can be as follows: Figure 12 As shown, it includes: at least one processor 31; and a memory 32 communicatively connected to the at least one processor 31; wherein the memory 32 stores instructions that can be executed by the at least one processor 31, and the instructions are executed by the at least one processor 31 so that the at least one processor 31 can execute a method for visual dynamic docking control of an AUV based on BP neural network and reinforcement learning as described in the above method embodiment.
[0065] The memory and processor can be connected using a bus. The bus can include any number of interconnected buses and bridges, connecting various circuits within one or more processors and the memory. The bus can also connect various other circuits, such as peripherals, voltage regulators, and power management circuits. These are well known in the art and will not be described further herein. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium.
[0066] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.
[0067] Correspondingly, another embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement an AUV visual dynamic docking control method based on BP neural network and reinforcement learning as described in the above method embodiment.
[0068] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (such as a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc., various media that can store program code.
[0069] It will be understood by those skilled in the art that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.
Claims
1. A visual dynamic docking control method for AUV based on BP neural network and reinforcement learning, suitable for docking AUV, characterized by: The method comprises: Construct a BP neural network and use the offline data set to train the BP neural network offline; A BP neural network position controller is constructed based on the offline trained BP neural network, and the BP neural network is trained online using the segmented PID control output to obtain a trained BP neural network position controller. The input of the BP neural network position controller is visual positioning information, visual loss rate, and visual delay, and the output is the propulsion voltage of the main thruster of the docking AUV. Construct a reinforcement learning network consisting of an action network and an evaluation network, and use offline datasets to perform offline training on the evaluation network in the reinforcement learning network; A reinforcement learning heading controller is constructed based on the offline-trained reinforcement learning network. The action network in the reinforcement learning network is trained online using the action outputs, resulting in a trained reinforcement learning heading controller. The inputs of the reinforcement learning heading controller are the target heading, visual loss rate, and visual delay, and the output is the propulsion voltage of the thrusters of the fully driven vehicle docked with the AUV. The trained BP neural network position controller and the trained reinforcement learning heading controller are deployed on the docking AUV to realize the control of the AUV's visual dynamic docking.
2. The AUV visual dynamic docking control method based on BP neural network and reinforcement learning according to claim 1 is characterized in that: Construct a BP neural network and use the offline data set to train the BP neural network offline, including: Design the network architecture of BP neural network, including input layer and output layer, and add hidden layer between the input layer and output layer to construct BP neural network; Obtain a discrete data set for training the BP neural network. The discrete data set for training the BP neural network contains a large number of labeled training samples. The training samples are input into the BP neural network. The loss function is calculated based on the labels marked on the training samples and the output of the BP neural network for the training samples. Based on the calculated loss function, the discrete training of the BP neural network is realized using the gradient descent method.
3. The AUV visual dynamic docking control method based on BP neural network and reinforcement learning according to claim 2 is characterized in that: The sources of training samples in the discrete data set used to train the BP neural network are divided into two parts. The first part is the output values corresponding to different inputs recorded when the docking AUV uses a segmented PID control method to achieve stage position approach. The second part is the operating experience recorded under human docking control when using remote control to achieve ROV-like operations using images transmitted back by optical fiber.
4. The AUV visual dynamic docking control method based on BP neural network and reinforcement learning according to claim 3 is characterized in that: A BP neural network position controller is constructed based on the offline trained BP neural network, and the BP neural network is trained online using the segmented PID control output, thereby obtaining a trained BP neural network position controller, including: The BP neural network completed by offline training is used as the processing core of the BP neural network position controller, and the input of the BP neural network position controller is designed to be the visual positioning information (V x ,V y ,V z ), visual loss rate σ and visual delay f t , the output is the propulsion voltage of the docking AUV main thruster, and the feedback is the updated relative distance between the docking AUV and the target AUV (x * ,y * ,z * ); Whether the output of the BP neural network position controller can complete the docking position control is used as the judgment basis. When the time exceeds the task time, the BP neural network in the BP neural network position controller is trained online using the segmented PID control output to obtain a trained BP neural network position controller.
5. The AUV visual dynamic docking control method based on BP neural network and reinforcement learning according to claim 1 is characterized in that: Construct a reinforcement learning network consisting of an action network and an evaluation network, and use offline datasets to perform offline training on the evaluation network in the reinforcement learning network, including: The action network and evaluation network are constructed separately, and the action network and evaluation network are connected and combined to construct a reinforcement learning network; the input of the action network is the target heading V ψ , visual loss rate σ and visual delay f t , the output is the propulsion voltage of the propeller of the full-drive vehicle docking with the AUV, the input of the evaluation network is the propulsion voltage of the propeller of the full-drive vehicle docking with the AUV, and the output is the evaluation reward value; Obtain a discrete data set for training the evaluation network. The discrete data set for training the evaluation network contains a large number of labeled training samples, where the labels are evaluation reward values. The training samples are input into the evaluation network, and the evaluation network is supervised for learning, thereby realizing offline training of the evaluation network.
6. The AUV visual dynamic docking control method based on BP neural network and reinforcement learning according to claim 5 is characterized in that: The training samples in the discrete data set used to train the evaluation network are obtained by the following steps: PID controllers are designed according to different visual loss rates and visual delays. The input, output, and evaluation reward value of each PID controller are recorded as training samples, and a discrete data set is formed for training the evaluation network. The evaluation reward value is determined based on the degree of change in the heading state before and after the action is performed. If the relative heading angle decreases after the action is performed, the evaluation reward value is positive. If the relative heading angle increases after the action is performed, the evaluation reward value is negative.
7. The AUV visual dynamic docking control method based on BP neural network and reinforcement learning according to claim 6 is characterized in that: A reinforcement learning heading controller is constructed based on the offline trained reinforcement learning network. The action network in the reinforcement learning network is trained online using the action output, thereby obtaining a trained reinforcement learning heading controller, including: The action network in the offline trained reinforcement learning network is used as the processing core of the reinforcement learning heading controller, and the input of the reinforcement learning heading controller is designed to be the target heading V ψ , visual loss rate σ and visual delay f t The output is the propulsion voltage of the propeller of the full-drive vehicle docked with the AUV, and the feedback is the updated target heading The output of the action network in the reinforcement learning heading controller is sent to the evaluation network to calculate the evaluation reward value, and the calculated evaluation reward value is used to perform online training of the action network using the gradient descent method to obtain a trained reinforcement learning heading controller.
8. A visual dynamic docking control system for AUV based on BP neural network and reinforcement learning, suitable for docking AUV, characterized by: The system comprises: BP neural network construction and offline training module, used to build BP neural network and use offline data sets to train BP neural network offline; The BP neural network position controller construction and online training module is used to construct a BP neural network position controller based on the offline trained BP neural network, and use the segmented PID control output to train the BP neural network online to obtain a trained BP neural network position controller. The input of the BP neural network position controller is visual positioning information, visual loss rate and visual delay, and the output is the propulsion voltage of the main thruster of the docking AUV; The reinforcement learning network construction and offline training module is used to build a reinforcement learning network consisting of an action network and an evaluation network, and to perform offline training on the evaluation network in the reinforcement learning network using an offline dataset; The reinforcement learning heading controller construction and online training module is used to construct a reinforcement learning heading controller based on the reinforcement learning network completed offline training, and use the action output to perform online training on the action network in the reinforcement learning network to obtain a trained reinforcement learning heading controller. The input of the reinforcement learning heading controller is the target heading, visual loss rate and visual delay, and the output is the propulsion voltage of the full-drive vehicle docked with the AUV. Deployment module, used to deploy the trained BP neural network position controller and the trained reinforcement learning heading controller to the docked AUV; The execution module is used to input the current visual positioning information, visual loss rate and visual delay of the docked AUV into the deployed BP neural network position controller, and input the current target heading, visual loss rate and visual delay of the docked AUV into the deployed reinforcement learning heading controller to achieve control of the AUV's visual dynamic docking.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the AUV visual dynamic docking control method based on BP neural network and reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement an AUV visual dynamic docking control method based on BP neural network and reinforcement learning as described in any one of claims 1 to 7.