Autonomous underwater vehicle control method and system based on incremental regularization network
By adopting incremental regularization networks and LoRA fine-tuning technology in autonomous underwater vehicles, the problems of inefficient computing efficiency and insufficient real-time performance in complex underwater environments are solved, and a more flexible and safe path planning is achieved.
Patent Information
- Application Number
- CN202510258559.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-06
AI Technical Summary
When facing complex underwater environments, traditional path planning algorithms have low computational efficiency and insufficient real-time performance, making it difficult to meet practical application needs.
The autonomous underwater vehicle control method based on incremental regularization network is adopted, and the model complexity is controlled through deep Q network training and the L2 regularization method is introduced. Combined with incremental path dynamic decision-making mechanism and LoRA fine-tuning technology, model parameters are updated in real time to adapt to environmental changes.
It improves the flexibility and real-time path planning of autonomous underwater vehicles in complex environments, reduces the risk of collision with dynamic underwater targets, and enhances the safety and reliability of the aircraft.
Smart Images

Figure CN120066095A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control technology, and particularly to the control of autonomous underwater vehicles. Background Art
[0002] With the increasing exploration of the ocean and development of resources, autonomous underwater vehicles (AUVs) are being used more and more widely in the fields of ocean scientific research, environmental monitoring, ocean resource exploration, etc. However, the underwater environment is complex and changeable, including factors such as ocean currents, marine organisms, and obstacles, which pose great challenges to the autonomous navigation and path planning of AUVs.
[0003] Existing autonomous underwater vehicle control technologies mostly rely on pre-set path planning methods, usually making decisions based on static environment models. Although these methods are effective under certain conditions, in a dynamic underwater environment, due to the inability to adapt to changing situations in real time, they often result in inflexible path planning and increase the risk of the vehicle colliding with underwater obstacles. In addition, traditional path planning algorithms are prone to problems such as low computational efficiency and insufficient real-time performance when facing complex environments, and it is difficult to meet the actual application requirements.
[0004] In order to improve the navigation ability of autonomous underwater vehicles, in recent years, artificial intelligence technologies such as deep learning and reinforcement learning have gradually been introduced into the field of path planning. By learning historical data, these algorithms can improve the adaptability to complex environments. However, existing deep reinforcement learning models are prone to being affected by overfitting during the training process, resulting in insufficient generalization ability of the models in actual applications. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems of low computational efficiency and insufficient real-time performance that traditional path planning algorithms are prone to when facing complex environments, and to provide a control method and system for autonomous underwater vehicles based on an incremental regularization network.
[0006] The present invention is realized through the following technical solutions. On the one hand, the present invention provides a control method for an autonomous underwater vehicle based on an incremental regularization network, and the method includes: Step 1: Obtain the AUV historical underwater navigation database, train it using a deep Q-network, and control the model complexity by introducing the L2 regularization method; Step 2: The autonomous underwater vehicle starts from the starting point and performs initial path planning according to the initial environmental information collected by the sensor; Step 3: Use the incremental path dynamic decision-making mechanism to train the deep reinforcement learning model to obtain the incremental path decision network model. The autonomous underwater vehicle sensor collects the surrounding environment data in real time, and uses the LoRA fine-tuning technology to fine-tune the parameters of the incremental path decision network model batch by batch; the LoRA is a fine-tuning technology for the low-rank decomposition model; Step 4: Use the fine-tuned incremental path decision network model in Step 3 to update the path planning of the autonomous underwater vehicle; Step 5: The autonomous underwater vehicle checks whether the return constraints are satisfied. If the return constraints are not satisfied, it returns to Step 3, continues underwater navigation, and uses the collected data for model fine-tuning and path update; if the return constraints are satisfied, it executes Step 6; Step 6: The autonomous underwater vehicle calls the incremental path decision network to generate a path planning scheme to return from the current position to the starting point, and returns according to the scheme; after the autonomous underwater vehicle arrives at the starting point, it uploads the underwater navigation data of this time to the historical underwater navigation database in the data center.
[0007] Further, in Step 1, the specific method of controlling the model complexity by introducing the L2 regularization method is to make the model smoother by adding the sum of the squares of the parameters to the loss function.
[0008] Further, in Step 1, the method of controlling the model complexity by introducing the L2 regularization method specifically includes: Step 1.1: Construct the state space representation of the autonomous underwater vehicle:
[0009] Among them, S represents the state space representation of the autonomous underwater vehicle, st represents the environmental information of the autonomous underwater vehicle at time step t when, t represents the current timestamp, tend represents the timestamp at the final moment when the navigation ends, end is the navigation end flag; represents the state vector composed of , and these four variables respectively represent the vector of the current position of the autonomous underwater vehicle, the current speed vector of the autonomous underwater vehicle, the current heading angle of the autonomous underwater vehicle, and the spatial information of the obstacles around the autonomous underwater vehicle; Step 1.2: Construct the action space representation of the autonomous underwater vehicle:
[0010] Among them, A represents the action space representation of the autonomous underwater vehicle, at represents the autonomous underwater vehicle at time stept The action space at t represents the current timestamp, tend the timestamp indicating the final moment of the navigation end, end is the navigation end flag; represents the state vector composed of These two variables respectively represent the vector representation of the change in the speed of the autonomous underwater vehicle and the vector representation of the change in the heading angle of the autonomous underwater vehicle; Step 1.3: Train the deep Q-network. During the iterative process based on deep reinforcement learning, calculate the estimated value of the future cumulative reward obtained in each attempt according to the following formula:
[0011] where, s t represents the environmental information of the autonomous underwater vehicle at time step t , s t+1 represents the environmental information of the autonomous underwater vehicle at time step t +1, t represents the current timestamp, a t represents the action space of the autonomous underwater vehicle at time step t , a’ represents the optimal action among all possible actions in the next state s t+1 ; max represents the function of taking the maximum value; represents the parameters of the deep Q-network, represents at t the estimated value of the future cumulative reward obtained by fitting the deep Q-network with parameters s t when executing the action a t in the environment ; r t represents the immediate reward obtained after executing the action s t in the state a t , represents the discount factor, whose meaning is to control the importance of future rewards; represents the parameters of the target network, which are used as parameters for delayed update to improve the stability of training; Step 1.4: During the training process of the deep Q-network, initialize its loss function according to the following formula:
[0012] Among them, is the initial value of the loss function, represents the parameters of the deep Q-network, represents at t the moment in the environment st when performing the action at , the estimated value of the future cumulative reward obtained by fitting the deep Q-network with parameters ; st represents the environmental information of the autonomous underwater vehicle at time step t , st +1 represents the environmental information of the autonomous underwater vehicle at time step t +1, t represents the current timestamp, at represents the action space of the autonomous underwater vehicle at time step t , rt represents the immediate reward obtained after performing the action st in the state at ; D represents the experience replay pool, represents the operation function for sampling from the experience replay pool D for training; yt represents the target Q value at time t , and the Q value represents the estimated value of the reward; Step 1.5: Introduce L2 regularization to correct the loss function; Step 1.6: Update the model parameters according to the following formula , to minimize the loss function:
[0013] Among them, represents the parameters of the deep Q-network, the symbol represents the update operation; α is the learning rate, used to control the step size of parameter update; represents the operation function for calculating the gradient of , represents the gradient of the loss function with respect to the parameter θ , used to update the model weights; represents the final loss function obtained after introducing L2 regularization for correction.
[0014] Furthermore, in Step 1.5, introduce L2 regularization to correct the loss function according to the following formula:
[0015] Among them, represents the final loss function obtained after introducing L2 regularization for correction, is the initial value of the loss function, represents the parameters of the deep Q-network, represents the regularization strength coefficient, which is used to control the weight size of L2 regularization; i represents the number value of the weight parameter, represents the weight parameter, represents the sum of the squares of all the weight parameters of the network .
[0016] Furthermore, in step 3, the parameters of the incremental path decision network model are fine-tuned batch by batch using the LoRA fine-tuning technique, which specifically includes: Step 3.1: During the dynamic decision-making process, the autonomous underwater vehicle continuously updates the Q-value function according to the data collected by the sensors in real time. The Q-value function represents the calculation function of the reward estimation value and makes the optimal path decision based on the current state. Among them, the incremental update mechanism of the Q-value function over time steps is used to adapt to the changing environment; the specific incremental update mechanism is: every time the sensor detects new information, the model will incrementally update the Q-value function according to these new data, gradually approaching the optimal path planning; Step 3.2: Perform low-rank decomposition operation according to the following formula:
[0017] where, W is the updated weight matrix obtained by using the low-rank decomposition operation; represents the original weight matrix, which represents the parameters of the model pre-training and is extracted from the deep Q-network with parameters ; represents the update amount of the model parameters; A and B are both low-rank matrices, d , r , k all represent the dimensions of the matrix; represents a matrix over the real number field, represents a d row r column matrix over the real number field, represents a r row k column matrix over the real number field; Step 3.3: The new environmental data collected by the sensor is passed to the incremental path decision network model in batches, and each batch will trigger a LoRA fine-tuning; each time a batch of data is collected, the operation of updating the low-rank matrix using this batch of data satisfies the following formula:
[0018] where, Represents the batch loss function used to update the parameters of LoRA fine-tuning; Represents the parameters of the deep Q-network; N represents the batch size, which means the number of data samples included in each batch; yt Represents the time step t The target Q-value at, where the Q-value represents the estimated value of the reward; Represents at t The time step when in the environment st Executing the action at The estimated value of the future cumulative reward obtained by fitting the deep Q-network with parameters ; i, j, k Both represent the number of the current data sample in the current batch; si Represents the environmental information of the autonomous underwater vehicle at the data sample i ; ai Represents the action space of the autonomous underwater vehicle at the data sample i ; Represents the target Q-value, which represents the actual reward value of the current state; A And B Are both low-rank matrices, Aj And Bk Respectively represent the column vectors of the low-rank matrices A And B In LoRA fine-tuning, which are used to represent some of the parameters for model update; Represents the regularization coefficient used to control the update amplitude of the low-rank matrix; Step 3.4: Incrementally update the Q-value function.
[0019] Furthermore, in Step 3.4, the Q-value function is incrementally updated according to the following formula:
[0020] Where, Represents the Q-value after incremental update, Represents the Q-value of the model before update; Represents the maximum Q-value among all possible actions in the next step; Represents the parameters of the deep Q-network after incremental update, Represents the parameters of the deep Q-network before incremental update, Represents the parameters of the deep Q-network when obtaining the maximum Q-value among all possible actions in the next step; max represents the function of taking the maximum value; Is the learning rate, which determines the update step size, Represents the immediate reward obtained by the autonomous underwater vehicle when in the state t At the time step st And executing the action at ;
[0021] Further, in step 5, the calculation formula for the return constraint of the autonomous underwater vehicle is as follows:
[0022] Among them, CON represents the return constraint. When its value is 0, it means that the return constraint is not satisfied. When its value is 1, it means that the return constraint is satisfied; represents the instruction constraint. When a return instruction is received from the control center, its value is 1, otherwise it is 0; represents the maximum navigation time constraint, t represents the current timestamp, represents the maximum navigation time; represents the battery constraint, represents the depth Q-network parameter; represents at time t the remaining battery power; is the estimated power consumption function for return, indicating that using the depth Q-network parameter with parameter at time t estimates the power consumption for its return to the starting point; represents the reserved battery power.
[0023] In a second aspect, the present invention provides an autonomous underwater vehicle control system based on an incremental regularization network for the method as described above. The system includes: A data acquisition module, configured to acquire historical data and real-time environmental data during the navigation of the autonomous underwater vehicle; A data transmission module, configured to obtain historical navigation data from the data center and transmit the navigation data back to the data center at the end of the navigation; A model training module, configured to perform training for path planning based on the acquired data using a depth reinforcement learning model introducing L2 regularization; A fine-tuning module, based on the LoRA fine-tuning technique, dynamically updates the model parameters to achieve real-time path planning adjustment; A path decision module, configured to output the motion path of the autonomous underwater vehicle according to the incremental path decision network, avoid dynamic obstacles and optimize the navigation path.
[0024] In a third aspect, the present invention provides a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, it executes the steps of an autonomous underwater vehicle control method based on an incremental regularization network as described above.
[0025] Fourthly, the present invention provides a computer-readable storage medium storing multiple computer instructions for causing a computer to execute an autonomous underwater vehicle control method based on an incremental regularization network as described above.
[0026] Advantages of the present invention: The present invention provides an autonomous underwater vehicle path planning method based on an incremental regularization network. When training a deep reinforcement learning model using historical data, the autonomous underwater vehicle control method based on an incremental path decision network introduces the L2 regularization method, thereby effectively controlling the complexity of the model while preventing overfitting and quickly realizing real-time path planning. An incremental path dynamic decision mechanism is introduced to extract the parameters of the trained deep reinforcement learning model and perform LoRA fine-tuning on the model parameters batch by batch, enabling the model to dynamically update its Q-value function during the movement of the autonomous underwater vehicle. Based on the deep reinforcement learning method, the present invention designs an incremental path decision network for the autonomous underwater vehicle, enabling it to make incremental decisions according to the newly obtained information during the movement process, having the ability of autonomous path planning to prevent collisions with underwater dynamic targets such as marine organisms, and reducing the recovery risk of the autonomous underwater vehicle.
[0027] The present invention is applicable to the autonomous path planning of autonomous underwater vehicles in complex environments. Description of the Drawings
[0028] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is a flowchart of the method of an embodiment of an autonomous underwater vehicle control method based on an incremental path decision network provided by the present invention; Figure 2 It is the planning result of an autonomous underwater vehicle with an incremental path decision network of the present invention. Detailed Embodiments
[0030] The following details the embodiments of the present invention. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0031] Embodiment 1. An autonomous underwater vehicle control method based on an incremental regularization network, the method includes: S1: Obtain the AUV historical underwater navigation database from the data center, where AUV stands for Autonomous Underwater Vehicle. Train using the Deep Q-Network (a deep reinforcement learning model), and control the model complexity by introducing the L2 regularization method to prevent model overfitting and enhance the generalization ability of the model in complex underwater environments. The meaning of L2 regularization is a method to make the model smoother by adding the sum of parameter squares to the loss function; S2: The Autonomous Underwater Vehicle starts from the starting point and performs initial path planning based on the initial environmental information collected by the sensors; S3: Introduce an incremental path dynamic decision-making mechanism. The sensors of the Autonomous Underwater Vehicle collect the surrounding environmental data in real-time, and use the LoRA fine-tuning technology to fine-tune the parameters of the incremental path decision network model in batches, so that the Q-value function of the model can be updated in real-time. The meaning of the Q-value function is a reward function simulated using the deep reinforcement learning model. The incremental path decision network model refers to a deep reinforcement learning model trained using the incremental path dynamic decision-making mechanism. LoRA refers to the low-rank decomposition model fine-tuning technology; S4: Use the incremental path decision network model fine-tuned by LoRA to update the path planning of the Autonomous Underwater Vehicle, so as to achieve incremental path dynamic decision-making for the newly collected batch of environmental information, avoid dynamic obstacles such as marine organisms, and optimize the path planning of the Autonomous Underwater Vehicle; It should be noted that the incremental path decision network model is the incremental path decision network model fine-tuned using the LoRA fine-tuning technology. After fine-tuning, this model will also perform an update operation on the path planning.
[0032] S5: The Autonomous Underwater Vehicle checks whether the return constraints are met. If the return constraints are not met, return to step S3, continue underwater navigation, and use the collected data for model fine-tuning and path update. If the return constraints are met, execute step S6; S6: The Autonomous Underwater Vehicle calls the incremental path decision network to generate a path planning scheme to return from the current position to the starting point, and returns according to the scheme. After the Autonomous Underwater Vehicle arrives at the starting point, it uploads the underwater navigation data of this time to the historical underwater navigation database in the data center.
[0033] In this embodiment, by introducing the L2 regularization method, the model complexity is controlled to prevent overfitting of the model and enhance the generalization ability of the model in a complex underwater environment; the LoRA fine-tuning technology is used to fine-tune the parameters of the incremental path decision network model batch by batch, so that the Q-value function of the model can be updated in real time. The Q-value function is a reward function simulated by using a deep reinforcement learning model; the incremental path decision network model updated after LoRA fine-tuning is used to update the path planning of the autonomous underwater vehicle, so as to realize the incremental path dynamic decision-making for the newly collected batch of environmental information, avoid dynamic obstacles such as marine organisms, and optimize the path planning of the autonomous underwater vehicle.
[0034] This embodiment provides a method for controlling an autonomous underwater vehicle that can quickly respond to changes in a dynamic environment. The method for controlling an autonomous underwater vehicle based on an incremental path decision network proposed in this embodiment aims to effectively reduce the model complexity and improve the real-time decision-making ability of the model by introducing the L2 regularization technology and the incremental path dynamic decision mechanism. This method can not only reduce the collision risk with underwater dynamic targets, but also achieve autonomous path planning in a complex environment, thereby enhancing the safety and reliability of the autonomous underwater vehicle.
[0035] Embodiment 2 is a further limitation on a method for controlling an autonomous underwater vehicle based on an incremental regularization network as described above. In this embodiment, the deep Q-network with L2 regularization in step S1 is further limited, specifically including: In order to improve the generalization performance of the model, the deep Q-network with L2 regularization in step S1 is constructed according to the following steps: Step 1-1: Construct the state space representation of the autonomous underwater vehicle according to the following formula:
[0036] where, S represents the state space representation of the autonomous underwater vehicle, s t represents the environmental information of the autonomous underwater vehicle at time step t t, t represents the current timestamp, t end represents the timestamp at the final moment when the navigation ends, end is the navigation end flag; represents the state vector composed of , and these four variables respectively represent the vector of the current position of the autonomous underwater vehicle, the current speed vector of the autonomous underwater vehicle, the current heading angle of the autonomous underwater vehicle, and the spatial information of the obstacles around the autonomous underwater vehicle; Step 1-2: Construct the action space representation of the autonomous underwater vehicle according to the following formula:
[0037] where, A represents the action space representation of the autonomous underwater vehicle, a t represents the action space of the autonomous underwater vehicle at time step t ; t represents the current timestamp, t end represents the timestamp of the final moment when the navigation ends, end is the navigation end flag; represents the state vector composed of , and these two variables respectively represent the vector representation of the change in the speed of the autonomous underwater vehicle and the vector representation of the change in the heading angle of the autonomous underwater vehicle; Step 1-3: Train the deep Q-network. During the iterative process based on deep reinforcement learning, calculate the estimated value of the future cumulative reward obtained in each attempt according to the following formula:
[0038] where, s t represents the environmental information of the autonomous underwater vehicle at time step t ; s t+1 represents the environmental information of the autonomous underwater vehicle at time step t +1; t represents the current timestamp, a t represents the action space of the autonomous underwater vehicle at time step t ; a’ represents the optimal action among all possible actions in the next state s t+1 ; max represents the function of taking the maximum value; represents the deep Q-network parameters, represents at t the estimated value of the future cumulative reward obtained by the deep Q-network with parameters s t fitted when executing the action a t in the environment ; r t represents the immediate reward obtained after executing the action s t in the state a t ; represents the discount factor, which means controlling the importance of future rewards; represents the parameters of the target network, which are used as parameters for delayed update to improve the stability of training; Step 1 - 4: During the training of the deep Q - network, initialize its loss function according to the following formula:
[0039] where, is the initial value of the loss function, represents the parameters of the deep Q - network, represents at t the estimated value of the future cumulative reward obtained by the deep Q - network with parameters s t when executing the action a t in the environment ; s t represents the environmental information of the autonomous underwater vehicle at time step t ; s t+1 represents the environmental information of the autonomous underwater vehicle at time step t +1; t represents the current timestamp, a t represents the action space of the autonomous underwater vehicle at time step t ; r t represents the immediate reward obtained after executing the action s t in the state a t ; D represents the experience replay pool, represents the operation function of sampling from the experience replay pool D for training; y t represents the target Q - value at time t , and the Q - value represents the estimated value of the reward.
[0040] Step 1 - 5: During the process of the deep Q - network calculating the loss function, introduce L2 regularization to correct the loss function, thereby preventing the network parameters from becoming too large and improving the generalization ability of the model: Step 1 - 6: During the process of the deep Q - network updating the model using the loss function, update the model parameters θ according to the following formula to minimize the loss function with L2 regularization:
[0041] Among them, represents the parameters of the deep Q-network, and the symbol represents the update operation; α is the learning rate, which is used to control the step size of parameter update; represents the operation function for calculating the gradient, represents the gradient of the loss function with respect to the parameter θ , which is used to update the model weights; represents the final loss function obtained by introducing L2 regularization for correction.
[0042] This embodiment provides a specific construction method of a deep Q-network with L2 regularization to improve the generalization performance of the model. Steps 1-3 are used to improve the stability of training; the formulas in Steps 1-5 introduce L2 regularization to correct the loss function, thereby preventing the network parameters from becoming too large and enhancing the generalization ability of the model; Step 1-6 is used to minimize the loss function with L2 regularization.
[0043] By introducing the L2 regularization technique, it aims to effectively reduce the model complexity. By introducing the L2 regularization method, the evaluation result of the comprehensive level has been improved by 21.0% compared with the DQN model. The DQN model refers to the Deep Q-Network model, which is a reinforcement learning algorithm model combining deep learning and Q-learning. The comprehensive level refers to the mean of the path planning ability and the obstacle avoidance ability. The path planning ability refers to the negative value of the path length planned by the model, and the larger the value, the shorter the path; the obstacle avoidance ability refers to the negative value of the number of times of contacting obstacles or risk areas in the path planned by the model, and the larger the value, the greater the probability that the autonomous underwater vehicle can avoid obstacles.
[0044] Embodiment 3 further limits a method for controlling an autonomous underwater vehicle based on an incremental regularization network as described above. In this embodiment, the formula for introducing L2 regularization is further limited, specifically including: In Step 1-5, introduce L2 regularization to correct the loss function according to the following formula:
[0045] Among them, represents the final loss function obtained by introducing L2 regularization for correction, is the initial value of the loss function, represents the parameters of the deep Q-network, represents the regularization strength coefficient, which is used to control the weight size of L2 regularization; i represents the number value of the weight parameter, represents the weight parameter, represents all the weight parameters of the network Sum of squares.
[0046] This embodiment provides a method for introducing L2 regularization to correct the loss function. The loss function for the training process of the deep Q-network is initialized using the formulas in steps 1 - 4, and then the L2 regularization formula in steps 1 - 5 is incorporated into the optimization process of minimizing the loss function of the deep Q-network in steps 1 - 6. This embodiment combines the method of combining the formulas in steps 1 - 4 and steps 1 - 6 (i.e., the deep Q-network training and optimization process) with the method of the formula in steps 1 - 5 (i.e., the L2 regularization process), which can more effectively introduce the L2 regularization method to control the model complexity and make the model smoother by adding the sum of parameter squares to the loss function.
[0047] Embodiment 4 further limits a method for controlling an autonomous underwater vehicle based on an incremental regularization network as described above. In this embodiment, the parameters of the incremental path decision network model are fine-tuned batch by batch using the LoRA fine-tuning technique in step 3, which is further limited as follows: To enhance the decision-making ability of the autonomous underwater vehicle in the underwater dynamic environment, in step S3, the parameters of the incremental path decision network model are fine-tuned batch by batch using the LoRA fine-tuning technique according to the following steps. The incremental path decision network model refers to a deep reinforcement learning model trained using an incremental path dynamic decision-making mechanism, and LoRA refers to the low-rank decomposition model fine-tuning technique: Step 3 - 1: During the dynamic decision-making process, the autonomous underwater vehicle continuously updates the Q-value function based on the data collected by the sensors in real time and makes an optimal path decision based on the current state. The key is that the Q-value function is incrementally updated over time steps to adapt to the changing environment. The meaning of the incremental update mechanism is that each time the sensor detects new information, the model incrementally updates the Q-value function based on this new data, gradually approaching the optimal path planning. The Q-value function represents the calculation function of the reward estimate value; Step 3 - 2: Perform low-rank decomposition operations according to the following formula:
[0048] Where, W is the updated weight matrix obtained by low-rank decomposition operation; represents the original weight matrix, which is the parameter of the model pre-training and is extracted from the deep Q-network with parameters ; represents the update amount of the model parameters; A and B are both low-rank matrices, d , r , kBoth represent the dimensions of the matrix; represents a matrix over the real number field, represents a d row r column matrix over the real number field, represents a r row k column matrix over the real number field; Step 3-3: LoRA fine-tuning is carried out batch by batch, that is, the newly collected environmental data by the sensor is passed to the incremental path decision network model in batches, and LoRA fine-tuning is triggered once for each batch. This mechanism ensures that the model can gradually adapt to the changing environment without the need for frequent full-scale updates. Each time a batch of data is collected, the operation of updating the low-rank matrix using this batch of data satisfies the following formula:
[0049] where, represents the batch loss function, which is used to update the parameters of LoRA fine-tuning, and the LoRA refers to the low-rank decomposition model fine-tuning technology. represents the deep Q-network parameters. N represents the batch size, which means the number of data samples included in each batch. yt represents the target Q-value at time t, and the Q-value represents the estimated value of the reward. represents at t time in the environment st when performing the action at with parameters the estimated value of the future cumulative reward obtained by fitting the deep Q-network. i, j, k Both represent the number of the current data sample in the current batch. si represents the environmental information of the autonomous underwater vehicle at the data sample i ; ai represents the action space of the autonomous underwater vehicle at the data sample i ; represents the target Q-value, which represents the actual reward value of the current state; A and B are both low-rank matrices, Aj and Bk respectively represent the column vectors of the low-rank matrices A and B in LoRA fine-tuning, which are used to represent some of the parameters for model update, represents the regularization coefficient, which is used to control the amplitude of the low-rank matrix update to avoid overfitting; Step 3-4: After each execution of LoRA fine-tuning, the Q-value function of the incremental path decision network is updated in real time, ensuring that the autonomous underwater vehicle can dynamically adjust its path planning strategy according to new information. The incremental update of the Q-value function means that whenever new environmental data is collected by the sensors of the autonomous underwater vehicle, the network parameters adjusted by LoRA fine-tuning are reflected in the Q-value function, enabling the model to better adapt to the current environment. The Q-value represents an estimated value of the reward.
[0050] This embodiment presents a method for fine-tuning the parameters of the incremental path decision network model in batches using LoRA fine-tuning technology. The incremental path decision network model refers to a deep reinforcement learning model trained using an incremental path dynamic decision mechanism. LoRA refers to the low-rank decomposition model fine-tuning technology, which is used to enhance the decision-making ability of the autonomous underwater vehicle in the underwater dynamic environment.
[0051] Among them, the mechanism of performing LoRA fine-tuning in batches in Step 3-3 ensures that the model can gradually adapt to the changing environment without the need for frequent full-scale updates; the regularization coefficient used to control the amplitude of the low-rank matrix update is used to avoid overfitting. After each execution of LoRA fine-tuning, the Q-value function of the incremental path decision network is updated in real time, ensuring that the autonomous underwater vehicle can dynamically adjust its path planning strategy according to new information. The incremental update mechanism of the Q-value function means that whenever new environmental data is collected by the sensors of the autonomous underwater vehicle, the network parameters adjusted by LoRA fine-tuning are reflected in the Q-value function, enabling the model to better adapt to the current environment.
[0052] By introducing LoRA fine-tuning technology, the aim is to improve the real-time decision-making ability of the model. By introducing LoRA fine-tuning technology, the evaluation result of the comprehensive level has been improved by 27.6% compared with the DQN model. The DQN model refers to the Deep Q-Network model, which is a reinforcement learning algorithm model that combines deep learning and Q-learning. The comprehensive level refers to the mean of the planning ability and the obstacle avoidance ability. The planning ability refers to the negative value of the path length planned by the model, and the larger the value, the shorter the path; the obstacle avoidance ability refers to the negative value of the number of times of contacting obstacles or risk areas in the path planned by the model, and the larger the value, the greater the probability that the autonomous underwater vehicle can avoid obstacles.
[0053] Embodiment 5, this embodiment further limits a method for controlling an autonomous underwater vehicle based on an incremental regularization network as described above. In this embodiment, the update formula of the Q-value function is further limited, specifically including: In Step 3-4, the Q-value function is incrementally updated according to the following steps:
[0054] Among them, represents the Q value after incremental update, represents the Q value of the model before update. represents the maximum Q value among all possible actions in the next step. represents the parameters of the deep Q network after incremental update, represents the parameters of the deep Q network before incremental update, represents the parameters of the deep Q network when obtaining the maximum Q value among all possible actions in the next step. st represents the environmental information of the autonomous underwater vehicle at time step t, st +1 represents the environmental information of the autonomous underwater vehicle at time step t + 1, where t represents the current time stamp, at represents the action space of the autonomous underwater vehicle at time step t, a’ represents the next state st +1 is the optimal action among all possible actions. max represents the function of taking the maximum value. is the learning rate, which determines the update step size, represents the autonomous underwater vehicle at time step t when in state st executes action at the immediate reward obtained thereafter.
[0055] This embodiment provides a method for incrementally updating the Q-value function. This method is combined with the LoRA low-rank decomposition method in step 3-2 and the network parameter update method in step 3-3, which can make the application of the LoRA technology in the deep Q network more effectively realize the control of the autonomous underwater vehicle in this application and improve the real-time decision-making ability of the model.
[0056] Embodiment 6 further limits the method for controlling an autonomous underwater vehicle based on an incrementally regularized network as described above. In this embodiment, the calculation formula for the return constraint of the autonomous underwater vehicle in step 5 is further limited, specifically including: To reduce the risk that the autonomous underwater vehicle cannot return, in step S5, the return constraint of the autonomous underwater vehicle is calculated according to the following formula:
[0057] Among them, CON represents the return constraint. When its value is 0, it means that the return constraint is not satisfied. When its value is 1, it means that the return constraint is satisfied. max is the function of taking the maximum value. represents the command constraint. When receiving a return command from the control center, its value is 1, otherwise it is 0; Represents the maximum navigation time constraint, t Represents the current timestamp, Represents the maximum navigation time; Represents the battery constraint, Represents the parameters of the deep Q-network; Represents at time t The remaining battery power; Is the estimated power consumption function for the return journey, indicating that using the parameters of the deep Q-network as At time t Estimate the power consumption when it returns to the starting point; Represents the reserved battery power to prevent the situation where it cannot return to the starting point due to unexpected additional power consumption during the return journey.
[0058] Embodiment 7. This embodiment is an example of an autonomous underwater vehicle control method based on an incremental regularization network as described above. As Figure 1 Shown, an autonomous underwater vehicle control method based on an incremental path decision network includes: S1: Obtain the AUV historical underwater navigation database from the data center. The meaning of AUV is an autonomous underwater vehicle. Train using the deep Q-network and control the model complexity by introducing the L2 regularization method to prevent model overfitting and enhance the generalization ability of the model in complex underwater environments. The meaning of L2 regularization is a method to make the model smoother by adding the sum of squares of parameters to the loss function; In the step S1, the deep Q-network with L2 regularization is constructed according to the following steps: Step 1-1: Construct the state space representation of the autonomous underwater vehicle according to the following formula:
[0059] Among them, S Represents the state space representation of the autonomous underwater vehicle, st Represents the environmental information of the autonomous underwater vehicle at time step t When, t Represents the current timestamp, tend Represents the timestamp at the final moment when the navigation ends, end Is the navigation end marker; Represents the state vector composed of These four variables respectively represent the vector of the current position of the autonomous underwater vehicle, the current speed vector of the autonomous underwater vehicle, the current heading angle of the autonomous underwater vehicle, and the spatial information of the obstacles around the autonomous underwater vehicle; Step 1-2: Construct the action space representation of the autonomous underwater vehicle according to the following formula:
[0060] Among them, A represents the action space representation of the autonomous underwater vehicle, at represents the action space of the autonomous underwater vehicle at time step t t, t represents the current timestamp, tend represents the timestamp of the final moment when the navigation ends, end is the navigation end flag; represents the state vector composed of These two variables respectively represent the vector representation of the change in the speed of the autonomous underwater vehicle and the vector representation of the change in the heading angle of the autonomous underwater vehicle; Steps 1 - 3: Train the deep Q - network. During the iterative process of deep reinforcement learning, calculate the estimated value of the future cumulative reward obtained for each attempt according to the following formula:
[0061] Among them, st represents the environmental information of the autonomous underwater vehicle at time step t t, st +1 represents the environmental information of the autonomous underwater vehicle at time step t t + 1, t represents the current timestamp, at represents the action space of the autonomous underwater vehicle at time step t t, a’ represents the optimal action among all possible actions in the next state st t + 1; max represents the function of taking the maximum value; represents the parameters of the deep Q - network, represents at t the moment in the environment st when performing the action at a, the estimated value of the future cumulative reward obtained by fitting the deep Q - network with parameters θ; rt represents the immediate reward obtained after performing the action st s and at a, represents the discount factor, whose meaning is to control the importance of future rewards; represents the parameters of the target network, which are used as parameters for delayed update to improve the stability of training; Steps 1 - 4: During the training process of the deep Q - network, initialize its loss function according to the following formula:
[0062] Among them, is the initial value of the loss function, represents the parameters of the deep Q-network, represents at t time in the environment st when performing the action at , the estimated value of the future cumulative reward obtained by fitting the deep Q-network with parameters ; st represents the environmental information of the autonomous underwater vehicle at time step t , st +1 represents the environmental information of the autonomous underwater vehicle at time step t +1, t represents the current timestamp, at represents the action space of the autonomous underwater vehicle at time step t , rt represents the immediate reward obtained after performing the action st in the state at ; D represents the experience replay pool, represents the operation function for sampling from the experience replay pool D for training; yt represents the target Q value at time t , and the Q value represents the estimated value of the reward.
[0063] Steps 1-5: In the process of calculating the loss function of the deep Q-network, introduce L2 regularization according to the following formula to correct the loss function, so as to prevent the network parameters from becoming too large and improve the generalization ability of the model:
[0064] Among them, represents the final loss function obtained after introducing L2 regularization for correction, is the initial value of the loss function, represents the parameters of the deep Q-network, represents the regularization strength coefficient, which is used to control the weight size of L2 regularization; i represents the number value of the weight parameter, represents the weight parameter, represents the sum of the squares of all the weight parameters of the network.
[0065] Steps 1-6: In the process of updating the model using the loss function in the deep Q-network, update the model parameters according to the following formula to minimize the loss function with L2 regularization:
[0066] Among them, Represents the parameters of the deep Q-network, The symbol represents the update operation; α $\alpha$ is the learning rate, used to control the step size of parameter update; represents the operation function for calculating the gradient, represents the gradient of the loss function with respect to the parameter θ $\theta$, used to update the model weights; $\mathcal{L}$ represents the final loss function obtained after introducing L2 regularization for correction.
[0067] S2: The autonomous underwater vehicle starts from the starting point and conducts initial path planning based on the initial environmental information collected by the sensors; S3: Introduce an incremental path dynamic decision-making mechanism. The sensors of the autonomous underwater vehicle continuously collect surrounding environmental data in real time, and use the LoRA fine-tuning technology to fine-tune the parameters of the incremental path decision-making network model batch by batch, so that the Q-value function of the model can be updated in real time. The meaning of the Q-value function is the reward function simulated by the deep reinforcement learning model. The meaning of the incremental path decision-making network model is the deep reinforcement learning model trained by the incremental path dynamic decision-making mechanism. The LoRA refers to the low-rank decomposition model fine-tuning technology; In step S3, according to the following steps, use the LoRA fine-tuning technology to fine-tune the parameters of the incremental path decision-making network model batch by batch. The meaning of the incremental path decision-making network model is the deep reinforcement learning model trained by the incremental path dynamic decision-making mechanism. The LoRA refers to the low-rank decomposition model fine-tuning technology: Step 3-1: In the dynamic decision-making process, the autonomous underwater vehicle continuously updates the Q-value function according to the data collected by the sensors in real time, and makes the optimal path decision based on the current state. The key lies in the incremental update mechanism of the Q-value function over time steps to adapt to the changing environment. The meaning of the incremental update mechanism is that every time the sensor detects new information, the model will incrementally update the Q-value function according to these new data, gradually approaching the optimal path planning. The Q-value function represents the calculation function of the reward estimate value; Step 3-2: Perform low-rank decomposition operation according to the following formula:
[0068] where, W $\tilde{W}$ is the updated weight matrix obtained by the low-rank decomposition operation; $W$ represents the original weight matrix, which represents the parameters of the model pre-training and is extracted from the deep Q-network with parameters $\theta$; $\Delta W$ represents the update amount of the model parameters; A $A$ and B $B$ are both low-rank matrices,d , r , k All represent the dimensions of the matrix; represents a matrix over the real number field, represents a d row r column matrix over the real number field, represents a r row k column matrix over the real number field; Step 3 - 3: LoRA fine - tuning is carried out batch - by - batch. That is, the newly acquired environmental data collected by the sensor is passed to the incremental path decision network model in batches, and LoRA fine - tuning is triggered once for each batch. This mechanism ensures that the model can gradually adapt to the changing environment without the need for frequent full - scale updates. Each time a batch of data is collected, the operation of updating the low - rank matrix using this batch of data satisfies the following formula:
[0069] where, represents the batch loss function, which is used to update the parameters of LoRA fine - tuning. The LoRA refers to the low - rank factorization model fine - tuning technique. represents the parameters of the deep Q - network; N represents the batch size, which means the number of data samples included in each batch; yt represents the moment t the target Q - value at which the Q - value represents an estimate of the reward; represents at t the moment in the environment st when the action at is executed, the estimated value of the future cumulative reward obtained by fitting the deep Q - network with parameters ; i, j, k All represent the numbers of the current data samples in the current batch; si represents the environmental information of the autonomous underwater vehicle at the data sample i ; ai represents the action space of the autonomous underwater vehicle at the data sample i ; represents the target Q - value, which represents the actual reward value of the current state; A and B are both low - rank matrices, Aj and Bk respectively represent the column vectors of the low - rank matrices A and B in LoRA fine - tuning, which are used to represent some of the parameters for model update; represents the regularization coefficient, which is used to control the amplitude of the low - rank matrix update to avoid overfitting; Step 3-4: After each execution of LoRA fine-tuning, the Q-value function of the incremental path decision network is updated in real time to ensure that the autonomous underwater vehicle can dynamically adjust its path planning strategy according to new information. The incremental update mechanism of the Q-value function means that whenever new environmental data is collected by the sensors of the autonomous underwater vehicle, the network parameters adjusted by LoRA fine-tuning will be reflected in the Q-value function, enabling the model to better adapt to the current environment. The Q-value represents an estimated value of the reward. The Q-value function is incrementally updated according to the following steps:
[0070] where, represents the Q-value after incremental update, represents the Q-value of the model before update. represents the maximum Q-value among all possible actions in the next step. represents the parameters of the deep Q-network after incremental update, represents the parameters of the deep Q-network before incremental update, represents the parameters of the deep Q-network when obtaining the maximum Q-value among all possible actions in the next step. st represents the environmental information of the autonomous underwater vehicle at time step t, st +1 represents the environmental information of the autonomous underwater vehicle at time step t+1, where t represents the current timestamp, at represents the action space of the autonomous underwater vehicle at time step t, a’ represents the optimal action among all possible actions in the next state st +1. max represents the function of taking the maximum value. is the learning rate, which determines the update step size.
[0071] Figure 2 is the path planning result of the autonomous underwater vehicle with an incremental path decision network according to one or more embodiments. As Figure 2 shown, the autonomous underwater vehicle starts from the starting point and goes to the end point. The gray area represents that the autonomous underwater vehicle detects the presence of obstacles at that position. The planning result based on the simulation experiment is plotted as a broken line in Figure 2 . In Figure 2 , cases 1, 2, 3, and 4 refer to different results of multiple path planning of the autonomous underwater vehicle based on the incremental path decision network, illustrating the convergence and robustness of the results.
[0072] S4: Use the updated incremental path decision network model after LoRA fine-tuning to update the path planning of the autonomous underwater vehicle, so as to achieve incremental path dynamic decision-making for the newly collected batch of environmental information, avoid dynamic obstacles such as marine organisms, and optimize the path planning of the autonomous underwater vehicle; In one or more embodiments disclosed in the present application, the evaluation results of the autonomous underwater vehicle based on the incremental path decision network model are shown in Table 1. The evaluation results include three dimensions: planning ability, obstacle avoidance ability, and comprehensive level. The values of each dimension are obtained by dividing by the maximum value in the corresponding row, so the results are all distributed in the interval of 0 and 1. The planning ability refers to the negative value of the path length planned by the model. The larger the value, the shorter the path. The obstacle avoidance ability refers to the negative value of the number of times the path planned by the model touches obstacles or risk areas. The larger the value, the greater the probability that the autonomous underwater vehicle avoids obstacles. The comprehensive level refers to the mean value of the planning ability and the obstacle avoidance ability. Table 1 shows the evaluation results of the method of the present application, the method of the present application - L2 regularization, the method of the present application - LoRA fine-tuning, and the DQN model in the three dimensions of planning ability, obstacle avoidance ability, and comprehensive level. The method of the present application - L2 regularization refers to the model obtained by deleting the L2 regularization module from the model proposed in the present application. The method of the present application - LoRA fine-tuning refers to the model obtained by deleting the LoRA fine-tuning module from the model proposed in the present application. The DQN model refers to the Deep Q-Network model, which is a reinforcement learning algorithm model that combines deep learning and Q-learning. It can be seen that by introducing the L2 regularization method or the LoRA fine-tuning technique, the evaluation results of the comprehensive level are respectively improved by 21.0% and 27.6% compared with the DQN model. The present invention combines the L2 regularization method and the LoRA fine-tuning technique, so that the evaluation result of the comprehensive level is improved by 46.1% compared with the DQN model, and the evaluation results of the planning ability and the obstacle avoidance ability are respectively improved by 24.7% and 67.5% compared with the DQN model. This shows that the present invention can effectively use the L2 regularization method to reduce the model complexity and improve the operation efficiency of the model, and at the same time effectively use the LoRA fine-tuning technique to improve the real-time decision-making ability of the model, so as to finally effectively improve the path planning ability of the autonomous underwater vehicle based on dynamic environment information, and can significantly improve the obstacle avoidance ability of the autonomous underwater vehicle in a dynamic environment.
[0073] Table 1 Evaluation Results of Autonomous Underwater Vehicle Based on Incremental Path Decision Network Model
[0074] S5: The autonomous underwater vehicle checks whether the return constraints are met. If the return constraints are not met, it returns to step S3 to continue underwater navigation and uses the collected data for model fine-tuning and path update. If the return constraints are met, it executes step S6; S6: The autonomous underwater vehicle calls the incremental path decision network to generate a path planning scheme to return from the current position to the starting point, and returns along the scheme. After the autonomous underwater vehicle arrives at the starting point, it uploads the underwater navigation data of this time to the historical underwater navigation database in the data center.
Claims
1. A control method for an autonomous underwater vehicle based on an incremental regularization network, characterized in that: The method comprises: Step 1: Obtain the AUV historical underwater navigation database, use the deep Q network for training, and control the model complexity by introducing the L2 regularization method; Step 2: The autonomous underwater vehicle starts from the starting point and performs initial path planning based on the initial environmental information collected by the sensor; Step 3: Using the incremental path dynamic decision mechanism to train the deep reinforcement learning model to obtain the incremental path decision network model, the autonomous underwater vehicle sensor collects surrounding environment data in real time, and uses the LoRA fine-tuning technology to fine-tune the parameters of the incremental path decision network model in batches; the LoRA is a low-rank decomposition model fine-tuning technology; Step 4: Use the incremental path decision network model fine-tuned in step 3 to update the path planning of the autonomous underwater vehicle; Step 5: The autonomous underwater vehicle checks whether the return constraint is met. If not, it returns to step 3, continues underwater navigation, and uses the collected data to fine-tune the model and update the path; if the return constraint is met, it executes step 6; Step 6: The autonomous underwater vehicle calls the incremental path decision network to generate a path planning plan from the current position back to the starting point, and returns according to the plan; after the autonomous underwater vehicle arrives at the starting point, it transmits the underwater navigation data back to the historical underwater navigation database of the data center.
2. The control method of an autonomous underwater vehicle based on an incremental regularization network according to claim 1, characterized in that: In step 1, controlling the model complexity by introducing the L2 regularization method is specifically to make the model smoother by adding the square sum of parameters to the loss function.
3. The control method of an autonomous underwater vehicle based on an incremental regularization network according to claim 1, characterized in that: In step 1, the model complexity is controlled by introducing the L2 regularization method, which specifically includes: Step 1.1: Construct the state-space representation of the autonomous underwater vehicle: in, S represents the state space representation of the autonomous underwater vehicle, st represents the autonomous underwater vehicle at time step t Environmental information at that time, t Represents the current timestamp, tend A timestamp indicating the final moment when the voyage ended, end Marking the end of a voyage; Indicated by The state vector is composed of four variables, which respectively represent the vector of the current position of the autonomous underwater vehicle, the current speed vector of the autonomous underwater vehicle, the current heading angle of the autonomous underwater vehicle, and the spatial information of obstacles around the autonomous underwater vehicle; Step 1.2: Construct the action space representation of the autonomous underwater vehicle: in, A represents the action space representation of the autonomous underwater vehicle, at represents the autonomous underwater vehicle at time step t The action space when t Represents the current timestamp, tend A timestamp indicating the final moment when the voyage ended, end Marking the end of a voyage; Indicated by The state vector is composed of two variables, which respectively represent the vector representation of the change of the speed of the autonomous underwater vehicle and the vector representation of the change of the heading angle of the autonomous underwater vehicle; Step 1.3: Train the deep Q network and calculate the estimated value of the future cumulative reward obtained for each attempt according to the following formula in the iterative process based on deep reinforcement learning: in, s t represents the autonomous underwater vehicle at time step t Environmental information at that time, s t+1 represents the autonomous underwater vehicle at time step t +1 environmental information, t Represents the current timestamp, a t represents the autonomous underwater vehicle at time step t The action space when a’ Indicates that in the next state s t+1 The best action among all possible actions; max represents the function that takes the maximum value; represents the deep Q network parameters, Indicated in t Always in the environment s t Next action a t When , the parameter is The estimated value of future cumulative rewards obtained by fitting the deep Q network; r t Indicates in status s t Execute an action a t After receiving the instant reward, Represents the discount factor, which means controlling the importance of future rewards; Represents the parameters of the target network, which are used as delayed update parameters to improve the stability of training; Step 1.4: During the training of the deep Q-network, initialize its loss function according to the following formula: in, is the initial value of the loss function, represents the deep Q network parameters, Indicated in t Always in the environment st Next action at When , the parameter is The estimated value of future cumulative rewards obtained by fitting the deep Q network; st represents the autonomous underwater vehicle at time step t Environmental information at that time, st +1 indicates that the autonomous underwater vehicle is at time step t +1 environmental information, t Represents the current timestamp, at represents the autonomous underwater vehicle at time step t The action space when rt Indicates in status st Execute an action at Immediate rewards after D Represents the experience replay pool, Indicates that the experience replay pool D Operation function for extracting samples for training; yt Indicates time t The target Q value under , wherein the Q value represents the estimated value of the reward; Step 1.5: Introduce L2 regularization to correct the loss function; Step 1.6: Update the model parameters according to the following formula , to minimize the loss function: in, represents the deep Q network parameters, The symbol indicates an update operation; α is the learning rate, which is used to control the step size of parameter update; Express The operation function that calculates the gradient, Represents the loss function for the parameter θ The gradient of is used to update the model weights; Represents the final loss function obtained after introducing L2 regularization.
4. The control method of an autonomous underwater vehicle based on an incremental regularization network according to claim 3 is characterized in that: In step 1.5, L2 regularization is introduced to correct the loss function according to the following formula: in, It represents the final loss function obtained after introducing L2 regularization. is the initial value of the loss function, represents the deep Q network parameters, represents the regularization strength coefficient, which is used to control the weight of L2 regularization; i represents the number value of the weight parameter, represents the weight parameter, Represents all weight parameters of the network The sum of squares.
5. The control method of an autonomous underwater vehicle based on an incremental regularization network according to claim 1, characterized in that: In step 3, the parameters of the incremental path decision network model are fine-tuned in batches using the LoRA fine-tuning technology, specifically including: Step 3.1: In the dynamic decision-making process, the autonomous underwater vehicle continuously updates the Q-value function according to the data collected by the sensor in real time. The Q-value function represents the calculation function of the reward estimation value, and makes the optimal path decision according to the current state. The Q-value function is incrementally updated over time to adapt to the changing environment. The incremental update mechanism is as follows: each time the sensor detects new information, the model will incrementally update the Q-value function according to the new data, and gradually approach the optimal path planning; Step 3.2: Perform low-rank decomposition operation according to the following formula: in, W is the updated weight matrix obtained by low-rank decomposition operation; Represents the original weight matrix, represents the parameters of the model pre-training, and the parameters are The deep Q network is extracted; Indicates the update amount of model parameters; A and B are all low-rank matrices, d , r , k Both represent the dimensions of the matrix; represents a matrix over the real number field, Indicates a d OK r A matrix over the real field of columns, Indicates a r OK k A matrix over the real field of columns; Step 3.3: The new environmental data collected by the sensor is passed to the incremental path decision network model in batches, and each batch triggers a LoRA fine-tuning; each time a batch of data is collected, the operation of updating the low-rank matrix using this batch of data satisfies the following formula: in, Represents the batch loss function, which is used to update the parameters of LoRA fine-tuning; represents the deep Q network parameters; N represents the batch size, which means the number of data samples contained in each batch; yt Indicates time t The target Q value under , wherein the Q value represents the estimated value of the reward; Indicated in t Always in the environment st Next action at When , the parameter is The estimated value of future cumulative rewards obtained by fitting the deep Q network; i, j, k Both represent the number of the current data sample in the current batch; si Represents the autonomous underwater vehicle in the data sample i The following environmental information, ai Represents the autonomous underwater vehicle in the data sample i The action space below; Represents the target Q value, which represents the actual reward value of the current state; A and B are all low-rank matrices, Aj and B Respectively represent the low-rank matrix in LoRA fine-tuning A and B , which is used to represent some parameters updated by the model; represents the regularization coefficient, which is used to control the amplitude of the low-rank matrix update; Step 3.4: Incrementally update the Q-value function.
6. The control method of an autonomous underwater vehicle based on an incremental regularization network according to claim 5, characterized in that: In step 3.4, the Q value function is incrementally updated according to the following formula: in, represents the Q value after incremental update, Represents the Q value of the model before updating; Indicates the maximum Q value of all possible actions in the next step; represents the deep Q network parameters after incremental update, represents the deep Q network parameters before incremental update, represents the deep Q network parameters when the maximum Q value among all possible actions for the next step is obtained; max represents the function that takes the maximum value; is the learning rate, which determines the update step size, represents the autonomous underwater vehicle at time step t When in state st Next action at The immediate reward obtained after that.
7. The control method of an autonomous underwater vehicle based on an incremental regularization network according to claim 1, characterized in that: In step 5, the calculation formula of the return constraint of the autonomous underwater vehicle is: in, CON Indicates the return constraint. When its value is 0, it means that the return constraint is not met. When its value is 1, it means that the return constraint is met. Indicates command constraint. When a return command is received from the control center, its value is 1, otherwise it is 0. represents the maximum flight time constraint, t Represents the current timestamp, Indicates the maximum flight time; represents the battery constraint, represents the parameters of the deep Q network; Indicates at time t The remaining power; is the return power consumption estimation function, which indicates that the parameters used are The parameters of the deep Q network, at time t Estimate the amount of power consumed when returning to the starting point; Indicates reserved power.
8. An autonomous underwater vehicle control system based on an incremental regularized network according to any one of claims 1 to 7, characterized in that: The system comprises: A data acquisition module, used to collect historical data and real-time environmental data during the navigation of the autonomous underwater vehicle; The data transmission module is used to obtain historical navigation data from the data center and transmit the navigation data back to the data center at the end of the navigation; Model training module, used to train path planning based on collected data using a deep reinforcement learning model with L2 regularization; Fine-tuning module, based on LoRA fine-tuning technology, dynamically updates model parameters to achieve real-time path planning adjustment; The path decision module is used to output the motion path of the autonomous underwater vehicle according to the incremental path decision network, avoid dynamic obstacles and optimize the navigation path.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor runs the computer program stored in the memory, the steps of the method according to any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of computer instructions, and the plurality of computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Autonomous underwater vehicle (AUV) adaptive fault diagnosis method based on discriminative feature learning method
CN110244689A
Path planning obstacle avoidance control method for autonomous underwater vehicle in large-scale continuous obstacle environment
CN112241176A
Underwater vehicle autonomous floating control method based on demonstration data reinforcement learning technology
CN113033118A
Autonomous underwater vehicle path planning method based on double neural network reinforcement learning
CN113064422A
Underwater robot real-time path planning method based on width reinforcement learning
CN119124175A
Cited By
Underwater robot navigation method and system, electronic equipment and storage medium
CN120846308A
A navigation method, system, electronic device and storage medium for an underwater robot
CN120846308B