A path planning method for a cable-driven manipulator used in underwater robot operations
The path planning of underwater robots through an adaptive multi-channel deep reinforcement learning network solves the problem of low path planning operation accuracy in the existing technology, and realizes the precise operation and autonomous operation capabilities of the robots in complex underwater environments.
Patent Information
- Application Number
- CN202510260084.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-06
AI Technical Summary
In the prior art, the operation accuracy of underwater rope-driven robot path planning is relatively low, especially in complex underwater environments, and it is difficult to achieve accurate position and attitude control.
Adaptive multi-channel deep reinforcement learning network is adopted to collect multi-source time series data through sensors on the robot for pre-processing, build the original network and twin network, and perform path planning and precise target planning to ensure that the robot can accurately reach the target position and complete maintenance actions.
It realizes precise path planning and operation of robots in complex underwater environments, improves operating accuracy and autonomous operation capabilities, and ensures the reliability and adaptability of robots in underwater tasks.
Smart Images

Figure CN119734283B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater robot path planning, and particularly to a path planning method for a cable-driven manipulator used in underwater robot operations. Background Art
[0002] An underwater cable-driven manipulator is an underwater intelligent device designed by imitating the functions of a human arm and wrist, aiming to complete complex underwater operation tasks such as detection, sampling, installation, and maintenance. In a complex underwater environment, it is easy to have errors in the position and attitude control of the end of the manipulator. Especially in the case of long cable transmission, problems such as tension and slip will lead to control deviations. In the face of dynamic water flow or complex terrain, the manipulator lacks sufficient environmental perception and adjustment capabilities. When underwater maintenance operations are required, the manipulator needs to accurately reach the location to be repaired.
[0003] Therefore, there is a need for a method that can enable the manipulator to achieve precise path planning and other tasks in a complex underwater environment, and improve the operation accuracy and autonomous operation ability of the underwater cable-driven manipulator. Summary of the Invention
[0004] The main purpose of the present invention is to provide a path planning method for a cable-driven manipulator used in underwater robot operations, so as to solve the problem of low operation accuracy in the path planning of underwater cable-driven manipulators in the prior art.
[0005] To achieve the above object, the present invention provides a path planning method for a cable-driven manipulator used in underwater robot operations, which specifically includes the following steps:
[0006] S1. Collect multi-source time series data by using sensors mounted on the manipulator.
[0007] S2. Preprocess the multi-source time series data, and the preprocessing includes: cleaning, formatting, normalizing, and standardizing.
[0008] S3. Construct an adaptive multi-channel deep reinforcement learning network, and the adaptive multi-channel deep reinforcement learning network includes: an origin network and a twin network connected to the origin network.
[0009] S4. Initialize the first experience replay module in the origin network, initialize the random noise, and obtain environmental information by using sensors, and transmit the environmental information to the adaptive multi-channel deep reinforcement learning network. The adaptive multi-channel deep reinforcement learning network performs path planning according to the environmental information, so that the manipulator reaches the target position.
[0010] S5. After the manipulator reaches the target position, continue to use the adaptive multi-channel deep reinforcement learning network for precise target planning to realize the repair action of the manipulator.
[0011] Furthermore, the origin network includes: a first experience replay module, an Actor1 network, a Critic network, and an Actor2 network connected in sequence; the Actor1 network includes: a first intelligent optimization module, a first online policy network, and a first target policy network connected in sequence; the Critic network includes: a first target Q network, a first online Q network, a first advantage evaluation and regularization module, and a first Adma optimizer connected in sequence; the Actor2 network includes: a second intelligent optimization module, a second online policy network, and a second target policy network; the first online Q network is connected to the first intelligent optimization module through a first AC interaction module, and the first online Q network is also connected to the second intelligent optimization module through a second AC interaction module.
[0012] The twin network includes: a second experience replay module, an Actor1' network, a Critic' network, and an Actor2' network connected in sequence; the Actor1' network includes: a third intelligent optimization module, a third online policy network, and a third target policy network connected in sequence; the Critic' network includes: a second target Q network, a second online Q network, a second advantage evaluation and regularization module, and a second Adma optimizer connected in sequence; the Actor2' network includes: a fourth intelligent optimization module, a fourth online policy network, and a fourth target policy network connected in sequence; the second online Q network is connected to the third intelligent optimization module through a third AC interaction module, and the second online Q network is also connected to the fourth intelligent optimization module through a fourth AC interaction module.
[0013] Furthermore, step S4 specifically includes the following steps:
[0014] S4.1, Transmit the environmental information to the first online policy network of the Actor1 network. The first online policy network processes the current state and generates an action :
[0015] ;
[0016] wherein, represents the first online policy network, is the current state, represents the first online policy network parameters.
[0017] S4.2, The Actor1 network transmits the action and the current state to the first online Q network of the Critic network. The first online Q network evaluates the current action and generates a decision value ;
[0018] ;
[0019] wherein, Represents the first online Q network, which are the parameters of the first online Q network.
[0020] S4.3, The Actor1 network simultaneously transmits the action to the environment module and obtains the environmental information at the t-th time step. The environment module includes: physical scene, task requirements, and external disturbances.
[0021] S4.4, The environment module transmits to the first target policy network of the Actor1 network. The first target policy network of the Actor1 network calculates and generates the corresponding action at the t-th time step;
[0022] ;
[0023] where, represents the first target policy network, which are the parameters of the first target policy network.
[0024] S4.5, The environment module simultaneously transmits the multi-dimensional array to the first experience replay module. After being judged by the first experience replay module, the multi-dimensional array is transmitted to the path experience replay pool in the first experience replay module; the multi-dimensional array includes: the current state , the environmental information at the t-th time step, the immediate reward and the action .
[0025] S4.6, Randomly extract a group from the multi-dimensional arrays in the path experience pool and transmit them to the Actor1 network and the Critic network respectively. The first target policy network of Actor1 transmits to the first target Q network of Critic. The first target Q network evaluates and generates the learning value ;
[0026] ;
[0027] where, represents the first target Q network, is the state at the t-th time step, is the action taken, which are the parameters of the first target Q network.
[0028] Furthermore, step S4 further includes: S4.7, updating the parameters of the Critic network; step S4.7 specifically includes the following steps:
[0029] S4.7.1, the first online Q-network calculates the target value : :
[0030] ;
[0031] where, is the target value, is the immediate reward, is the discount factor, is the action output by the first online policy network at ;
[0032] S4.7.2, calculate the advantage function of the Critic network :
[0033] ;
[0034] is the estimated value of the first online Q-network taking action at state ; is the state value function:
[0035] ;
[0036] where, is the parameter set, represents the probability of taking action at state ; is the policy; is the estimated value of the online Q-network taking action at state ; is the parameter set.
[0037] S4.7.3, calculate the loss function of the first target Q-network :
[0038] ;
[0039] represents the batch size, represents the advantage function, represents the prediction value of taking action at state ; represents the sum of the absolute values of the advantage functions of all samples in the batch.
[0040] S4.7.4, For select L2 regularization:
[0041] ;
[0042] Wherein, is L2 regularization, is the regularization strength.
[0043] ;
[0044] Wherein, represents the overall loss function of the Critic network, is the loss function of the first online Q network.
[0045] S4.7.5, Calculate the regularization gradient :
[0046] .
[0047] S4.7.6, Add the regularization gradient to the original gradient:
[0048] + ;
[0049] Wherein, is the total gradient, is the gradient of the loss function, is the regularization gradient.
[0050] S4.7.7, Use the first Adam optimizer to optimize the Q network parameters .
[0051] Furthermore, step S4 further includes: S4.8, Optimize the Actor1 network parameters. Step S4.8 specifically includes the following steps:
[0052] S4.8.1, Pass and to the first AC interaction module for data processing to generate a reinforcement signal .
[0053] S4.8.2, Pass to the first intelligent optimization module for parameter optimization;
[0054] Calculate the loss function:
[0055] .
[0056] Calculate the gradient:
[0057] ;
[0058] is the loss function with respect to the learning rate gradient. is the gradient of the online policy network with respect to the learning rate gradient. is the output of the online policy network in the state below, is variance.
[0059] S4.8.3, Update the Actor1 network parameters:
[0060] ;
[0061] wherein, is the learning rate.
[0062] Furthermore, step S4 further includes the following steps:
[0063] S4.9, Every time S4.1 to S4.8 are looped once, calculate the distance from the manipulator to the target point in the environment module. If the distance is less than D, where D is a positive number, it means that the data in Actor1 and Critic at this time is excellent data.
[0064] S4.10, Copy the excellent data layer by layer and copy them into the corresponding layers of the Actor1' network and the Critic' network in the twin network respectively; after the excellent data is completely copied into the twin network, continue to calculate in the Actor1' network and the Critic' network in the twin network, and repeat S4.1 to S4.8; until the maximum number of updates is reached, determine whether the path planning is slow.
[0065] S4.11, If it is judged that the calculation is slow, discard the current data, transfer to the original network, and use the excellent data in S4.9 for calculation, repeating S4.1 to S4.10.
[0066] S4.14, If it is judged that the calculation is not slow, continue to use the data in the twin network for calculation, repeating S4.1 to S4.10 until the target location is reached.
[0067] Furthermore, step S5 is specifically:
[0068] Input the environmental information of the current time step into the Actor2 network and the Critic network in the origin network, and repeat steps S4.1 to S4.14; among them, the second intelligent optimization module in the Actor2 network corresponds to the first intelligent optimization module in the Actor1 network, the second online policy network corresponds to the first online policy network, the second target policy network corresponds to the first target policy, and the second AC interaction module corresponds to the first AC interaction module.
[0069] The present invention has the following beneficial effects:
[0070] The method provided by the present invention can accurately predict the specific posture, accurate position and complete motion trajectory of the underwater manipulator during the execution of tasks in a complex marine environment, ensuring the accuracy and reliability of the operation.
[0071] The present invention designs an adaptive multi-channel deep reinforcement learning network, which can focus more on the optimization of specific tasks, improve the learning efficiency, and at the same time can dynamically adjust the weights and parameters of the Actor1 network, Actor2 network, Actor1' network and Actor2' network according to different task requirements to adapt to different environmental changes and task requirements. Improve its generalization ability. By training the Actor1 network, Actor2 network, Actor1' network and Actor2' network respectively, the network can have stronger generalization ability on specific tasks and improve the adaptability and robustness in different scenarios.
[0072] Classify and manage the data stored in the experience replay pool. Classified calls can improve the sample efficiency. By specifically extracting specific types of experiences, such as those that lead to high rewards or failures, the agent can train more frequently in critical situations, thus learning important policy knowledge faster and avoiding wasting computing resources on less critical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0074] Figure 1 Shows a flowchart of a path planning method for a cable-driven manipulator for underwater robot operations of the present invention.
[0075] Figure 2 Shows a structural diagram of the origin network of the present invention.
[0076] Figure 3 The structural diagram of the twin network of the present invention is shown.
[0077] Figure 4 The training effects of the method provided by the present invention at different stages are shown.
[0078] Figure 5 The simulation diagram of the path planning of the cable-driven manipulator using the method provided by the present invention is shown. Detailed implementation manners
[0079] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0080] As Figure 1 shown, a path planning method for a cable-driven manipulator for underwater robot operation specifically includes the following steps:
[0081] S1. Collect multi-source time series data using the sensors mounted on the manipulator; the sensors include a speedometer, an accelerometer, a sonar system, and a pressure sensor, and the collected multi-source time series data includes position, speed, acceleration, and the pressure received.
[0082] S2. Preprocess the multi-source time series data, and the preprocessing includes: cleaning, formatting, normalizing, and standardizing.
[0083] S3. Construct an adaptive multi-channel deep reinforcement learning network, and the adaptive multi-channel deep reinforcement learning network includes: a source network and a twin network connected to the source network.
[0084] S4. Initialize the first experience replay module in the source network, initialize the random noise, and obtain environmental information using the sensors, and transmit the environmental information to the adaptive multi-channel deep reinforcement learning network. The adaptive multi-channel deep reinforcement learning network performs path planning according to the environmental information to make the manipulator reach the target position.
[0085] S5. After the manipulator reaches the target position, continue to use the adaptive multi-channel deep reinforcement learning network to perform precise target planning to realize the maintenance action of the manipulator.
[0086] Specifically, as Figure 2As shown in the figure, the origin network includes: a first experience replay module, an Actor1 network, a Critic network, and an Actor2 network connected in sequence; the Actor1 network includes: a first intelligent optimization module, a first online policy network, and a first target policy network connected in sequence; the Critic network includes: a first target Q network, a first online Q network, a first advantage evaluation and regularization module, and a first Adma optimizer connected in sequence; the Actor2 network includes: a second intelligent optimization module, a second online policy network, and a second target policy network; the first online Q network is connected to the first intelligent optimization module through a first AC interaction module, and the first online Q network is also connected to the second intelligent optimization module through a second AC interaction module.
[0087] As Figure 3 shown in the figure, the twin network includes: a second experience replay module, an Actor1' network, a Critic' network, and an Actor2' network connected in sequence; the Actor1' network includes: a third intelligent optimization module, a third online policy network, and a third target policy network connected in sequence; the Critic' network includes: a second target Q network, a second online Q network, a second advantage evaluation and regularization module, and a second Adma optimizer connected in sequence; the Actor2' network includes: a fourth intelligent optimization module, a fourth online policy network, and a fourth target policy network connected in sequence; the second online Q network is connected to the third intelligent optimization module through a third AC interaction module, and the second online Q network is also connected to the fourth intelligent optimization module through a fourth AC interaction module.
[0088] First, path planning is performed on the manipulator in a large range. In this process, it mainly involves the Actor1 network, the Critic network, the first AC interaction module, the second AC interaction module, and the first experience replay module in the origin network, as well as the Actor1' network, the Critic' network, the third AC interaction module, the fourth AC interaction module, and the second experience replay module in the twin network; the structures of the twin network and the origin network are the same.
[0089] The first AC interaction module is mainly used to receive the parameters of the Actor1 network and the Critic network, perform processing and calculation, and then transfer them to the Actor1 network for updating the parameters of the Actor1 network; the second AC interaction module is the same.
[0090] Specifically, step S4 specifically includes the following steps:
[0091] S4.1, transfer the environmental information to the first online policy network of the Actor1 network, and the first online policy network processes the current state to generate an action :
[0092] ;
[0093] Among them, represents the first online policy network, is the current state, represents the parameters of the first online policy network.
[0094] S4.2, the Actor1 network passes the action and the current state to the first online Q-network of the Critic network. The first online Q-network evaluates the current action and generates a decision value ;
[0095] ;
[0096] Among them, represents the first online Q-network, are the parameters of the first online Q-network.
[0097] S4.3, the Actor1 network simultaneously passes the action to the environment module and obtains the environmental information at the time step. The environment module includes: physical scene, task requirements, and external disturbances; the physical scene is the underwater environment where the manipulator is located; the task requirement is to perform path planning for the manipulator to reach the repair location; the external disturbance is underwater noise, etc.
[0098] S4.4, the environment module passes to the first target policy network of the Actor1 network. The first target policy network of the Actor1 network calculates and generates the corresponding action at the time step
[0099] ;
[0100] Among them, represents the first target policy network, are the parameters of the first target policy network.
[0101] S4.5, the environment module simultaneously passes a multi-dimensional array to the first experience replay module. After being judged by the first experience replay module, the multi-dimensional array is passed to the path experience replay pool in the first experience replay module; the multi-dimensional array includes: the current state , the environmental information at the time step, the immediate reward and the action .
[0102] The first and second experience replay modules are each composed of two experience replay pools, namely the path experience replay pool and the maintenance experience replay pool. After the first and second experience replay modules receive external information data, they first judge the current manipulator state. If the target point has not been reached, the multi-dimensional array is stored in the path experience replay pool; if the target point has been reached, the multi-dimensional array is stored in the maintenance experience replay pool.
[0103] S4.6, randomly extract a group of multi-dimensional arrays from the path experience pool and pass them to the Actor1 network and the Critic network respectively. The first target policy network of Actor1 will pass to the first target Q network of Critic, and the first target Q network will evaluate and generate a learning value ;
[0104] ;
[0105] Among them, represents the first target Q network, is the state at the time step, is the action taken, is the parameter of the first target Q network.
[0106] Specifically, step S4 also includes: S4.7, update the Critic network parameters; step S4.7 specifically includes the following steps:
[0107] S4.7.1, the first online Q network calculates the target value according to :
[0108] ;
[0109] Among them, is the target value, is the immediate reward, is the discount factor, used to balance the immediate reward and future rewards, is the target Q network, used to evaluate the value of the action given by the policy network in the next state , is the action output by the first online policy network in ;
[0110] S4.7.2, calculate the advantage function of the Critic network:
[0111] ;
[0112] Among them, is the estimated value of the first online Q-network taking an action in state ; is the state-value function:
[0113] ;
[0114] where is a set of parameters determined by the set of parameters ; represents the probability of taking an action in state ; is the policy; determined by the policy ; is the estimated value of the online Q-network taking an action in state ; is a set of parameters; determined by the set of parameters ;
[0115] S4.7.3. Calculate the loss function of the first target Q-network :
[0116] ;
[0117] is a weighted mean squared error used to measure the difference between the predicted value and the target value; represents the batch size, i.e., the number of samples used to update the network. represents the advantage function, indicating the advantage of taking an action in state relative to a random action. represents the prediction of taking an action in state calculated from the online Q-network parameters ; represents the sum of the absolute values of the advantage functions of all samples in the batch; used as a normalization factor to ensure that the sum of the weights is 1.
[0118] S4.7.4. Select L2 regularization for :
[0119] ;
[0120] where is the L2 regularization, which is based on the sum of the squares of the model parameters ; is the regularization strength, which is a hyperparameter used to control the relative importance of the regularization term in the total loss function. The larger the value of, the heavier the penalty on the model complexity. are the online Q-network parameters, which can be the weights w or the biases b. This is for all The sum of squares is calculated. For , its square is added to the sum.
[0121] ;
[0122] represents the overall loss function of the Critic network. is the loss function of the first online Q-network.
[0123] S4.7.5, Calculate the regularization gradient :
[0124] .
[0125] S4.7.6, Add the regularization gradient to the original gradient:
[0126] + ;
[0127] where, is the total gradient, is the gradient of the loss function, is the regularization gradient.
[0128] S4.7.7, Use the first Adam optimizer to optimize the Q-network parameters .
[0129] Specifically, step S4 also includes: S4.8, Optimize the Actor1 network parameters. Step S4.8 specifically includes the following steps:
[0130] S4.8.1, Pass and to the first AC interaction module for data processing to generate the reinforcement signal .
[0131] The first AC interaction module has two inputs, which are the actions of Actor1 and ;
[0132] The data is processed in the first AC interaction module as follows:
[0133] Pass two inputs to the first hidden layer of the first AC interaction module, concatenate the two inputs, and generate an output;
[0134] Pass the output of the first hidden layer to the second hidden layer of the first AC interaction module for further processing and generate an output;
[0135] Pass the output of the second hidden layer to the fusion layer of the first AC interaction module, concatenate the outputs of the two hidden layers, and generate an output;
[0136] Pass the output of the fusion layer to the output layer of the first AC interaction module to finally generate a reinforcement signal 。
[0137] The first AC interaction module is composed of a neural network. The adaptive multi-channel deep reinforcement learning network added to the first AC interaction module can be optimized in aspects such as information fusion, learning stability, policy adaptability, exploration efficiency, sample efficiency, and generalization ability, thus showing more excellent performance in reinforcement learning tasks.
[0138] S4.8.2, pass to the first intelligent optimization module for parameter optimization.
[0139] The first intelligent optimization module is mainly used to update the parameters of the Actor1 network. The first intelligent optimization module receives the output from the first AC interaction module , calculates the loss function of the Actor1 network, and after calculating the gradient, is used to update the parameters, increasing the efficiency and accuracy of network update and exploration.
[0140] Calculate the loss function:
[0141] ;
[0142] Calculate the gradient:
[0143] ;
[0144] Among them, is the loss function with respect to the learning rate . The gradient is a vector that contains the contribution of each parameter to the loss function and is used to guide the update of the parameters. is the output of the above AC interaction module; is the gradient of the online policy network with respect to the learning rate . This gradient measures the impact of parameter changes on the policy output. is the output of the online policy network in the state , is The variance. Variance measures the dispersion of a probability distribution and is commonly used to calculate Gaussian noise for the gradient. is the variance of the difference between the output of the online policy network and the actual action. This difference measures the deviation between the prediction of the online policy network and the actual execution.
[0145] S4.8.3, Update the parameters of the Actor1 network:
[0146] ;
[0147] where is the learning rate, which controls the step size of parameter update, and is the gradient of the loss function of the Actor1 network.
[0148] Specifically, step S4 also includes the following steps:
[0149] S4.9, Every time S4.1 to S4.8 are looped once, calculate the distance from the manipulator to the target point in the environment module. If the distance is less than D, where D is a positive number, it means the data in Actor1 and Critic at this time are excellent data; the distance from the manipulator to the target point is stored in the task requirements in the environment module.
[0150] S4.10, Copy the excellent data layer by layer and copy them to the corresponding layers of the Actor1' network and the Critic' network in the twin network respectively; after the excellent data are completely copied to the twin network, continue to calculate in the Actor1' network and the Critic' network in the twin network, repeating S4.1 to S4.8; until the maximum number of updates is reached, judge whether the path planning is slow.
[0151] S4.11, If it is judged that the calculation is slow, discard the current data, transfer to the original network, and use the excellent data in S4.9 for calculation, repeating S4.1 to S4.10.
[0152] S4.14, If it is judged that the calculation is not slow, continue to use the data in the twin network for calculation, repeating S4.1 to S4.10 until reaching the target location.
[0153] Specifically, step S5 is specifically as follows:
[0154] Input the environmental information of the current time step into the Actor2 network and the Critic network in the original network, repeating steps S4.1 to S4.14; among them, the second intelligent optimization module in the Actor2 network corresponds to the first intelligent optimization module in the Actor1 network, the second online policy network corresponds to the first online policy network, the second target policy network corresponds to the first target policy network, and the second AC interaction module corresponds to the first AC interaction module.
[0155] In step S5, the manipulator conducts path planning within a small range, mainly responsible for planning the path of the end of the manipulator near the maintenance point to assist the end effector in achieving the maintenance action. During this process, it mainly involves the Actor2 network, Critic network, first AC interaction module, second AC interaction module, and first experience replay module in the origin network, as well as the Actor2' network, Critic' network, third AC interaction module, fourth AC interaction module, and experience replay module in the twin network; the structures of the twin network and the origin network are the same.
[0156] Figure 4 It shows the training effects of the deep reinforcement learning algorithm at different stages, including the initial stage, the middle stage, and the late stage. Figure 4 The original return in [reference] is coherent noise. In the middle stage, it shows a steady increase; while in the late stage, a better strategy is found and the performance tends to be stable. This proves the effectiveness and reliability of the method provided by the present invention in path planning tasks.
[0157] As Figure 5 shown, the motion trajectory of the robotic arm clearly demonstrates its ability to avoid circular obstacles. Starting from the starting point, the robotic arm skillfully bypasses the obstacles under the guidance of the deep reinforcement learning algorithm. Throughout the movement process, the robotic arm always maintains a stable speed and precise control, ensuring the safety and efficiency of the movement. Until the robotic arm successfully reaches the end point and completes the predetermined task.
[0158] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or replacements made by those skilled in the art within the substantial scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for path planning of a rope-driven manipulator for underwater robot operation, characterized in that: The specific steps include: S1, collects multi-source time series data using sensors mounted on the manipulator; S2, preprocessing of multi-source time series data, including cleaning, formatting, normalization and normalization; S3, constructing an adaptive multi-channel deep reinforcement learning network, which includes: an original network and a twin network connected to the original network; S4, initialize the first experience replay module in the source network, initialize the random noise, and use the sensor to obtain environmental information, and pass the environmental information to the adaptive multi-channel deep reinforcement learning network. The adaptive multi-channel deep reinforcement learning network performs path planning based on the environmental information to enable the manipulator to reach the target position; S5, when the manipulator reaches the target position, it continues to use the adaptive multi-channel deep reinforcement learning network for precise target planning to achieve the maintenance action of the manipulator; The original network includes: a first experience playback module, an Actor1 network, a Critic network, and an Actor2 network connected in sequence; the Actor1 network includes: a first intelligent optimization module, a first online strategy network, and a first target strategy network connected in sequence; the Critic network includes: a first target Q network, a first online Q network, a first advantage evaluation and regularization module, and a first Adma optimizer connected in sequence; the Actor2 network includes: a second intelligent optimization module, a second online strategy network, and a second target strategy network; the first online Q network is connected to the first intelligent optimization module through a first AC interaction module, and the first online Q network is also connected to the second intelligent optimization module through a second AC interaction module; The twin network includes: a second experience playback module, an Actor1' network, a Critic' network and an Actor2' network connected in sequence; the Actor1' network includes: a third intelligent optimization module, a third online strategy network and a third target strategy network connected in sequence; the Critic' network includes: a second target Q network, a second online Q network, a second advantage evaluation and regularization module and a second Adma optimizer connected in sequence; the Actor2' network includes: a fourth intelligent optimization module, a fourth online strategy network and a fourth target strategy network connected in sequence; the second online Q network is connected to the third intelligent optimization module through the third AC interaction module, and the second online Q network is also connected to the fourth intelligent optimization module through the fourth AC interaction module; Step S4 specifically includes the following steps: S4.1, passing the environmental information to the first online policy network of the Actor 1 network, the first online policy network processes the current state and generates an action : ; in, Represents the first online strategy network, is the current state, represents the first online strategy network parameters; S4.2, Actor1 network will act and current status The first online Q network passed to the Critic network evaluates the current action and generates a decision value ; ; in, Representing the First Online Q Network, are the parameters of the first online Q network; S4.3, Actor1 network simultaneously takes action Pass it to the environment module to get the Environmental information at a time step ,The environment module includes: physical scenes, task requirements and external disturbances; S4.4, the environment module will The first target strategy network of Actor1 network is passed to the first target strategy network of Actor1 network, which calculates and generates the corresponding Actions in time steps ; ; in, represents the first target strategy network, are the parameters of the first target policy network; S4.5, the environment module simultaneously passes the multidimensional array to the first experience replay module, and then after the first experience replay module makes a judgment, passes the multidimensional array to the path experience replay pool in the first experience replay module; the multidimensional array includes: the current state , No. Environmental information at a time step , instant rewards and actions ; S4.6, randomly select a group of multidimensional arrays from the path experience pool through the sampling strategy, and pass them to the Actor1 network and the Critic network respectively. The first target strategy network of Actor1 will The first target Q network passed to Critic, the first target Q network Evaluate and generate learning values ; ; in, represents the first target Q network, It is The state of the time step, is the action taken, are the parameters of the first target Q network; Step S4 also includes: S4.7, updating the Critic network parameters; Step S4.7 specifically includes the following steps: S4.7.1, the first online Q network according to Calculate target value : ; in, is the target value, It’s an instant reward. is the discount factor, Is the first online strategy network in The action of outputting S4.7.2, Calculating the Advantage Function of the Critic Network : ; Is the first online Q network in the state Take action The estimated value of is the state value function: ; in, is the parameter set, Indicates in status Take action The probability of For strategy; For online Q network in state Take action The estimated value of is the parameter set; S4.7.3, calculate the loss function of the first target Q network : ; Indicates the batch size, represents the advantage function, Indicates in status Take action Prediction value, Represents the sum of the absolute values of the advantage functions of all samples in the batch; S4.7.4, Yes Choose L2 regularization: ; in, is L2 regularization, is the regularization strength; ; in, represents the overall loss function of the Critic network, is the loss function of the first online Q network; S4.7.5, Calculate the regularization gradient : ; S4.7.6, add the regularized gradient to the original gradient: + ; in, is the total gradient, is the gradient of the loss function, is the regularization gradient; S4.7.7, Optimize Q network parameters using the first Adam optimizer ; Step S4 also includes: S4.8, optimizing the network parameters of Actor 1, and step S4.8 specifically includes the following steps: S4.8.1, and The data is transmitted to the first AC interaction module for data processing to generate an enhanced signal. ; S4.8.2, The data are transmitted to the first intelligent optimization module for parameter optimization; Calculate the loss function: ; Compute the gradient: ; is the loss function about The gradient of Is an online strategy network about The gradient of Is an online strategy network in the state The output of the following, yes The variance of S4.8.3, update Actor1 network parameters: ; in, is the learning rate; Step S4 also includes the following steps: S4.9, each time S4.1 to S4.8 are cycled once, the distance from the robot in the environment module to the target point is calculated. If the distance is less than D, and D is a positive number, it means that the data in Actor1 and Critic at this time are excellent data; S4.10, copy the excellent data layer by layer, and copy them into the corresponding layers of the Actor1' network and the Critic' network in the twin network respectively; after the excellent data is completely copied to the twin network, the Actor1' network and the Critic' network in the twin network continue to calculate, and repeat S4.1 to S4.8; until the maximum number of updates is reached, determine whether the path planning is slow; S4.11, if it is judged that the calculation is slow, then discard the current data, switch to the original network, use the excellent data in S4.9 to calculate, and repeat S4.1 to S4.10; S4.14, if it is judged that the calculation is not slow, continue to use the data in the twin network for calculation, and repeat S4.1 to S4.10 until the target location is reached.
2. A method for path planning of a rope-driven manipulator for underwater robot operation according to claim 1, characterized in that: Step S5 is specifically: inputting the environmental information of the current time step into the Actor2 network and the Critic network in the source network to repeat steps S4.1 to S4.14; wherein, the second intelligent optimization module in the Actor2 network corresponds to the first intelligent optimization module in the Actor1 network, the second online strategy network corresponds to the first online strategy network, the second target strategy network corresponds to the first target strategy network, and the second AC interaction module corresponds to the first AC interaction module.
Citation Information
Patent Citations
Robot path navigation method and system based on deep reinforcement learning
CN111487864A
Amphibious vehicle path planning method and system based on fuzzy logic rule and reinforcement learning
CN119512108A
Robot motion planning method and device based on multi-modal information fusion
CN119550335A