Target tracking control method and device for unmanned ship based on deep reinforcement learning

Through the unmanned boat target tracking and control method based on deep reinforcement learning, the target control model is trained using multimodal state information and reward function, and the servo angle and propeller speed instructions are output, the problem of low success rate of unmanned boat target tracking is solved, and precise control and efficient tracking are achieved.

CN120491646APending Publication Date: 2025-08-15CHINA SHIP SCIENTIFIC RESEARCH CENTER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510618830.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing unmanned boat target tracking algorithm has low control accuracy, resulting in low success rate of target ship tracking and poor adaptability and practicality.

Method used

The target tracking control method based on deep reinforcement learning is adopted. By obtaining the multimodal state information of the unmanned boat and the target ship, the pre-trained target control model outputs control instructions for the servo angle and propeller speed, combined with the reward function and SAC algorithm for training, a multimodal state space is constructed, and control instructions are generated to achieve accurate tracking of the unmanned boat.

Benefits of technology

The success rate of unmanned boat tracking target ships is improved, the adaptability and practicality of the model is enhanced, and the target control model completed is efficient in the training and has fast real-time response speed, which solves the problem of low control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491646A_ABST
    Figure CN120491646A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned ship target tracking control method and device based on deep reinforcement learning, and relates to the technical field of unmanned ship target tracking, and the method comprises the steps: obtaining to-be-used data; constructing a multi-modal state space based on the to-be-used data; and inputting the multi-modal state space into a pre-trained target control model to obtain a control instruction which is output by the target control model and carries the steering engine rotation angle and the propeller rotation speed. The method is used for solving the problem that the target ship tracking success rate is low due to the fact that the control accuracy of the unmanned ship is low in the prior art, the unmanned ship is accurately controlled, and the target ship tracking success rate of the unmanned ship is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of unmanned boat target tracking, and in particular to a target tracking control method and device for an unmanned boat based on deep reinforcement learning. Background Art

[0002] When performing patrol reconnaissance, fire strikes, and water law enforcement tasks, unmanned boats need to obtain information about target ships in order to achieve the purpose of continuous tracking.

[0003] Existing target tracking algorithms used by unmanned vehicles (UAVs) are based on traditional control theory, such as pure tracking algorithms, LQR algorithms, and MPC algorithms. However, the control effects of these algorithms are significantly affected by parameters, requiring frequent and difficult parameter adjustments. Furthermore, these algorithms rarely consider the characteristics of UAVs (such as complex motion constraints, limited steering capabilities, and limited power system flexibility). As a result, these algorithms have poor adaptability and practicality for UAVs, resulting in low control accuracy and a high risk of target vessel tracking failure.

[0004] It can be seen that the existing target tracking algorithm cannot meet the target tracking requirements of unmanned boats. Therefore, how to solve the target tracking problem of unmanned boats is an important issue that needs to be solved urgently in the industry. Summary of the Invention

[0005] In response to the above-mentioned problems and technical needs, the applicant has proposed a target tracking control method and equipment for an unmanned boat based on deep reinforcement learning, so as to solve the problem of low target ship tracking success rate caused by low control accuracy of the unmanned boat in the existing technology, realize precise control of the unmanned boat, and improve the success rate of the unmanned boat tracking the target ship.

[0006] The present invention provides a method for target tracking and control of an unmanned boat based on deep reinforcement learning, the method comprising:

[0007] Acquiring data to be used, wherein the data to be used includes: first state information of the unmanned boat, relative state information between the unmanned boat and the target ship being tracked by the unmanned boat, ocean environment information corresponding to the unmanned boat and the target ship, and obstacle information;

[0008] Construct a multimodal state space based on the data to be used;

[0009] Inputting the multimodal state space into a pre-trained target control model to obtain a control instruction output by the target control model that carries a steering gear angle and a propeller speed;

[0010] The target control model includes an actor network and a critic network. The target control model is trained based on multimodal state space samples and control instruction samples in combination with a preset reward function and a SAC algorithm.

[0011] The control instruction is used to control the unmanned boat to continuously track the target ship, and the reward function is obtained based on the first state information, the relative state information, the ocean environment information and the obstacle information.

[0012] According to the target tracking control method of the unmanned boat based on deep reinforcement learning provided by the embodiment of the present application, the reward function includes:

[0013]

[0014] Among them, R t represents the reward value, ω1 represents the distance preservation reward weight coefficient, k1 represents the weight corresponding to the distance in the distance preservation reward, d t Indicates the relative distance between the target ship and the unmanned boat, d opt represents the optimal tracking distance, k2 represents the weight of speed in the distance keeping reward, v rel represents the relative speed between the unmanned boat and the target ship, ω2 represents the weight coefficient of the field of view maintenance reward, represents the azimuth deviation, ω3 represents the obstacle avoidance reward weight coefficient, N represents the number of obstacles, represents the shortest distance from the unmanned boat to the i-th obstacle, ε represents the preset minimum value, ω4 represents the energy consumption optimization reward weight coefficient, and n represents the rotation speed of the unmanned boat.

[0015] According to the target tracking control method of an unmanned watercraft based on deep reinforcement learning provided by an embodiment of the present application, before inputting the multimodal state space into a pre-trained target control model and obtaining the control instructions output by the target control model that carry the steering gear angle and propeller speed, the method further includes:

[0016] The actor network of the target control model is used to interact with the simulation environment to generate simulation data samples, wherein the simulation data samples include a tag value, and the tag value is used to indicate whether the unmanned boat successfully tracks the target ship;

[0017] Based on the tag value, the simulation data sample is stored in the success experience pool or the failure experience pool, and a sample priority is configured for each simulation data sample in the success experience pool and the failure experience pool;

[0018] Extracting simulation data samples from the success experience pool and the failure experience pool based on a hybrid sampling method and sample priority, and determining the extracted simulation data samples as training sample data, wherein the training sample data includes multimodal state space samples and control instruction samples;

[0019] The target control model is trained based on training sample data, preset reward function and SCA algorithm.

[0020] According to the target tracking control method of an unmanned boat based on deep reinforcement learning provided by an embodiment of the present application, a sample priority is configured for each simulation data sample in the success experience pool and the failure experience pool, including:

[0021] Configure sample priorities for simulation data samples in the successful experience pool based on a first weight calculation formula;

[0022] The first weight calculation formula includes:

[0023]

[0024] Among them, ω i represents the sample priority of the i-th simulation data sample in the successful experience pool, Q target,i Represents the Q value calculated by the target value network for the i-th simulation data sample, Q current,i represents the Q value calculated by the current value network for the i-th simulation data sample, and ε represents the preset minimum value;

[0025] configuring sample priorities for simulation data samples in the failure experience pool based on a second weight calculation formula;

[0026] The second weight calculation formula includes:

[0027] ω j =|R t -Q current,j |;

[0028] Among them, ω j represents the sample priority of the jth simulation data sample in the failure experience pool, Q current,j Represents the Q value calculated by the current value network for j simulation data samples, R t Indicates the reward value.

[0029] According to the target tracking control method of an unmanned boat based on deep reinforcement learning provided by an embodiment of the present application, simulation data samples are extracted from the success experience pool and the failure experience pool respectively based on a hybrid sampling method and sample priority, including:

[0030] Determine the sample extraction ratio of the successful experience pool based on the sample priority of each simulation data sample in the successful experience pool and the failed experience pool and a preset ratio calculation formula;

[0031] The ratio calculation formula includes:

[0032]

[0033] Among them, β represents the sample extraction ratio of the successful experience pool, 1-β represents the sample extraction ratio of the failed experience pool, P represents the number of simulation data samples in the successful experience pool, and Q represents the number of simulation data samples in the failed experience pool.

[0034] According to the target tracking control method of an unmanned boat based on deep reinforcement learning provided by an embodiment of the present application, the actor network includes a policy network, and the critic network includes: two value networks and two target value networks;

[0035] The strategy network is connected to the value network, and a value network is connected to a target value network;

[0036] The training process of the target control model includes:

[0037] Obtain multimodal state space samples, control instruction samples, and reward functions, where the state space samples are temporally sequential.

[0038] The multimodal state space samples, control instruction samples and reward function are input into the target control model. Based on the SAC algorithm, the value network and the policy network are used to obtain the first Q value based on the reward value output by the reward function, and the target value network and the policy network are used to obtain the second Q value based on the reward value output by the reward function; the mean square error of the first Q value and the second Q value is calculated, and the mean square error is used to perform a gradient descent operation on the first network parameter of the value network; the preset policy loss function is used to perform a gradient ascent operation on the second network parameter of the policy network; and the preset soft update mechanism is used to optimize the third network parameter of the target value network to obtain the final target control model.

[0039] According to the target tracking control method of an unmanned boat based on deep reinforcement learning provided by an embodiment of the present application, a first Q value is obtained based on a reward value output by a value network and a policy network, and a second Q value is obtained based on a reward value output by a reward function by using a target value network and a policy network, including:

[0040] The policy network generates an action probability distribution corresponding to the predicted action based on the multimodal state space samples;

[0041] The target value network and value grid respectively obtain the Q value based on the action probability distribution, the reward value obtained by the reward function, and the Q value calculation formula;

[0042] The Q value calculation formula includes:

[0043] Q(s′,a′)=R t +γ(min i=1,2 Q(s′,a′)-k3logπ φ (a′|s′));

[0044] Among them, R t represents the reward value, γ represents the reward discount factor, which is a constant, Q(s′,a′) represents the Q value corresponding to the i-th value network or the i-th target value network, k3 represents the temperature coefficient, which is a constant, π φ (a′|s′) represents the action probability distribution, a′ represents the control instruction sample at the next moment, and s′ represents the state space sample at the next moment;

[0045] The Q value includes a first Q value and a second Q value.

[0046] According to the target tracking control method of an unmanned boat based on deep reinforcement learning provided by an embodiment of the present application, a gradient descent operation is performed on the first network parameter of the value network using the mean square error, including:

[0047] Inputting the mean square error into a predetermined gradient descent calculation formula to obtain the optimized first network parameters output by the gradient descent calculation formula;

[0048] Among them, the gradient descent calculation formula includes:

[0049]

[0050] Among them, θ i represents the first network parameter of the i-th value network, η Q represents the learning rate of the value network, represents the gradient of the first network parameter of the i-th value network, L Q (θ i ) represents the mean square error.

[0051] According to the target tracking control method of the unmanned boat based on deep reinforcement learning provided by the embodiment of the present application, a gradient ascent operation is performed on the second network parameter of the policy network using a preset policy loss function, including:

[0052] Use the policy loss function to calculate the prediction loss of the policy network;

[0053] Among them, the strategy loss function includes:

[0054]

[0055] Among them, L π (φ) represents the predicted loss, k3 represents the temperature coefficient, which is a constant, logπ φ (a|s) represents the entropy regularization term, represents the minimized Q value, a represents the control instruction sample at the current moment, and s represents the state space sample at the current moment;

[0056] Input the predicted loss into a predetermined gradient ascent calculation formula to obtain the optimized second network parameters output by the gradient ascent calculation formula;

[0057] Among them, the gradient ascent calculation formula includes:

[0058]

[0059] Among them, φ i Represents the second network parameter of the policy network, η π represents the learning rate of the policy network, represents the gradient of the second grid parameter.

[0060] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the target tracking control method for an unmanned boat based on deep reinforcement learning as described above are implemented.

[0061] An embodiment of the present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the target tracking control method of an unmanned boat based on deep reinforcement learning as described above are implemented.

[0062] The target tracking control method and device of an unmanned boat based on deep reinforcement learning provided in the embodiment of the present application obtains a state space by obtaining first dynamic information based on the unmanned boat, second dynamic information of the target ship tracked by the unmanned boat, and ocean environment information corresponding to the unmanned boat and the target ship. It can be seen that the present application fully considers the operating dynamics of the unmanned boat, the operating dynamics of the target ship and the ocean-related information of the environment in which the two are located, and provides an effective data basis for controlling the unmanned boat to continuously and effectively track the target ship; then, the state space is input into a pre-trained target control model to obtain a control instruction carrying a steering gear angle and a propeller speed output by the target control model; the control instruction is used to control the unmanned boat to continuously track the target ship Tracking the target ship, the reward function is obtained based on the first dynamic information and the second dynamic information. The present application trains the target control model based on the SAC algorithm, the reward function based on the first dynamic information and the second dynamic information, and the sample data. It uses deep reinforcement learning to strengthen the model, and does not require frequent parameter adjustment in the subsequent model application process, which improves the adaptability of the target control model and its practicality in the field of unmanned boat tracking. In addition, the trained target control model has high computational efficiency and fast real-time response speed, which solves the problem of low target ship tracking success rate due to low control accuracy of the unmanned boat in the existing technology, realizes precise control of the unmanned boat, and improves the success rate of the unmanned boat in continuously tracking the target ship. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0064] Figure 1 This is one of the flow charts of the target tracking control method of an unmanned boat based on deep reinforcement learning provided in an embodiment of the present application;

[0065] Figure 2 This is a schematic diagram of a scenario in which an unmanned boat tracks a target ship, as provided in an embodiment of the present application;

[0066] Figure 3 This is the second flow chart of the target tracking control method of an unmanned boat based on deep reinforcement learning provided in an embodiment of the present application;

[0067] Figure 4 is a schematic structural diagram of the target control model provided in an embodiment of the present application;

[0068] Figure 5 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0069] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0070] The target tracking problem for unmanned vessels is an optimal control problem subject to multiple constraints, including motion constraints, obstacle constraints, energy consumption constraints, ocean wave constraints, and detection range constraints. To continuously and effectively track the target vessel, the unmanned vessel must maintain a certain distance from the target vessel while keeping the target vessel within its field of view.

[0071] In order to solve the above problems, the embodiment of the present application provides a target tracking control method for an unmanned boat based on deep reinforcement learning. The method can be applied to smart terminals, servers, and controllers of unmanned boats. This application uses the application of this method in the controller of an unmanned boat as an example to illustrate. This is an example and is not used to limit the scope of protection of this application. Some other descriptions in the embodiments are also examples and will not be explained one by one later. Figure 1As shown, the method includes:

[0072] Step 101: Obtain data to be used.

[0073] The data to be used include: first state information of the unmanned boat, relative state information between the unmanned boat and the target ship being tracked by the unmanned boat, ocean environment information corresponding to the unmanned boat and the target ship, and obstacle information.

[0074] The first state information includes: the position, heading angle, longitudinal speed, lateral speed and steering angular velocity of the unmanned boat, etc. The relative state information includes: the relative distance between the target ship and the unmanned boat, azimuth deviation and speed deviation, etc.

[0075] Specifically, the relative state information can be obtained by monitoring the second state information of the target ship, such as the position, heading angle, longitudinal speed, lateral speed and steering angular speed of the target ship, in combination with the first state information.

[0076] Among them, the ocean environment information includes: vector-represented ocean current velocity, wind field disturbance torque and effective wave height, etc.

[0077] The obstacle information includes the relative distance and direction angle between the unmanned boat and the obstacle.

[0078] Among them, the first state information of the unmanned boat is obtained based on the ship's inertial navigation equipment, and the second state information, marine environment information and obstacle information of the target ship are obtained using ship-borne radar and AIS equipment.

[0079] Specifically, since the unmanned boat and the target ship are dynamic, the data to be used are time-series.

[0080] Step 102: construct a multimodal state space based on the data to be used.

[0081] Step 103: Input the multimodal state space into the pre-trained target control model to obtain a control instruction output by the target control model that carries the steering gear angle and propeller speed.

[0082] Among them, the target control model includes an actor network and a critic network. The target control model is based on multimodal state space samples and control instruction samples, and is trained in combination with a preset reward function and SAC algorithm.

[0083] The control instructions are used to control the unmanned boat to continuously track the target ship, and the reward function is obtained based on the first state information, relative state information, ocean environment information and obstacle information.

[0084] Among them, the control instructions are equivalent to the action space of the unmanned boat, including the steering gear angle and propeller speed.

[0085] The steering gear angle is represented by δ, and its value range is δ∈[-δ max ,δ max ], the propeller speed is represented by n, and its value range is n∈[0,n max ]. Among them, δ max Indicates the maximum steering angle of the unmanned boat, n max Indicates the maximum propeller speed of the unmanned boat.

[0086] Specifically, the present application sets a value range for the servo angle and propeller speed, and limits the servo angle and propeller speed output by the target control model. If they exceed the corresponding value range, the limit value corresponding to the value range is used as the final servo angle and / or propeller speed.

[0087] Specifically, the present application predicts the action space at the next moment based on the multimodal state space corresponding to the current moment.

[0088] Among them, through Figure 2 The scenario of the unmanned boat tracking the target ship is illustrated.

[0089] Among them, Figure 2 In the figure, A represents the unmanned boat, B represents the target ship, C1, C2 and C3 represent different obstacles respectively. The largest sector corresponding to A is the detection area, and the circle corresponding to B represents the range corresponding to the safe driving distance of the target ship. It represents the angle between the line connecting the unmanned boat and the target ship and the center line of the field of view, d represents the straight-line distance between the unmanned boat and the target ship, and L represents the maximum detection range of the unmanned boat.

[0090] The target tracking control method of an unmanned boat based on deep reinforcement learning provided in an embodiment of the present application obtains a state space by obtaining first dynamic information based on the unmanned boat, second dynamic information of the target ship tracked by the unmanned boat, and ocean environment information corresponding to the unmanned boat and the target ship. It can be seen that the present application fully considers the operating dynamics of the unmanned boat, the operating dynamics of the target ship and the ocean-related information of the environment in which the two are located, and provides an effective data basis for controlling the unmanned boat to continuously and effectively track the target ship; then, the state space is input into a pre-trained target control model to obtain a control instruction carrying a steering gear angle and a propeller speed output by the target control model; the control instruction is used to control the unmanned boat to continuously track the target ship The target ship, the reward function is obtained based on the first dynamic information and the second dynamic information. The present application trains the target control model based on the SAC algorithm, the reward function based on the first dynamic information and the second dynamic information, and the sample data. It uses deep reinforcement learning to strengthen the model, and does not require frequent parameter adjustment during the subsequent model application process, thereby improving the adaptability of the target control model and its practicality in the field of unmanned boat tracking. In addition, the trained target control model has high computational efficiency and fast real-time response speed, which solves the problem of low target ship tracking success rate due to low control accuracy of the unmanned boat in the existing technology, realizes precise control of the unmanned boat, and improves the success rate of the unmanned boat in continuously tracking the target ship.

[0091] In a specific embodiment, the reward function is shown in formula (1):

[0092]

[0093] Among them, R t represents the reward value, ω1 represents the distance preservation reward weight coefficient, k1 represents the weight corresponding to the distance in the distance preservation reward, d t represents the relative distance between the target ship and the unmanned boat, dopt represents the optimal tracking distance, k2 represents the weight of speed in the distance keeping reward, v rel represents the relative speed between the unmanned boat and the target ship, ω2 represents the weight coefficient of the field of view maintenance reward, represents the azimuth deviation, ω3 represents the obstacle avoidance reward weight coefficient, N represents the number of obstacles, represents the shortest distance from the unmanned boat to the i-th obstacle, ε represents the preset minimum value, ω4 represents the energy consumption optimization reward weight coefficient, and n represents the rotation speed of the unmanned boat.

[0094] Among them, the azimuth deviation is the angle between the line connecting the unmanned boat and the target ship and the center line of the field of view.

[0095] Specifically, Used to represent the distance keeping reward, d optPreset the distance based on the actual tracking task. For example, the optimal tracking distance is determined by the minimum safe distance to the target ship and the maximum detection distance of the unmanned vehicle. Specifically, the median of the minimum safe distance and the maximum detection distance is used as the optimal tracking distance. The distance retention reward represents the tracking distance and the tracking distance change rate. The closer the tracking distance is to the optimal tracking distance and the smaller the tracking distance change rate, the greater the distance reward value.

[0096] Specifically, Used to represent the visual field maintenance reward. The closer the line is to the center line, the greater the reward value.

[0097] in, Used to represent obstacle avoidance rewards. The farther the UAV is from the obstacle, the greater the reward value.

[0098] Among them, |n| is used to represent the energy loss reward. The higher the speed, the more energy is consumed and the smaller the reward value.

[0099] In addition, the reward value for tracking failure is a preset value, and users can set it according to their actual needs.

[0100] In a specific embodiment, the multimodal state space is input into a pre-trained target control model, and before the target control model outputs a control instruction carrying the steering gear angle and propeller speed, the target control model is trained. For specific implementation, see Figure 3 :

[0101] Step 301: Use the actor network of the target control model and the simulation environment to interact and generate simulation data samples.

[0102] The simulation data sample includes a tag value, which is used to indicate whether the unmanned boat successfully tracks the target ship.

[0103] Step 302: based on the tag value, store the simulation data sample in a success experience pool or a failure experience pool, and configure a sample priority for each simulation data sample in the success experience pool and the failure experience pool.

[0104] Step 303 : extracting simulation data samples from the success experience pool and the failure experience pool respectively based on the hybrid sampling method and the sample priority, and determining the extracted simulation data samples as training sample data.

[0105] The training sample data includes multimodal state space samples and control instruction samples.

[0106] Step 304 : training the target control model based on the training sample data, the preset reward function, and the SCA algorithm.

[0107] The specific implementation of generating simulation data samples is as follows:

[0108] Specifically, the actor network and the simulation environment are used for online interaction to generate simulation data samples, and the simulation data samples are placed in the success experience pool and the failure experience pool respectively based on the interaction results.

[0109] The simulation results include marker values.

[0110] Among them, the simulation environment is used to simulate the operation of an unmanned boat tracking a target ship. The simulation data sample is a series of data generated when the unmanned boat tracks the target ship in the simulation environment, including: the first simulation state information of the unmanned boat, the second simulation state information of the target ship, the simulation environment information corresponding to the unmanned boat and the target ship, and the simulation obstacle information.

[0111] The interaction result is used to characterize whether the unmanned boat successfully tracks the target ship. The simulation data samples corresponding to the successful tracking situation are stored in the successful experience pool, and the simulation data samples corresponding to the failed tracking situation are stored in the failed experience pool.

[0112] In a specific embodiment, the specific implementation of configuring the sample priority for each simulation data sample in the success experience pool and the failure experience pool includes:

[0113] A sample priority is configured for the simulation data samples in the successful experience pool based on the first weight calculation formula.

[0114] The first weight calculation formula is shown in formula (2):

[0115]

[0116] Among them, ω i represents the sample priority of the i-th simulation data sample in the successful experience pool, Q target,i Represents the Q value calculated by the target value network for the i-th simulation data sample, Q current,i represents the Q value calculated by the current value network for the i-th simulation data sample, and ε represents the preset minimum value.

[0117] A sample priority is configured for the simulation data samples in the failure experience pool based on the second weight calculation formula.

[0118] The second weight calculation formula is shown in formula (3):

[0119] ω j =|R t -Q current,j |……………………(3)

[0120] Among them, ω j represents the sample priority of the jth simulation data sample in the failure experience pool, Q current,jRepresents the Q value calculated by the current value network for j simulation data samples, R t Represents the reward value.

[0121] Among them, the sampling weight and sampling priority are positively correlated.

[0122] The smaller the Q-value error, the larger the corresponding sampling weight, and the higher the sampling priority. Based on this, the model can avoid focusing on high-error samples and prevent overfitting.

[0123] The larger the prediction deviation, the greater the corresponding sampling weight and the higher the sampling priority. Based on this, the model pays attention to such unexpected failure samples to quickly correct the model's misjudgment.

[0124] In a specific embodiment, the specific implementation of extracting simulation data samples from the success experience pool and the failure experience pool based on the hybrid sampling method and sample priority includes:

[0125] The sample extraction ratio of the successful experience pool is determined based on the sample priority of each simulation data sample in the successful experience pool and the failed experience pool and a preset ratio calculation formula.

[0126] The ratio calculation formula is shown in formula (4):

[0127]

[0128] Among them, β represents the sample extraction ratio of the successful experience pool, 1-β represents the sample extraction ratio of the failed experience pool, P represents the number of simulation data samples in the successful experience pool, and Q represents the number of simulation data samples in the failed experience pool.

[0129] This application obtains training sample data from the experience pool through a mixed sampling method, and then uses the training sample data to train the target control model.

[0130] In a specific embodiment, Figure 4 As shown, the actor network includes a strategy network, and the critic network includes: two value networks and two target value networks; the strategy network is connected to the value network, and one value network is connected to a target value network.

[0131] In a specific embodiment, the training process of the target control model includes:

[0132] Obtain multimodal state space samples, control instruction samples, and a reward function. Input the multimodal state space samples, control instruction samples, and reward function into the target control model. Based on the SAC algorithm, use the value network and policy network to obtain a first Q value based on the reward value output by the reward function, and use the target value network and policy network to obtain a second Q value based on the reward value output by the reward function. Calculate the mean squared error between the first Q value and the second Q value, and use the mean squared error to perform a gradient descent operation on the first network parameter of the value network. Use a preset policy loss function to perform a gradient ascent operation on the second network parameter of the policy network. Use a preset soft update mechanism to optimize the third network parameter of the target value network to obtain the final target control model.

[0133] Among them, the state space samples are time-series.

[0134] In a specific embodiment, the specific implementation of obtaining a first Q value based on a reward value output by a reward function using a value network and a policy network, and obtaining a second Q value based on a reward value output by a reward function using a target value network and a policy network includes:

[0135] The policy network generates an action probability distribution corresponding to the predicted action based on the multimodal state space samples; the target value network and the value grid obtain the Q value based on the action probability distribution, the reward value obtained by the reward function, and the Q value calculation formula respectively.

[0136] The Q value calculation formula is shown in formula (5):

[0137] Q(s′,a′)=R t +γ(min i=1,2 Q(s′,a′)-k3logπ φ (a′|s′))…(5)

[0138] Among them, R t represents the reward value, γ represents the reward discount factor, which is a constant, Q(s′,a′) represents the Q value corresponding to the i-th value network or the i-th target value network, k3 represents the temperature coefficient, which is a constant and is used to control the weight of the entropy term and balance exploration and utilization, π φ (a′|s′) represents the action probability distribution, a′ represents the control instruction sample at the next moment, and s′ represents the state space sample at the next moment.

[0139] The Q value includes a first Q value and a second Q value.

[0140] In a specific embodiment, the specific implementation of performing a gradient descent operation on the first network parameter of the value network using the mean square error includes:

[0141] The mean square error is input into a predetermined gradient descent calculation formula to obtain the optimized first network parameters output by the gradient descent calculation formula.

[0142] The gradient descent calculation formula is shown in formula (6):

[0143]

[0144] Among them, θ i represents the first network parameter of the i-th value network, η Q represents the learning rate of the value network, represents the gradient of the first network parameter of the i-th value network, L Q (θ i ) represents the mean square error.

[0145] The mean square error is obtained by formula (7):

[0146]

[0147] in, Indicates the first Q value.

[0148] In a specific embodiment, a gradient ascent operation is performed on the second network parameter of the policy network using a preset policy loss function, including:

[0149] The policy loss function is used to calculate the prediction loss of the policy network.

[0150] Among them, the policy loss function is shown in formula (8):

[0151]

[0152] Among them, L π (φ) represents the prediction loss, k3 represents the temperature coefficient, which is a constant used to control the weight of the entropy term and balance exploration and utilization, logπ φ (a|s) represents the entropy regularization term, Represents the minimized Q value. The Q value here includes the first Q value. By taking the minimum value, the problem of overestimation of a single value network is solved. a represents the control instruction sample at the current moment, and s represents the state space sample at the current moment.

[0153] The prediction loss is input into a predetermined gradient ascent calculation formula to obtain the optimized second network parameters output by the gradient ascent calculation formula.

[0154] The gradient ascent calculation formula is shown in formula (9):

[0155]

[0156] Among them, φi represents the second network parameter of the policy network, ηπ represents the learning rate of the policy network, represents the gradient of the second grid parameter.

[0157] In a specific embodiment, the specific implementation of optimizing the third network parameter of the target value network using the preset soft update mechanism includes:

[0158] The third network parameter is updated using a soft update calculation formula.

[0159] The soft update calculation formula is shown in formula (10):

[0160] θ target,i =τθ i +(1-τ)θ target,i ,i=1,2………(10)

[0161] Among them, θ target,i represents the third network parameter of the i-th target value network, τ represents the update rate of the target value network, and θ i Represents the first network parameter of the i-th value network.

[0162] This application utilizes a neural network-based reinforcement learning model that requires less frequent parameter adjustments than traditional algorithms and offers greater adaptability. Furthermore, the SAC algorithm has been improved by introducing adaptive hybrid sampling of success and failure experiences, which increases sample utilization and training efficiency. Furthermore, the trained target control model boasts high computational efficiency and a strong real-time response speed, improving the accuracy of unmanned vehicle control.

[0163] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other via the communication bus 504. The processor 501 may call the logic instructions in the memory 503 to execute the target tracking control method for the unmanned vehicle based on deep reinforcement learning.

[0164] In addition, the logic instructions in the above-mentioned memory 503 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0165] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the target tracking control method of the unmanned boat based on deep reinforcement learning provided by the above methods.

[0166] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the target tracking control method of the unmanned boat based on deep reinforcement learning provided in the above-mentioned embodiments.

[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0169] Finally, it should be noted that the above description is merely a preferred embodiment of the present application and the present application is not limited to the above embodiments. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included within the scope of protection of the present application.

Claims

1. A target tracking control method for an unmanned boat based on deep reinforcement learning, characterized in that: The method comprises: Acquiring data to be used, wherein the data to be used includes: first state information of the unmanned boat, relative state information between the unmanned boat and the target ship being tracked by the unmanned boat, ocean environment information corresponding to the unmanned boat and the target ship, and obstacle information; Construct a multimodal state space based on the data to be used; Inputting the multimodal state space into a pre-trained target control model to obtain a control instruction output by the target control model that carries a steering gear angle and a propeller speed; The target control model includes an actor network and a critic network. The target control model is trained based on multimodal state space samples and control instruction samples in combination with a preset reward function and a SAC algorithm. The control instruction is used to control the unmanned boat to continuously track the target ship, and the reward function is obtained based on the first state information, the relative state information, the ocean environment information and the obstacle information.

2. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 1 is characterized in that: The reward function includes: Among them, R t represents the reward value, ω1 represents the distance preservation reward weight coefficient, k1 represents the weight corresponding to the distance in the distance preservation reward, d t Indicates the relative distance between the target ship and the unmanned boat, d opt represents the optimal tracking distance, k2 represents the weight of speed in the distance keeping reward, v rel represents the relative speed between the unmanned boat and the target ship, ω2 represents the weight coefficient of the field of view maintenance reward, represents the azimuth deviation, ω3 represents the obstacle avoidance reward weight coefficient, N represents the number of obstacles, represents the shortest distance from the unmanned boat to the i-th obstacle, ε represents the preset minimum value, ω4 represents the energy consumption optimization reward weight coefficient, and n represents the rotation speed of the unmanned boat.

3. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 1 or 2, characterized in that: Before inputting the multimodal state space into a pre-trained target control model to obtain a control instruction outputted by the target control model and carrying a steering gear angle and a propeller speed, the method further includes: The actor network of the target control model is used to interact with the simulation environment to generate simulation data samples, wherein the simulation data samples include a tag value, and the tag value is used to indicate whether the unmanned boat successfully tracks the target ship; Based on the tag value, the simulation data sample is stored in the success experience pool or the failure experience pool, and a sample priority is configured for each simulation data sample in the success experience pool and the failure experience pool; Extracting simulation data samples from the success experience pool and the failure experience pool based on a hybrid sampling method and sample priority, and determining the extracted simulation data samples as training sample data, wherein the training sample data includes multimodal state space samples and control instruction samples; The target control model is trained based on training sample data, preset reward function and SCA algorithm.

4. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 3 is characterized in that: Configure the sample priority for each simulation data sample in the success experience pool and failure experience pool, including: Configure sample priorities for simulation data samples in the successful experience pool based on a first weight calculation formula; The first weight calculation formula includes: Among them, ω i represents the sample priority of the i-th simulation data sample in the successful experience pool, Q target,i Represents the Q value calculated by the target value network for the i-th simulation data sample, Q current,i represents the Q value calculated by the current value network for the i-th simulation data sample, and ε represents the preset minimum value; configuring sample priorities for simulation data samples in the failure experience pool based on a second weight calculation formula; The second weight calculation formula includes: oh j =|R t -Q current,j |; Among them, ω j represents the sample priority of the jth simulation data sample in the failure experience pool, Q current,j Represents the Q value calculated by the current value network for j simulation data samples, R t Indicates the reward value.

5. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 4 is characterized in that: Based on the hybrid sampling method and sample priority, simulation data samples are extracted from the successful experience pool and the failed experience pool respectively, including: Determine the sample extraction ratio of the successful experience pool based on the sample priority of each simulation data sample in the successful experience pool and the failed experience pool and a preset ratio calculation formula; The ratio calculation formula includes: Among them, β represents the sample extraction ratio of the successful experience pool, 1-β represents the sample extraction ratio of the failed experience pool, P represents the number of simulation data samples in the successful experience pool, and Q represents the number of simulation data samples in the failed experience pool.

6. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 1 or 2, characterized in that: The actor network includes a strategy network, and the critic network includes: two value networks and two target value networks; The strategy network is connected to the value network, and a value network is connected to a target value network; The training process of the target control model includes: Obtain multimodal state space samples, control instruction samples, and reward functions, where the state space samples are temporally sequential. The multimodal state space samples, control instruction samples and reward function are input into the target control model. Based on the SAC algorithm, the value network and the policy network are used to obtain the first Q value based on the reward value output by the reward function, and the target value network and the policy network are used to obtain the second Q value based on the reward value output by the reward function; the mean square error of the first Q value and the second Q value is calculated, and the mean square error is used to perform a gradient descent operation on the first network parameter of the value network; the preset policy loss function is used to perform a gradient ascent operation on the second network parameter of the policy network; and the preset soft update mechanism is used to optimize the third network parameter of the target value network to obtain the final target control model.

7. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 6, characterized in that: Obtaining a first Q value based on a reward value output by a reward function using a value network and a policy network, and obtaining a second Q value based on a reward value output by a reward function using a target value network and a policy network, including: The policy network generates an action probability distribution corresponding to the predicted action based on the multimodal state space samples; The target value network and value grid respectively obtain the Q value based on the action probability distribution, the reward value obtained by the reward function, and the Q value calculation formula; The Q value calculation formula includes: Q(s′,a′)=R t +γ(min i=1,2 Q(s′,a′)-k3logπ φ (a′|s′)); Among them, R t represents the reward value, γ represents the reward discount factor, which is a constant, Q(s′,a′) represents the Q value corresponding to the i-th value network or the i-th target value network, k3 represents the temperature coefficient, which is a constant, π φ (a′|s′) represents the action probability distribution, a′ represents the control instruction sample at the next moment, and s′ represents the state space sample at the next moment; The Q value includes a first Q value and a second Q value.

8. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 6, characterized in that: The gradient descent operation of the first network parameter of the value network is performed using the mean square error, including: Inputting the mean square error into a predetermined gradient descent calculation formula to obtain the optimized first network parameters output by the gradient descent calculation formula; Among them, the gradient descent calculation formula includes: Among them, θ i represents the first network parameter of the i-th value network, η Q represents the learning rate of the value network, represents the gradient of the first network parameter of the i-th value network, L Q (θ i ) represents the mean square error.

9. The target tracking control method of an unmanned vehicle based on deep reinforcement learning according to claim 6, characterized in that: The preset policy loss function is used to perform a gradient ascent operation on the second network parameter of the policy network, including: Use the policy loss function to calculate the prediction loss of the policy network; Among them, the strategy loss function includes: Among them, L π (φ) represents the predicted loss, k3 represents the temperature coefficient, which is a constant, logπ φ (a|s) represents the entropy regularization term, represents the minimized Q value, a represents the control instruction sample at the current moment, and s represents the state space sample at the current moment; Input the predicted loss into a predetermined gradient ascent calculation formula to obtain the optimized second network parameters output by the gradient ascent calculation formula; Among them, the gradient ascent calculation formula includes: Among them, φ i Represents the second network parameter of the policy network, η π represents the learning rate of the policy network, represents the gradient of the second grid parameter.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the target tracking control method of the unmanned boat based on deep reinforcement learning are implemented as described in any one of claims 1 to 9.