An adaptive control method, device, equipment and medium for unmanned ship formation

Through the adaptive control method of unmanned ship formation, the decision table is updated using the exploration probability and decision probability, and the problem of single coupling between the model and the environment in the existing technology is solved, and dynamic response to different environments and more effective decision execution is achieved.

CN114995446BActive Publication Date: 2025-05-16SHENYANG INST OF AUTOMATION GUANGZHOU CHINESE ACAD OF SCI +1

Patent Information

Application Number
CN202210696649.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-05-16
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

The existing multi-unmanned ship fleet control method of the pilot follow-up method has a single coupling model and the environment, and it is impossible to make action decisions based on different environments using historical data.

Method used

Adaptive control method of unmanned ship formation is adopted to control the pilot ship through preset task path planning, update the status information of the pilot ship and the following ship, build the overall status information of the unmanned ship formation, and update the decision table and execute the decision table according to the exploration probability and decision probability.

Benefits of technology

The decision-making probability of dynamically adjusting the unmanned ship formation according to environmental changes is solved, the problem of single coupling between the model and the environment in the prior art is solved, and the decision-making can be made more effectively using historical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114995446B_ABST
    Figure CN114995446B_ABST
Patent Text Reader

Abstract

The present application relates to the field of unmanned ship control technology, and discloses an adaptive control method, device, equipment and medium for an unmanned ship formation, the method comprising: controlling the pilot ship to act based on a preset mission path planning; updating the state information of the pilot ship after the pilot ship acts; obtaining the state information of each follower ship, and constructing the overall state information of the unmanned ship formation based on the updated state information of the pilot ship and the state information of each follower ship; based on the current overall state information, exploring potential better decisions according to the exploration probability or executing the optimal decision in the decision table with the decision probability. The method solves the defect that the existing pilot-follower method for controlling a multi-unmanned ship formation has a single coupling between the model and the environment, and cannot use historical data to select action decisions for different environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of unmanned ship control, and in particular to an adaptive control method, device, equipment and medium for an unmanned ship formation. Background Art

[0002] In recent years, with the attention paid by various countries to marine science and technology, surface ships are developing towards intelligence and unmanned operation. Surface unmanned ships are playing an increasingly important role in scientific research, water monitoring, marine environmental exploration, military and other civilian fields. However, the marine water environment is complex and changeable, and there are many unknown disturbances. When facing multi-task operations, a single unmanned ship cannot achieve the goal well. Therefore, the collaborative control operation of multiple unmanned ships has become a research hotspot of various research institutions due to its high efficiency, wide coverage, and strong robustness. In the field of collaborative control of multiple unmanned ships, formation control is the most important. At present, the most commonly used formation method is the pilot-follow formation method. In the pilot-follow unmanned ship formation control system, its stability will be affected by complex external environmental disturbances, such as changes in weather and climate, changes in water levels, and the influence of wind and waves, which brings great difficulties to the stable control of the unmanned ship formation.

[0003] Patent No. CN114089761A discloses a ROS-based leader-follower multi-unmanned ship formation control system and method thereof, including a control module, an inertial measurement module, a GPS module, a power module, a power module, and a wireless transmission module, wherein the control module is connected to the inertial measurement module, the GPS module, and the power module; the position information is sent to the host computer through the serial port, and the host computer transmits the current position and attitude information of the unmanned ship to the remote computer through wifi. The parameter setting of the model in this patent depends on the existing environment, and in actual applications, the model operation effect is highly sensitive to the parameters of the formation control model, and the operation effect is poor for an environment that does not match the parameters.

[0004] With respect to the above-mentioned related technologies, the existing multi-unmanned ship formation control method of the leader-follower method has the defect of single coupling between the model and the environment and cannot use historical data to select action decisions for different environments. Summary of the invention

[0005] In order to solve the defects of the existing multi-unmanned ship formation control method of the leader-follower method, which has a single coupling between the model and the environment and cannot use historical data to select action decisions for different environments, the present application provides an adaptive control method device, equipment and medium for an unmanned ship formation.

[0006] In a first aspect, the present application provides an adaptive control method for an unmanned ship formation.

[0007] The present application is realized by the following technical solutions: an adaptive control method for an unmanned ship formation, comprising:

[0008] Control the pilot ship to take action based on the preset mission path planning;

[0009] After the pilot boat takes action, updating the status information of the pilot boat;

[0010] Acquire the status information of each following ship, and form the overall status information of the unmanned ship formation based on the updated status information of the pilot ship and the status information of each following ship;

[0011] Based on the current overall state information, explore potential better decisions according to the exploration probability or execute the best decision in the decision table according to the decision probability;

[0012] The decision table includes the state information of the unmanned ship formation in the historical data, and different decisions and corresponding decision probabilities under each state. The optimal decision is the decision with the highest decision probability; the decision probability is updated after each action.

[0013] In a preferred example, the present application can be further configured as follows: the decision probability is the ratio of the reward value obtained by taking the decision π(A|S) in the state S of the determined unmanned ship formation to the total reward value, and the decision probability P π The expression of (A|S) is:

[0014]

[0015] in, To adjust the course, is the speed adjustment, γ is the attenuation coefficient and γ<1, T is the current moment, t is the moment when the state is S in the history, R(S t , A t ) is state S t Take action A based on decision π(A|S) t The reward value obtained, G(T) is the sum of the reward values ​​obtained for all execution decisions in state S.

[0016] In a preferred example, the present application can be further configured as follows: based on the current overall state information, exploring potential better decisions according to the exploration probability and executing the optimal decision in the decision table according to the decision probability, the sum of the decision probability of the optimal decision and the exploration probability is 1.

[0017] In a preferred example, the present application can be further configured as follows: the total reward value is the sum of all the reward values ​​in the historical data, and the calculation formula is:

[0018] G(T)=∑ 从0到T状态为S γT-t R(S t , A t ).

[0019] In a preferred example, the present application can be further configured as follows: the reward value is calculated by a preset reward function, and the reward function is established by the following four evaluation indicators: the concentration degree of the unmanned ship formation r1, the closeness degree of the speed of each following ship and the pilot ship r2, the closeness degree of the heading angle of each following ship and the pilot ship r3, and the closeness degree of the unmanned ship formation and the target point r4;

[0020] The expression of the reward function is as follows:

[0021]

[0022]

[0023] Wherein i and j are respectively the i-th unmanned ship and the j-th unmanned ship in the unmanned ship formation and i≠j, n is the total number of ships in the unmanned ship formation, α1, α2, α3 and α4 are weight coefficients and α1+α2+α3+α4=1.

[0024] In a preferred example, the present application can be further configured as follows: the average distance r1 between the unmanned ships and the proximity r4 between the unmanned ship formation and the target point are both introduced to avoid collision. The calculation formula is as follows:

[0025]

[0026]

[0027] Where Δx(t) ij is the horizontal coordinate difference between the i-th and j-th unmanned ships in the unmanned ship formation at time t, Δy(t) ij is the vertical coordinate difference between the i-th and j-th unmanned ships in the unmanned ship formation at time t, g(t) i is the distance between the i-th unmanned ship in the unmanned ship formation and the target point.

[0028] In a preferred example, the present application can be further configured as follows: after the pilot ship takes action, updating the status information of the pilot ship is specifically by using the pilot ship navigation to measure and update the position, speed, heading and distance information from the target point of the pilot ship; obtaining the status information of each follower ship and constituting the overall status information of the unmanned ship formation based on the updated status information of the pilot ship is specifically by using the follower ship navigation to measure the position, speed, heading and distance information from the target point of each follower ship, and the status information of the pilot ship and the status information of each follower ship together constitute the status information of the unmanned ship formation.

[0029] In a second aspect, the present application provides an adaptive control device for an unmanned ship formation.

[0030] The present application is realized by the following technical solution: an adaptive control device for an unmanned ship formation, comprising:

[0031] Control module: used to control the pilot ship to perform actions based on the preset mission path planning;

[0032] The first state measurement module is used to update the state information of the pilot ship after the pilot ship takes action;

[0033] The second state measurement module is used to obtain the state information of each following ship, and to form the overall state information of the unmanned ship formation based on the updated state information of the pilot ship;

[0034] Decision module: used to explore potential better decisions according to the exploration probability based on the current overall state information or to execute the optimal decision in the decision table according to the decision probability.

[0035] In a third aspect, the present application provides a computer device.

[0036] The present application is implemented through the following technical solutions: a computer device, including a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any one of the above-mentioned adaptive control methods for unmanned ship formations.

[0037] In a fourth aspect, the present application provides a computer-readable storage medium.

[0038] The present application is implemented through the following technical solutions: a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned adaptive control methods for unmanned ship formations are implemented.

[0039] By adopting the above technical solution, the decision probability is updated through the rewards of historical actions, which enables the simultaneous execution of the existing model while updating the decision probability to update the model. According to the continuous changes in the environment, the decision probability of the unmanned ship formation is dynamically adjusted, which solves the defect of the existing multi-unmanned ship formation control method of the leading and following method, which has a single coupling between the model and the environment and cannot use historical data to select action decisions for different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A schematic flow chart of a method for controlling a multi-unmanned ship formation based on a pilot-follower method provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0041] This specific embodiment is merely an explanation of the present application and is not a limitation of the present application. After reading this specification, those skilled in the art may make modifications to the present embodiment without any creative contribution as needed, but such modifications are protected by the patent law as long as they are within the scope of the claims of the present application.

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0043] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.

[0044] The following is combined with Figure 1 This application is described in further detail.

[0045] The present application embodiment discloses a multi-unmanned ship formation control method based on the pilot-follower method. Figure 1 ,The adaptive control method for the unmanned ship formation includes the following steps:

[0046] Step S1: Control the pilot ship to take action based on the preset mission path planning;

[0047] Step S2: after the pilot boat takes action, updating the state information of the pilot boat;

[0048] The pilot boat sails in the corresponding direction based on the control signal. After the pilot boat sails, the position coordinate x(t) of the pilot boat is measured using the i-th pilot boat at time t. i and y(t) i , speed v(t) i 、Heading θ(t) i and distance to the target point g(t) i Information and updates.

[0049] Step S3: Acquire the status information of each following ship, and construct the overall status information of the unmanned ship formation based on the updated status information of the pilot ship and the status information of each following ship;

[0050] Use follower ship navigation to measure the position, speed, heading and distance from the target point of each follower ship. The status information of the lead ship and the status information of each follower ship together constitute the status information of the unmanned ship formation.

[0051] Step S4: based on the current overall state information, exploring potential better decisions according to the exploration probability or executing the best decision in the decision table according to the decision probability;

[0052] The decision table includes the state information of the unmanned ship formation in the historical data, and the different decisions and corresponding decision probabilities in each state; the decision probability is updated after each action, and the decision taken after each action will generate a new reward value and update the total reward value. The decision probability will be calculated and updated based on the new total reward value. The decision probability of the optimal decision is ε, and the exploration probability is 1-ε. According to the probability 1-ε, feasible decisions are explored to obtain potential better decisions, and the optimal decision in the decision table is executed with probability ε.

[0053] Select the decision that meets the preset conditions from the decision table, and the follower ship takes action based on the decision that meets the preset conditions. Specifically, select the decision with the highest probability among the decisions corresponding to the current state from the decision table, and the follower ship takes action based on the selected decision, including adjusting the course. and speed adjustment

[0054] Decision π(A, S) refers to the decision in state S t Under this condition, take action according to decision π So that the state S t Change to state S t+1 However, different environments bring different disturbances, so the state S under the same decision t+1 It is not a definite amount.

[0055] Decision probability P π (A|S) is an indicator for evaluating the quality of decisions. The better the decision, the higher its decision probability. The decision with the highest decision probability is the optimal decision. The specific calculation is: from time 0 to time T now, the sum of the rewards obtained by all actions in state S and behavior A is divided by G(T), that is, the sum of the rewards obtained by all actions in state S. For each state S, the sum of the decision probabilities of all decisions is 1.

[0056] Decision probability P π The expression of (A|S) is:

[0057]

[0058] in, To adjust the course, is the speed adjustment, γ is the attenuation coefficient and γ<1, T is the current moment, t is the moment when the state is S in the history, R(S t , A t ) is state S t Take action A based on decision π(A|S) t The reward value obtained, G(T) is the sum of the reward values ​​obtained for all execution decisions in state S.

[0059] The decision probability is to combine the same reward values ​​for all actions taken in the same state and calculate the ratio of them to the total reward value to obtain the decision probability of the action in the state.

[0060] For example, for a state S1, if three different actions a, b, and c have been taken and the number of times is different, then the reward values ​​of all state S1 actions a are combined, the reward values ​​of all state S1 actions b are combined, and the reward values ​​of state S1 actions c are combined, and their decision probabilities are calculated separately.

[0061] The total reward value is the sum of all reward values ​​in the historical data, and the calculation formula is:

[0062] G(T)=∑ 从0到T状态为S γ T-t R(S t , A t )

[0063] Wherein, γ is the attenuation coefficient, γ<1.

[0064] Because the environment has disturbance information, the same state, the same decision, and the same action in different environments will result in different states at the next moment, that is, the reward value of feedback will also be different. The sum of the rewards obtained for all states S and actions A is used as the basis for updating, and then through the attenuation mechanism γ, the farther away from T, the greater the attenuation, which makes the weight of the older environmental information smaller, that is, the smaller the impact on the decision probability, so the updated environmental information has a higher influence. The update of decision probability is to constantly adapt to the environment, and even if the environment changes, it has the adaptive ability to learn from the environment through feedback.

[0065] The reward value is calculated by a preset reward function, which is established by the following four evaluation indicators: the concentration of the unmanned ship formation r1, the speed of each following ship and the pilot ship r2, the heading angle of each following ship and the pilot ship r3, and the proximity of the unmanned ship formation to the target point r4;

[0066] The expression of the reward function is as follows:

[0067]

[0068]

[0069] Wherein i and j are respectively the i-th unmanned ship and the j-th unmanned ship in the unmanned ship formation and i≠j, n is the total number of ships in the unmanned ship formation, α1, α2, α3 and α4 are weight coefficients and α1+α2+α3+α4=1.

[0070] r1 is the concentration degree of the unmanned ship formation, which is used to evaluate the concentration degree of each ship in the unmanned ship formation. The expression is:

[0071]

[0072] Where Δx(t) ij is the horizontal coordinate difference between the i-th and j-th unmanned ships in the unmanned ship formation at time t, Δy(t) ij is the vertical coordinate difference between the i-th and j-th unmanned ships in the unmanned ship formation at time t. In order to avoid collision, the maximum safety distance s is introduced in r1 to ensure that the distance between every two ships must be greater than the maximum safety distance;

[0073] r2 is the speed closeness between each following ship and the pilot ship, expressed as:

[0074]

[0075] Where, Δv(t) i is the speed difference between each following ship and the leading ship in the unmanned ship formation;

[0076] r3 is the closeness of the heading angles of each following ship and the pilot ship, expressed as:

[0077]

[0078] Where Δθ(t) i is the heading difference between each following ship and the leading ship in the unmanned ship formation;

[0079] r4 is the proximity between the unmanned ship formation and the target point, and its expression is:

[0080]

[0081] Among them, g(t) i is the distance between each ship in the unmanned ship formation and the target point. In order to avoid collision, a maximum safety distance s is introduced in r4 to ensure that the distance between each ship and the target point must be greater than the safety distance.

[0082] For the four evaluation indicators r1, r2, r3 and r4, the sigmoid function is used After smoothing, four weight coefficients α1, α2, α3 and α4 are used to achieve normalization, so that R(St , A t ) ranges from -1 to 1. A negative number means that the state gets worse after the decision is made; a positive number means that the state gets better after the decision is made. The worst evaluation is -1; the best evaluation is 1.

[0083] After each action at a moment, the decision probability and the preset decision table have been updated based on the decision at the current moment, and steps S1 to S4 are repeated at the next moment based on the updated decision table.

[0084] The four evaluation indicators r1, r2, r3 and r4 are used to evaluate the quality of each action, form the reward value of the action, and merge the same items for different actions in the same state. The reward value obtained by taking action in the state is then normalized by the total establishment function to form a probability, which is the decision probability, and is used as the decision probability of different actions in different states at different moments in the future. At each moment, the unmanned ship formation will make decisions and take actions based on the decision probability obtained using historical data, and will calculate the reward value of this action to update the total reward value and decision probability. The decision probability is updated by the reward value of historical actions, which realizes the synchronization of executing the existing model while updating the decision probability to update the model, and dynamically adjusts the decision probability of the unmanned ship formation according to the changing environment.

[0085] The implementation principle of the embodiment is as follows: at each time point, the pilot ship acts according to the preset planned route, and the state of the pilot ship is obtained by the pilot ship navigation. After the state of the pilot ship changes, the overall state information of the entire unmanned fleet is obtained. According to the preset decision probability and the overall state information of the unmanned ship formation at the current moment, each follower ship will explore feasible decisions according to the probability 1-ε to obtain potential better decisions, and take actions with the probability ε to execute the best decision in the decision table. After each follower ship takes action, the reward value of the action is calculated to evaluate the pros and cons of the action, and the total reward value is updated. Then, based on the newly obtained reward value and the total reward value, the decision probability is updated, and the preset decision table is updated. From the first moment to the completion of the entire action of the unmanned ship formation, each action updates the decision probability through the reward value of the historical action to determine the decision taken in the corresponding state, realizing the simultaneous execution of the existing model while updating the decision probability to update the model. According to the continuous changes in the environment, the decision probability of the unmanned ship formation is dynamically adjusted, which solves the defect that the existing multi-unmanned ship formation control method of the leading and following method has the model parameter setting dependent on the existing environment and cannot use historical data to select action decisions for different environments.

[0086] An embodiment of the present application also provides an adaptive control device for an unmanned ship formation.

[0087] The adaptive control device of the unmanned ship formation includes:

[0088] Control module: used to control the pilot ship to perform actions based on the preset mission path planning;

[0089] The first state measurement module is used to update the state information of the pilot ship after the pilot ship takes action;

[0090] The second state measurement module is used to obtain the state information of each following ship, and to form the overall state information of the unmanned ship formation based on the updated state information of the pilot ship;

[0091] Decision module: used to explore potential better decisions according to the exploration probability based on the current overall state information or to execute the optimal decision in the decision table according to the decision probability.

[0092] For the specific definition of an adaptive control device for an unmanned ship formation, please refer to the definition of an adaptive control method for an unmanned ship formation mentioned above, which will not be repeated here. Each module in the above-mentioned adaptive control device for an unmanned ship formation can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0093] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, any of the above-mentioned adaptive control methods for unmanned ship formations is implemented.

[0094] In one embodiment, a computer-readable storage medium is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0095] Step S1: Control the pilot ship to take action based on the preset mission path planning;

[0096] Step S2: after the pilot ship takes action, updating the state information of the pilot ship;

[0097] Step S3: Acquire the status information of each following ship, and form the overall status information of the unmanned ship formation based on the updated status information of the pilot ship;

[0098] Step S4: Based on the current overall state information, explore potential better decisions according to the exploration probability or execute the optimal decision in the decision table with the decision probability.

[0099] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0100] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0101] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.

Claims

1. An adaptive control method for an unmanned ship formation, characterized in that: include: Control the pilot ship to take action based on the preset mission path planning; After the pilot boat takes action, updating the status information of the pilot boat; Acquire the status information of each following ship, and form the overall status information of the unmanned ship formation based on the updated status information of the pilot ship and the status information of each following ship; Based on the current overall state information, explore potential better decisions according to the exploration probability or execute the best decision in the decision table according to the decision probability; The decision table includes the state information of the unmanned ship formation in the historical data, and different decisions and corresponding decision probabilities under each state. The optimal decision is the decision with the highest decision probability. The decision probability is updated after each action. The decision probability is the ratio of the reward value obtained by taking the decision π(A|S) in the state S of the unmanned ship formation in the past to the total reward value. The decision probability P π The expression of (A|S) is: in, To adjust the course, is the speed adjustment, γ is the attenuation coefficient and γ<1, T is the current moment, t is the moment when the state is S in the history, R(S t , A t ) is state S t Take action A based on decision π(A|S) t The reward value obtained, G(T) is the sum of the reward values ​​obtained by all execution decisions in state S; The reward value is calculated by a preset reward function, and the reward function is established by the following four evaluation indicators: the concentration degree of the unmanned ship formation r1, the closeness degree of the speed of each following ship and the pilot ship r2, the closeness degree of the heading angle of each following ship and the pilot ship r3, and the closeness degree of the unmanned ship formation and the target point r4; The expression of the reward function is as follows: Wherein i and j are respectively the i-th unmanned ship and the j-th unmanned ship in the unmanned ship formation and i≠j, n is the total number of ships in the unmanned ship formation, α1, α2, α3 and α4 are weight coefficients and α1+α2+α3+α4=1.

2. The adaptive control method of the unmanned ship formation according to claim 1, characterized in that: Based on the current overall state information, the potential better decision is explored according to the exploration probability, or the optimal decision in the decision table is executed with the decision probability, and the sum of the decision probability of the optimal decision and the exploration probability is 1.

3. The adaptive control method of the unmanned ship formation according to claim 1, characterized in that: The total reward value is the sum of all the reward values ​​in the historical data, and the calculation formula is: G(T)=∑ 从0到T状态为s γ T-t R(S t ,A t )。 4. The adaptive control method of the unmanned ship formation according to claim 1, characterized in that: The concentration degree r1 of the unmanned ship formation and the proximity degree r4 of the unmanned ship formation and the target point are both introduced into the maximum safety distance s to avoid collision. The calculation formula is as follows: Where Δx(t) ij is the horizontal coordinate difference between the i-th and j-th unmanned ships in the unmanned ship formation at time t, Δy(t) ij is the vertical coordinate difference between the i-th and j-th unmanned ships in the unmanned ship formation at time t, g(t) i is the distance between the i-th unmanned ship in the unmanned ship formation and the target point.

5. The adaptive control method of unmanned ship formation according to claim 1, characterized in that: The updating of the status information of the pilot ship after the pilot ship takes action is specifically, using the pilot ship navigation to measure and update the position, speed, heading and distance information from the target point of the pilot ship; the obtaining of the status information of each follower ship and forming the overall status information of the unmanned ship formation based on the updated status information of the pilot ship is specifically, using the follower ship navigation to measure the position, speed, heading and distance information from the target point of each follower ship, and the status information of the pilot ship and the status information of each follower ship together constitute the status information of the unmanned ship formation.

6. An adaptive control device for an unmanned ship formation, characterized in that: include: Control module: used to control the pilot ship to perform actions based on the preset mission path planning; The first state measurement module is used to update the state information of the pilot ship after the pilot ship takes action; The second state measurement module is used to obtain the state information of each follower ship, and to form the overall state information of the unmanned ship formation based on the updated state information of the pilot ship and the state information of each follower ship; Decision module: used to explore potential better decisions according to the exploration probability or execute the best decision in the decision table according to the decision probability based on the current overall state information; The decision table includes the state information of the unmanned ship formation in the historical data, and different decisions and corresponding decision probabilities under each state. The optimal decision is the decision with the highest decision probability. The decision probability is updated after each action. The decision probability is the ratio of the reward value obtained by taking the decision π(A|S) in the state S of the unmanned ship formation in the past to the total reward value. The decision probability P π The expression of (A|S) is: in, To adjust the course, is the speed adjustment, γ is the attenuation coefficient and γ<1, T is the current moment, t is the moment when the state is S in the history, R(S t , A t ) is state S t Take action A based on decision π(A|S) t The reward value obtained, G(T) is the sum of the reward values ​​obtained by all execution decisions in state S; The reward value is calculated by a preset reward function, and the reward function is established by the following four evaluation indicators: the concentration degree of the unmanned ship formation r1, the closeness degree of the speed of each following ship and the pilot ship r2, the closeness degree of the heading angle of each following ship and the pilot ship r3, and the closeness degree of the unmanned ship formation and the target point r4; The expression of the reward function is as follows: Wherein i and j are respectively the i-th unmanned ship and the j-th unmanned ship in the unmanned ship formation and i≠j, n is the total number of ships in the unmanned ship formation, α1, α2, α3 and α4 are weight coefficients and α1+α2+α3+α4=1.

7. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Unmanned ship formation path tracking method based on deep reinforcement learning

    CN111694365A

  • Ship formation path active coordination system and method considering maneuverability difference

    CN112462777A

Cited By

  • Under-actuated multi-unmanned ship system man-in-the-loop self-triggering formation control method

    CN120821281A

  • An under-actuated multi-unmanned ship system human-in-the-loop self-triggered formation control method

    CN120821281B