Path planning method, device and equipment for underwater unmanned vehicle, and storage medium

By introducing dynamic composite reward value update action-value table, the data acquisition path of underwater unmanned vehicles is optimized, and the problems of low throughput and efficiency of autonomous underwater vehicles are solved, and more efficient data acquisition is achieved.

CN120333468AActive Publication Date: 2025-07-18JILIN UNIVERSITY

Patent Information

Application Number
CN202510817977.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The throughput and efficiency of existing autonomous underwater vehicles during data acquisition are relatively low.

Method used

The dynamic compound reward value is used to update the preset action-value table. By obtaining the node access sequence, calculating the dynamic compound reward value and updating the target action-value table, the data acquisition path is optimized.

Benefits of technology

It improves the data acquisition throughput and efficiency of underwater unmanned vehicles, reduces the probability of packet loss and reduces the number of retransmissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120333468A_ABST
    Figure CN120333468A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of path planning, and provides a path planning method, device and equipment for an underwater unmanned vehicle, and a storage medium, and the method comprises the steps: obtaining a node access sequence of the underwater unmanned vehicle; determining the current state of the underwater unmanned vehicle by taking the first acquisition node as a target acquisition node; executing a current action corresponding to the current state to obtain a next state and a node residual data volume corresponding to the next state; calculating a dynamic composite reward value; judging whether the residual data volume of the node is 0; if the residual data volume of the node is not 0 and is within a preset communication range, updating a corresponding expected return value in a preset action-value table by using the dynamic composite reward value to obtain a target action-value table; and obtaining a data acquisition path of the underwater unmanned vehicle based on the target action-value table. Therefore, the throughput and efficiency of data acquisition of the unmanned underwater vehicle can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of path planning, and particularly relates to a path planning method, device, equipment and storage medium for an underwater unmanned vehicle. Background Art

[0002] With the development of marine science and the progress of artificial intelligence, autonomous underwater vehicles are developing towards self-learning and self-adaptation. At present, most of the underwater unmanned vehicles used for deep-water exploration are underactuated underwater unmanned vehicles, which generally only include a stern thruster, and steering and pitching are achieved through vector propulsion or rudders.

[0003] Path planning is one of the core issues in the field of underactuated underwater unmanned vehicles, and is also an important prerequisite for underwater unmanned vehicles to achieve autonomous decision-making, running through the whole process of underwater operation of underwater unmanned vehicles. Currently, commonly used path planning algorithms include artificial potential field method, ant colony algorithm, genetic algorithm, etc. With the rapid development of machine learning technology, reinforcement learning algorithm has become a popular solution to solve the path planning problem.

[0004] However, the throughput and efficiency of existing autonomous underwater vehicles during data collection are relatively low. Summary of the Invention

[0005] Embodiments of the present application provide a path planning method, device, computer equipment and storage medium for an underwater unmanned vehicle, aiming to solve the problem that the throughput and efficiency of existing autonomous underwater vehicles during data collection are relatively low.

[0006] In a first aspect, embodiments of the present application provide a path planning method for an underwater unmanned vehicle, the method comprising: Obtain a node access sequence of the underwater unmanned vehicle, where the node access sequence includes N collection nodes; Take the first collection node as the target collection node, and determine the current state of the underwater unmanned vehicle at the target collection node; Execute the current action corresponding to the current state to obtain the next state of the underwater unmanned vehicle at the target collection node and the remaining data volume of the node corresponding to the next state; Update a preset initial remaining data volume of the node by using the remaining data volume of the node corresponding to the next state; Calculate a dynamic composite reward value of the current action, where the dynamic composite reward value includes valid information of data packets; Judge whether the updated remaining data volume of the node is 0; If the remaining data volume of the node is not 0 and within the preset communication range, the corresponding expected return value in the preset action-value table is updated using the dynamic composite reward value to obtain a target action-value table; Based on the target action-value table, a data acquisition path for the underwater unmanned vehicle is obtained.

[0007] In some possible implementation manners, the judging whether the updated remaining data volume of the node is 0 further includes: If the updated remaining data volume of the node is 0, the next acquisition node of the target acquisition node is set as the new target acquisition node, and the process returns to the step of determining the current state of the underwater unmanned vehicle at the target acquisition node.

[0008] In some possible implementation manners, calculating the dynamic composite reward value of the current action includes: Obtaining a first distance reward between the underwater unmanned vehicle and the target acquisition node, a direction angle reward between the underwater unmanned vehicle and the target acquisition node, a second distance reward between the underwater unmanned vehicle and the next acquisition node of the target acquisition node, an effective reward for the proportion of valid data in the received data packet by the underwater unmanned vehicle, and a final reward; Based on the first distance reward, the direction angle reward, the second distance reward, the effective reward, and the final reward, the dynamic composite reward value of the current action is obtained.

[0009] In some possible implementation manners, the effective reward for the proportion of valid data in the data packet received by the underwater unmanned vehicle is calculated using the following formula: ; where, is the effective reward, represents the sum of the data payloads of the target acquisition node collected by the underwater unmanned vehicle, r represents the number of data packets collected by the underwater unmanned vehicle, j is the index of the received data packet, is the total amount of data of the target acquisition node , is the weight coefficient, and C is the proportion of valid data in the data packet; where, , k is the length of the valid data, and n is the total length of the data packet.

[0010] In some possible implementation manners, the expected return value is updated using the following formula: ; where, represents the current state The action taken next of value is the learning rate indicating the maximum reward estimate value in the next state is the maximum reward estimate value in the next state is the action taken in the next state is the action taken in the next state is the discount factor indicating the dynamic composite reward value

[0011] In some possible implementation manners, determining the current state of the underwater unmanned vehicle at the target acquisition node includes: Obtaining the absolute position of the underwater unmanned vehicle relative to the origin of the geodetic coordinate system, the pitch angle and yaw angle of the underwater unmanned vehicle, and the signal-to-noise ratio of the underwater unmanned vehicle at the current position coordinates and adjacent position coordinates around; Based on the absolute position, the pitch angle, the yaw angle, and the signal-to-noise ratio, obtaining the state space of the underwater unmanned vehicle; Selecting the current state of the underwater unmanned vehicle at the target acquisition node from the state space

[0012] In some possible implementation manners, the remaining data volume of the node corresponding to the next state is obtained in the following manner: Obtaining the payload of the data packet collected in the next state; Based on the payload and the preset initial remaining data volume of the node, obtaining the remaining data volume of the node corresponding to the next state

[0013] In a second aspect, an embodiment of the present application further provides a path planning device for an underwater unmanned vehicle, which includes units for executing the above method

[0014] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor, and a computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented

[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, and the storage medium stores a computer program, and when the computer program is executed by a processor, the above method can be implemented

[0016] The embodiments of the present application provide a path planning method, device, equipment and storage medium for an underwater unmanned vehicle. By introducing a dynamic composite reward value in the embodiments of the present application, since the dynamic composite reward value contains valid information of data packets and can dynamically update the corresponding expected return value in a preset action-value table, the expected return value in the preset action-value table can be continuously optimized as the environment changes and the task progresses, rather than being static. In this way, the underwater unmanned vehicle can perform data collection under good communication quality conditions, thereby reducing the probability of data packet loss, reducing the number of retransmissions, and improving the throughput of data collection. Moreover, the underwater unmanned vehicle can select an optimal data collection path from the target action-value table obtained after dynamic update, thereby improving the data collection efficiency. Description of the Drawings

[0017] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0020] Figure 1 Schematic flowchart of the first embodiment of a path planning method for an underwater unmanned vehicle provided by the embodiments of the present application.

[0021] Figure 2 Schematic diagram of RS coding provided by the embodiments of the present application.

[0022] Figure 3 Schematic diagram of the structure of a data packet provided by the embodiments of the present application.

[0023] Figure 4 Schematic diagram of the structure of a computer device provided by the embodiments of the present application.

[0024] Explanation of the reference numerals in the drawings: Processor 111, communication interface 112, memory 113, communication bus 114. Detailed Embodiments

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0026] The following disclosure provides many different embodiments or examples for implementing different structures of this application. To simplify the disclosure of this application, components and settings of specific examples are described below. Of course, they are merely examples and are not intended to limit this application. In addition, this application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0027] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0028] It should also be understood that the terms used in this specification of this application are merely for the purpose of describing specific embodiments and are not intended to limit this application. As used in this specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0029] It should be further understood that the term "and / or" used in this specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0030] As used in this specification and the appended claims, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0031] To solve the technical problem of low throughput and efficiency in the data collection process of autonomous underwater vehicles in the prior art, the present application provides a path planning method, device, equipment, and storage medium for an underwater unmanned vehicle, which can enable the underwater unmanned vehicle to perform dynamic collection, thereby improving the throughput and efficiency of data collection of the underwater unmanned vehicle.

[0032] Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of a path planning method for an underwater unmanned vehicle provided by an embodiment of the present application. The method includes: Step 110: Obtain the node access sequence of the underwater unmanned vehicle.

[0033] Among them, the node access sequence includes N collection nodes.

[0034] Step 120: Take the first collection node as the target collection node, and determine the current state of the underwater unmanned vehicle at the target collection node.

[0035] Step 130: Execute the current action corresponding to the current state to obtain the next state of the underwater unmanned vehicle at the target collection node and the remaining data volume of the node corresponding to the next state.

[0036] Step 140: Update the preset initial remaining data volume of the node by using the remaining data volume of the node corresponding to the next state.

[0037] Step 150: Calculate the dynamic composite reward value of the current action.

[0038] Among them, the dynamic composite reward value contains the valid information of the data packet.

[0039] Step 160: Determine whether the updated remaining data volume of the node is 0.

[0040] Step 170: If the remaining data volume of the node is not 0 and within the preset communication range, update the corresponding expected return value in the preset action-value table by using the dynamic composite reward value to obtain the target action-value table.

[0041] Among them, the preset action-value table is a two-dimensional array, and each cell in the preset action-value table represents the expected return of taking action in state . Usually initialized to zero or other small random values, that is .

[0042] Step 180: Based on the target action-value table, obtain the data collection path of the underwater unmanned vehicle.

[0043] Thus, in this embodiment, by introducing a dynamic composite reward value, since the dynamic composite reward value contains the valid information of the data packet and can dynamically update the corresponding expected return value in the preset action-value table, the expected return value in the preset action-value table can be continuously optimized with the change of the environment and the progress of the task, rather than being static. Thus, the underwater unmanned vehicle can perform data collection under the condition of better communication quality, thereby reducing the probability of data packet loss, reducing the number of retransmissions, and improving the throughput of data collection. Moreover, the underwater unmanned vehicle can select an optimal data collection path from the target action-value table obtained after dynamic update, thereby improving the data collection efficiency.

[0044] In some possible implementation manners, determining whether the remaining data volume of the updated node is 0 further includes: If the remaining data volume of the updated node is 0, then set the next collection node of the target collection node as the new target collection node, and return to the step of determining the current state of the underwater unmanned vehicle at the target collection node.

[0045] In some possible implementation manners, step 140, that is, calculating the dynamic composite reward value of the current action, includes: Step 141: Obtain the first distance reward between the underwater unmanned vehicle and the target collection node, the direction angle reward between the underwater unmanned vehicle and the target collection node, the second distance reward between the underwater unmanned vehicle and the next collection node of the target collection node, the valid reward of the proportion of valid data received by the underwater unmanned vehicle for the data packet, and the final reward.

[0046] Step 142: Based on the first distance reward, the direction angle reward, the second distance reward, the valid reward, and the final reward, obtain the dynamic composite reward value of the current action.

[0047] For steps 141 and 142, the dynamic composite reward value can be calculated using the following formula: .

[0048] Wherein, represents the dynamic composite reward value, represents the first distance reward between the underwater unmanned vehicle and the target collection node (i.e., the i-th collection node) , represents the direction angle reward between the underwater unmanned vehicle and the target collection node , represents the second distance reward between the underwater unmanned vehicle and the next node , Represents the effective reward for the proportion of valid data in the data packet received by the underwater unmanned vehicle, Indicates the final reward.

[0049] In some possible implementation manners, The following formula can be used for calculation: .

[0050] Wherein, Represents the Euclidean distance between the position of the underwater unmanned vehicle and the target acquisition node at time slot t, Represents the Euclidean distance between the position of the underwater unmanned vehicle and the next node at time slot t + 1, is the weight coefficient.

[0051] If is positive, it indicates that the underwater unmanned vehicle is closer to the target acquisition node , and the first distance reward is larger. On the contrary, the smaller the value of, the smaller the first distance reward .

[0052] In some possible implementation manners, The following formula can be used for calculation: .

[0053] Wherein, Represents and the included angle between, is the weight coefficient, the smaller it is, the closer the path length selected by the underwater unmanned vehicle is to the straight-line distance at this time, and the larger the direction angle reward .

[0054] In some possible implementation manners, The following formula can be used for calculation: .

[0055] Wherein, Represents the Euclidean distance between the position of the underwater unmanned vehicle and the next node at time slot t; Represents the Euclidean distance between the position of the underwater unmanned vehicle and the next node at time slot t + 1; is the weight coefficient.

[0056] If If it is positive, it indicates that the underwater unmanned vehicle is closer to the next node and the second distance reward is greater.

[0057] In some possible embodiments, the effective reward for the ratio of valid data in the data packets received by the underwater unmanned vehicle can be calculated using the following formula: .

[0058] Wherein, is the effective reward, represents the sum of the data payloads of the target acquisition nodes collected by the underwater unmanned vehicle , r represents the number of data packets collected by the underwater unmanned vehicle, j is the index of the received data packets, is the total amount of data of the target acquisition node , is the weight coefficient, C is the ratio of valid data in the data packet; Wherein, , k is the length of the valid data, and n is the total length of the data packet.

[0059] When , that is, when the sum of the currently collected data payloads is less than 2 / 3 of the total node data volume, the larger the ratio C of valid data, the larger the effective reward ; when , the larger the ratio C of valid data, the smaller the effective reward .

[0060] Since the underwater acoustic channel is a key factor affecting data communication performance, its unique propagation characteristics have an important impact on data transmission. In the three-dimensional space where the underwater unmanned vehicle performs data acquisition tasks, the bit error rate and throughput of the data packets received by the underwater unmanned vehicle are usually used as indicators to judge the efficiency of the acquisition task, and the bit error rate and packet loss rate are mainly determined by the signal-to-noise ratio (SNR) of the received signal, and the SNR depends on the underwater acoustic channel environment. For an underwater acoustic signal with a frequency of f, according to the sonar equation, the SNR of the signal at the receiving end can be expressed as: .

[0061] Wherein, is the transmitting sound source level, is the propagation loss, is the ocean ambient noise, is the frequency bandwidth of the underwater acoustic signal, is the directivity index, usually set to 0 dB.

[0062] The Bit Error Rate (BER) refers to the ratio of the number of error bits received during the communication process to the total number of transmitted bits. There is a close relationship between the Bit Error Rate and the Signal-to-Noise Ratio. According to the theoretical formula, the relationship between the Bit Error Rate (BER) and the Signal-to-Noise Ratio (SNR) can be expressed as: .

[0063] Where, represents the Gaussian error function. This formula indicates that the Bit Error Rate decreases as the Signal-to-Noise Ratio increases. That is to say, the higher the Signal-to-Noise Ratio, the lower the Bit Error Rate, and the higher the system throughput.

[0064] The throughput refers to the amount of data passing through the network per unit time and is expressed as: .

[0065] Where, is the throughput, is the total length of the data collected by the underwater unmanned vehicle, is the total duration of the collection task.

[0066] To address the above problems, this application adopts the channel coding method of Reed-Solomon (RS), abbreviated as the RS channel coding method. RS channel coding is a powerful Forward Error Correction (FEC) coding technology and plays a key role in communication systems. It is particularly outstanding in combating channel noise, interference, and data loss, and is good at dealing with burst errors and high-noise environments.

[0067] Specifically, as Figure 2 shown, the form of RS coding is usually denoted as RS(n, k), where n represents the total number of symbols after coding (including redundancy), which is generally a fixed value; k represents the number of symbols of the effective data; represents the number of added redundant symbols; t represents the maximum number of symbol errors that can be detected and corrected.

[0068] In RS coding, the codeword length exists in units of bytes, and each symbol contains 8 bits. Therefore, the Symbol Error Rate (SER) is introduced to measure the reliability of data transmission at the symbol level. The general formulas for the Bit Error Rate (BER) and the Symbol Error Rate (SER) are: .

[0069] Then the derivation formula for the Symbol Error Rate (SER) and the Signal-to-Noise Ratio (SNR) is: .

[0070] In this application, the total length of the data packet is a fixed value, that is, n is a fixed value; in order to adapt to different channel environments, when the number of symbol errors during transmission does not exceed t, RS decoding can successfully recover the original data. Therefore, to ensure successful decoding, the following conditions need to be met: 。

[0071] 。

[0072] Therefore, the effective data length of the data packet based on RS coding can be dynamically set to: 。

[0073] Among them, k changes dynamically according to the channel quality. Therefore, dynamic coding can be performed by controlling the value of the effective data k. When ( is the minimum amount of data that can be transmitted), that is, the current communication channel condition is too poor to meet the effective transmission requirements. To reduce the energy consumption waste caused by ineffective communication, this application can set the minimum transmission data volume to , that is, when k < , the sensor node does not transmit data until the channel quality is restored.

[0074] Due to the combined effects of underwater acoustic transmission limitations such as limited bandwidth, slow underwater sound speed, large transmission attenuation, and the complex marine environment, the underwater acoustic channel has characteristics such as low propagation rate, large time delay, and high bit error rate. Based on this, this application can effectively correct symbol errors and improve the reliability during data transmission by adopting the RS channel coding method and sending the data packet in the form of RS coding.

[0075] In some possible implementation manners, the length of the data packet of this application can be assumed to be 256 bytes, and OFDM modulation is adopted to enhance the anti-multipath ability, and the symbol duration ≥ 20 ms (adapting to the long delay of the underwater acoustic channel).

[0076] Exemplarily, the data packet structure can be as Figure 3 shown.

[0077] As Figure 3 shown, the effective data k includes a preamble, a packet header, a data payload d, and a packet tail.

[0078] Among them, the value of the data payload d is , and B is in bytes.

[0079] In some possible implementation manners, the remaining data volume of the node corresponding to the next state can be calculated based on the data payload d (i.e., the effective payload), specifically refer to the following: 1) Obtain the effective payload of the data packet collected in the next state; 2) Obtain the remaining data volume of the node corresponding to the next state based on the payload and the preset initial remaining data volume of the node.

[0080] Among them, the remaining data volume of the node can be calculated by the following formula: .

[0081] Taking the updated next state as for illustration, then is the remaining data volume of the node after the next update, is the remaining data volume of the node before the next update, where can be the preset initial remaining data volume of the node , or can be the value after multiple updates, d can be the payload of the data packet collected in the state .

[0082] In some possible implementation manners, the expected return value is updated by the following formula: .

[0083] Among them, represents the value of taking the action in the current state , is the learning rate, represents the maximum reward estimation value in the next state , is the action taken in the next state , is the discount factor, represents the dynamic composite reward value.

[0084] Among them, if the underwater unmanned vehicle does not receive the data packet sent by the target acquisition node , then is 0.

[0085] In some possible implementation manners, if the underwater unmanned vehicle reaches the end range, R5 = 5, if the underwater unmanned vehicle does not reach the end range, then is 0.

[0086] In some possible implementation manners, determining the current state of the underwater unmanned vehicle at the target acquisition node includes: Step 121: Obtain the absolute position of the underwater unmanned vehicle relative to the origin of the geodetic coordinate system, the pitch angle and yaw angle of the underwater unmanned vehicle, and the signal-to-noise ratio of the underwater unmanned vehicle at the current position coordinates and adjacent position coordinates around it.

[0087] Step 122: Based on the absolute position, the pitch angle, yaw angle, and the signal-to-noise ratio, obtain the state space of the underwater unmanned vehicle.

[0088] In some possible implementation manners, the state space S of the underwater unmanned vehicle includes the position coordinates of the underwater unmanned vehicle relative to the geodetic coordinate system and the angular attitude, and can be specifically expressed as follows: 。

[0089] Among them, , represents the absolute position of the underwater unmanned vehicle relative to the origin of the geodetic coordinate system, , represents the pitch angle of the underwater unmanned vehicle (longitudinal tilt motion centered on the y-axis), represents the yaw angle of the underwater unmanned vehicle (yaw motion rotating around the z-axis), represents the signal-to-noise ratio of the current position coordinates of the underwater unmanned vehicle.

[0090] The underwater unmanned vehicle receives and sends data through an underwater transceiver. Since the underwater transceiver is fixedly connected to the underwater unmanned vehicle and the distance is short, the underwater transceiver coordinate system in this embodiment approximately coincides with the underwater unmanned vehicle body coordinate system.

[0091] Step 123: Select the current state of the underwater unmanned vehicle at the target acquisition node from the state space of the underwater unmanned vehicle.

[0092] In some possible implementation manners, the selecting and executing the current action corresponding to the current state, and taking the action as the target action, includes: Step 131: Obtain the first action space selected by the underwater unmanned vehicle in the yaw angle, and the second action space selected by the underwater unmanned vehicle in the pitch angle.

[0093] Step 132: Based on the first action space and the second action space, obtain the target action space of the underwater unmanned vehicle.

[0094] Step 133: Select the current action corresponding to the current state from the target action space.

[0095] Since the turning attitude of an underwater unmanned vehicle is usually changed by vector propulsion or a rudder, it can only achieve forward, left, and right turns, as well as up and down pitches on the horizontal plane. Based on this, the present invention selects the target action space of the underwater unmanned vehicle as .

[0096] Wherein, , represents the action space that the underwater unmanned vehicle can select at the yaw angle, that is, the first action space, is a 90° left turn, is a 45° left turn, represents going straight, is a 45° right turn, is a 90° right turn, , represents the action space that the underwater unmanned vehicle can select at the pitch angle, that is, the second action space, is a 45° upward pitch, represents going straight, is a 45° downward pitch.

[0097] Based on the above embodiments, the path planning method of the underwater unmanned vehicle provided by this application mainly includes the following steps: S10: Initialize the Q table.

[0098] Wherein, the Q table is usually initialized to zero or other small random values, that is .

[0099] S11: Initialize the environmental parameters.

[0100] Wherein, the environmental parameters include the acquisition area, the pose of the underwater unmanned vehicle, the number of sensor nodes N, the position Sn, the total amount of data D, the access sequence V of the sensor nodes, and the remaining data amount of the nodes.

[0101] For example, let the preset initial remaining data amount of the node be D.

[0102] S12: Initialize the index i = 0 of the node access sequence V.

[0103] Wherein, , represents the target acquisition node of the AUV, and the position of the target acquisition node is .

[0104] Each acquisition node has the functions of environmental data monitoring, sending, and receiving. The total amount of data of each acquisition node is different. Each acquisition node broadcasts data packets at the beginning of each time slot, and the size of each data packet sent is the same.

[0105] For example, it can be used Indicates the total amount of data of the target acquisition node, using to indicate the remaining data volume of the target acquisition node.

[0106] S13: Based on the current state of the AUV select an action .

[0107] S14: The AUV obtains the next state , and updates the remaining data volume of the node .

[0108] For S13 and S14, it is mainly to select and execute actions in each time slot. The underwater unmanned vehicle needs to select an action according to the current state . For example, the underwater unmanned vehicle executes the selected action as , observes the obtained dynamic composite reward value and the new state , and updates the remaining data volume of the node under the state .

[0109] S15: Calculate the dynamic composite reward value of the current action .

[0110] In the Q-learning algorithm, the dynamic composite reward value is used to measure whether the action is correct. According to the influence of the channel quality and its own maneuverability characteristics during the navigation of the underwater unmanned vehicle, this application introduces the dynamic composite reward value .

[0111] Among them, the calculation of the dynamic composite reward value can refer to the above embodiments, and this application will not elaborate here.

[0112] S16: Judge whether the remaining data volume of the target acquisition node is not zero.

[0113] If so, execute S17; if not, execute S27.

[0114] S17: Judge whether the underwater unmanned vehicle is within the communication range.

[0115] If so, execute S18; if not, execute S20.

[0116] S18: Update the Q table, let .

[0117] S19: Determine whether i is equal to N.

[0118] If so, execute S20; if not, return to S14.

[0119] S20: Determine whether the Q-table converges.

[0120] If so, execute S21; if not, return to S12.

[0121] S21: According to the converged Q-table, simulate and obtain the position coordinates of the underwater unmanned vehicle, and generate the optimal operation path of the underwater unmanned vehicle.

[0122] In some embodiments, after the data collection of all acquisition nodes is completed, a Non-Uniform Rational B-Spline (NURBS) curve can be introduced to smooth the data collection path output by Q-learning, so that the output data collection path better conforms to the navigation maneuverability of the AUV.

[0123] Corresponding to the above path planning method of the underwater unmanned vehicle, the present application also provides a path planning device for the underwater unmanned vehicle. The path planning device for the underwater unmanned vehicle includes a unit for executing the above path planning method of the underwater unmanned vehicle, and the path planning device for the underwater unmanned vehicle can be configured in a desktop computer, a tablet computer, a laptop computer, and other terminals.

[0124] As Figure 4 shown, an embodiment of the present application provides a computer device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114. The memory 113 is used to store a computer program. In an embodiment of the present application, when the processor 111 executes the program stored on the memory 113, it implements the path planning method of the underwater unmanned vehicle provided by any one of the foregoing method embodiments, including: Obtain the node access sequence of the underwater unmanned vehicle, where the node access sequence includes N acquisition nodes; Take the first acquisition node as the target acquisition node, and determine the current state of the underwater unmanned vehicle at the target acquisition node; Execute the current action corresponding to the current state to obtain the next state of the underwater unmanned vehicle at the target acquisition node and the remaining data volume of the node corresponding to the next state; Update the preset initial remaining data volume of the node by using the remaining data volume of the node corresponding to the next state. Calculate the dynamic composite reward value of the current action, where the dynamic composite reward value includes the valid information of the data packet; Determine whether the remaining data volume of the updated node is 0; If the remaining data volume of the node is not 0 and within the preset communication range, use the dynamic composite reward value to update the corresponding expected return value in the preset action-value table to obtain the target action-value table; Based on the target action-value table, obtain the data acquisition path of the underwater unmanned vehicle.

[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0126] Therefore, the embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the path planning method of the underwater unmanned vehicle provided in any one of the foregoing method embodiments, including: Obtain the node access sequence of the underwater unmanned vehicle, where the node access sequence includes N acquisition nodes; Take the first acquisition node as the target acquisition node, and determine the current state of the underwater unmanned vehicle at the target acquisition node; Execute the current action corresponding to the current state to obtain the next state of the underwater unmanned vehicle at the target acquisition node and the remaining data volume of the node corresponding to the next state; Update the preset initial remaining data volume of the node by using the remaining data volume of the node corresponding to the next state; Calculate the dynamic composite reward value of the current action, where the dynamic composite reward value includes the valid information of the data packet; Determine whether the remaining data volume of the updated node is 0; If the remaining data volume of the node is not 0 and within the preset communication range, use the dynamic composite reward value to update the corresponding expected return value in the preset action-value table to obtain the target action-value table; Based on the target action-value table, obtain the data acquisition path of the underwater unmanned vehicle.

[0127] The storage medium is a physical and non-transitory storage medium, which can be various physical storage media such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. The computer-readable storage medium can be non-volatile or volatile.

[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0129] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0130] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of this application can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application.

[0132] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0133] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, provided that these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these modifications and variations.

[0134] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of various equivalent modifications or replacements, and these modifications or replacements should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A path planning method for an underwater unmanned vehicle, characterized in that, The method includes: Obtaining a node access sequence of an underwater unmanned vehicle, where the node access sequence includes N acquisition nodes; Taking the first acquisition node as the target acquisition node and determining the current state of the underwater unmanned vehicle at the target acquisition node; Executing the current action corresponding to the current state to obtain the next state of the underwater unmanned vehicle at the target acquisition node and the remaining data volume of the node corresponding to the next state; Updating a preset initial remaining data volume of the node by using the remaining data volume of the node corresponding to the next state; Calculating a dynamic composite reward value of the current action, where the dynamic composite reward value includes valid information of data packets; Judging whether the updated remaining data volume of the node is 0; If the remaining data volume of the node is not 0 and within a preset communication range, updating the corresponding expected return value in a preset action-value table by using the dynamic composite reward value to obtain a target action-value table; Based on the target action-value table, obtaining a data acquisition path of the underwater unmanned vehicle.

2. The method according to claim 1, wherein The judging whether the updated remaining data volume of the node is 0 further includes: If the updated remaining data volume of the node is 0, making the next acquisition node of the target acquisition node be the new target acquisition node, and returning to the step of determining the current state of the underwater unmanned vehicle at the target acquisition node.

3. The method according to claim 1, characterized in that, Calculating the dynamic composite reward value of the current action includes: Obtaining a first distance reward between the underwater unmanned vehicle and the target acquisition node, a direction angle reward between the underwater unmanned vehicle and the target acquisition node, a second distance reward between the underwater unmanned vehicle and the next acquisition node of the target acquisition node, a valid reward for the proportion of valid data of the data packets received by the underwater unmanned vehicle, and a final reward; Based on the first distance reward, the direction angle reward, the second distance reward, the valid reward and the final reward, obtaining the dynamic composite reward value of the current action.

4. The method according to claim 3, wherein The valid reward for the proportion of valid data of the data packets received by the underwater unmanned vehicle is calculated by the following formula: ; Among them, is the effective reward, represents the sum of the data payloads of the target acquisition nodes that have been collected by the underwater unmanned vehicle, r represents the number of data packets that have been collected by the underwater unmanned vehicle, j is the index of the received data packet, is the total amount of data of the target acquisition node , is the weight coefficient, and C is the proportion of the effective data of the data packet; Among them, , where k is the length of valid data and n is the total length of the data packet.

5. The method according to claim 1, wherein The expected return value is updated by the following formula: ; Among them, represents taking an action in the current state of value, is the learning rate, represents the maximum reward estimate in the next state is the action taken in the next state is the discount factor, represents the dynamic composite reward value.

6. The method according to claim 1, wherein The determining the current state of the underwater unmanned vehicle at the target acquisition node includes: Obtaining the absolute position of the underwater unmanned vehicle relative to the origin of the geodetic coordinate system, the pitch angle and the yaw angle of the underwater unmanned vehicle, and the signal-to-noise ratio of the underwater unmanned vehicle at the current position coordinates and the adjacent position coordinates around; Based on the absolute position, the pitch angle, the yaw angle and the signal-to-noise ratio, obtaining the state space of the underwater unmanned vehicle; Selecting the current state of the underwater unmanned vehicle at the target acquisition node from the state space.

7. The method according to claim 1, characterized in that, The remaining data volume of the node corresponding to the next state is obtained in the following manner: Obtaining the payload of the data packets acquired in the next state; Based on the payload and the preset initial remaining data volume of the node, obtaining the remaining data volume of the node corresponding to the next state.

8. An underwater unmanned vehicle path planning device, characterized in that, It includes a unit for executing the method according to any one of claims 1-7.

9. A computer device, characterized in that, The computer device includes a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program, the method described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program. When the computer program is executed by a processor, the method described in any one of claims 1-7 can be implemented.

Citation Information

Patent Citations

  • Method and device for indicating path of underwater autonomous vehicle

    CN116295449A

  • Path planning method, device and equipment of unmanned ship and storage medium

    CN117664126A

  • Unmanned surface vessel path planning method and device based on EAS-Double DQN

    CN118466509A

  • Unmanned ship path planning method based on improved A star algorithm

    CN118859938A

Cited By

  • Route determination method of underwater acoustic sensor network, equipment and medium

    CN122437808A

  • A method, device and medium for determining routing of an underwater acoustic sensor network

    CN122437808B