A UAV Search Path Planning Method and Device Based on D2D Communication
By building a three-dimensional environmental model in the drone and using D2D communication and consistency algorithms to make distributed collaborative decisions, the problem of inefficiency in traditional drone search path planning is solved, and efficient collaborative search of drones in complex environments is achieved.
Patent Information
- Application Number
- CN202510698307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Traditional UAV search path planning methods are inefficient in complex environments, and there are spectrum conflicts in the information sharing and collaborative decision-making of multiple UAVs, which affects communication quality and efficiency.
Using a drone search path planning method based on D2D communication, each unmanned agency builds a three-dimensional environmental model, shares information through D2D communication, and uses improved deep deterministic strategy gradient model and consistency algorithm to make distributed collaborative decisions, and sets up spectrum perception and allocation modules to avoid spectrum conflicts.
It improves the intelligence level of drone search path planning, enhances the autonomy and flexibility of drones, reduces spectrum conflicts, and improves search operation efficiency and decision-making accuracy.
Smart Images

Figure CN120220477B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of UAV search path planning, and particularly relates to a UAV search path planning method and device based on D2D communication. Background Art
[0002] With the rapid development of UAV technology, UAVs are increasingly widely used in fields such as search, rescue, and reconnaissance. Especially in complex environments such as mountainous areas and urban building clusters, the search efficiency and capabilities of UAVs are particularly important. Most traditional UAV search path planning methods rely on pre-set maps and paths. However, in actual applications, due to the complex and ever-changing environment, pre-set paths often cannot meet the actual needs, resulting in low search efficiency and even possible omission of target areas.
[0003] In addition, when using multiple UAVs for distributed cooperative search, how to ensure information sharing and cooperative decision-making among UAVs, so as to determine the search path for UAVs becomes a key issue. Traditional UAV cooperation methods usually rely on a ground control station for centralized control, which not only increases communication latency but also limits the autonomy and flexibility of UAVs. While distributed UAV cooperative decision-making schemes improve the autonomy and flexibility of UAVs to a certain extent, the algorithms for solving the global optimal solution are often very complex, greatly affected by environmental changes, and may even deviate from the global optimal solution in decision-making results.
[0004] In recent years, D2D (Device-to-Device) communication technology has been widely used in UAV communication due to its advantages such as reducing latency, improving energy efficiency, and increasing network capacity. However, when UAVs share information through D2D communication, there is a problem of spectrum conflict, which becomes one of the key factors restricting its application. Specifically, when multiple UAVs perform D2D communication on the same frequency band, due to the limited spectrum resources, serious spectrum conflicts may occur, which in turn affects communication quality and efficiency. This conflict is particularly significant when the number of UAVs is large or the communication environment is complex, and may lead to interruption or delay of information transmission, thus affecting information sharing and cooperative decision-making among UAVs and unable to achieve UAV search path planning.
[0005] Therefore, there is an urgent need for a new type of UAV search path planning method and device based on D2D communication, which can achieve information sharing and distributed cooperative decision-making among UAVs, improve the intelligent level of UAV search path planning, and thus improve the efficiency of UAV search operations. Summary of the Invention
[0006] In view of the defects existing in the above-mentioned prior art, the present invention provides a method for planning a search path of an unmanned aerial vehicle (UAV) based on device-to-device (D2D) communication. The number of UAVs is greater than or equal to two, and each UAV serves as an independent agent. The method includes the following steps:
[0007] S1: Each UAV collects surrounding environment information in real time through sensors, and constructs a three-dimensional environment model corresponding to itself based on the surrounding environment information. The three-dimensional environment model includes at least terrain information and obstacle position information, and the three-dimensional environment model is dynamically updated according to the surrounding environment information collected by the UAV.
[0008] S2: Based on the three-dimensional environment model, each UAV uses an improved deep deterministic policy gradient model for preliminary path planning to determine the preliminary optimal path of the UAV from the current position to the target area in the current three-dimensional environment model.
[0009] S3: The UAVs share information through D2D communication. The information includes the preliminary optimal path, terrain information, obstacle information, and the battery state information of the UAVs. The UAV includes a D2D communication unit, and the D2D communication unit includes a spectrum sensing module and a spectrum allocation module.
[0010] S4: Based on the shared information, the UAVs perform distributed collaborative decision-making based on a consensus algorithm to ensure that multiple UAVs maintain the same search direction and search speed during the search process, while avoiding collisions and duplicate searches. According to the results of the distributed collaborative decision-making, the search strategy of each UAV is determined, where the search strategy includes the final search path.
[0011] S5: The UAVs perform searches according to the search strategy, and update the three-dimensional environment model, then return to step S2 to re-plan the search path until multiple UAVs complete the search task.
[0012] In step S2, the improved deep deterministic policy gradient model is obtained through pre-training, and the training includes the following steps:
[0013] S21: Define the state space S and action space A of the UAV. The state space S = {P, V, G}, where P = (x, y, z) represents the position information of the UAV in three-dimensional space, and x, y, and z respectively represent the positions of the UAV on three axes. represents the speed information of the UAV in three-dimensional space, and respectively represent the speeds of the UAV on three axes. , representing the attitude information of the UAV, , and respectively represent the roll angle, pitch angle and yaw angle of the UAV; the action space A represents all possible actions taken by the UAV during flight, and the action vector is an element in the action space A, , where represents the speed change of the UAV in three-dimensional space, , , respectively represent the speed changes of the UAV in three axes, , and respectively represent the roll angle change rate, pitch angle change rate and yaw angle change rate of the UAV;
[0014] S22: Initialize the Actor network and the Critic network, where the Actor network is , and the Critic network is , where s represents the UAV state, , and respectively represent the weight parameters of the Actor network and the Critic network; the experience replay pool D is used to store tuples of state, action, reward and next state; the target network parameters and , as and 's copies, are used to stabilize the training process; set the reward function to , where is the immediate reward of the state ;
[0015] S23: Store the tuple into the experience replay pool D;
[0016] S234: Sample a batch of tuples from the experience replay pool D;
[0017] S235: Calculate the target Q value:
[0018] , ;
[0019] where, is the target Q value corresponding to this batch of tuples; is the discount factor, and are the target networks;
[0020] S236: Update the Critic network using the mean squared error loss function:
[0021] ;
[0022] Update the Actor network using the policy gradient method:
[0023] ;
[0024] where, and are both learning rates;
[0025] S237: After each time step, update the parameters of the target network using the soft update rule:
[0026] ;
[0027] ;
[0028] where, is the soft update coefficient, ; represents assignment;
[0029] Repeat the above steps until the Actor network converges.
[0030] In the said step S234, sampling a batch of tuples from the experience replay pool D specifically includes:
[0031] For each stored experience tuple , calculate its priority ;
[0032] ; ; ;
[0033] where, is the correction parameter;
[0034] Store the priority together with the experience tuple in the experience replay pool;
[0035] Sample a batch of tuples from the experience replay pool according to the priority probability distribution, and the priority probability distribution is calculated by the following formula:
[0036] ;
[0037] where, ext represents the parameter for adjusting the influence degree of the priority, and m and n represent the m-th and n-th experience tuples.
[0038] The spectrum sensing module is used to monitor and analyze the wireless communication environment of the UAV to determine the spectrum sensing result, where the spectrum sensing result includes spectrum occupancy, signal strength, and interference level;
[0039] The spectrum allocation module is used to determine the communication spectrum of the UAV according to the spectrum sensing result.
[0040] The spectrum allocation module is used to determine the communication spectrum of the UAV according to the spectrum sensing result, specifically including:
[0041] S31: The spectrum allocation module exchanges the spectrum sensing result with other UAVs through a wireless communication link, so as to obtain multiple spectrum sensing results;
[0042] S32: Based on the multiple spectrum sensing results, construct a spectrum similarity matrix, where each element in the spectrum similarity matrix represents the similarity between two spectrum resources;
[0043] S33: According to the spectral clustering algorithm and the spectrum similarity matrix, cluster the spectrum resources to obtain a spectrum clustering result;
[0044] S34: Based on the spectrum clustering result, according to the preset communication priority, determine the communication spectrum of the UAV itself;
[0045] S35: Send the communication spectrum to other UAVs through the wireless communication link for negotiation. In the case of reaching an agreement, the spectrum allocation module determines the communication spectrum as the final available communication spectrum of the UAV; if the negotiation fails, update the spectrum sensing result according to the communication spectrum, and return to step S31.
[0046] In step S4, the consensus algorithm is the Raft algorithm. The UAVs perform distributed collaborative decision-making based on the consensus algorithm to ensure that multiple UAVs maintain the same search direction and search speed during the search process, and at the same time avoid collisions and duplicate searches. According to the distributed collaborative decision result, determine the search strategy of each UAV, including:
[0047] The leader writes the received preliminary optimal path, terrain information, obstacle information, and the battery status information of the UAV as a log entry into the local log; where the leader is determined by voting among the multiple UAVs;
[0048] After the log is committed, the leader integrates the preliminary optimal path, terrain information, obstacle information, and the battery status information of each UAV, and determines whether there is a conflict in the preliminary optimal path of each UAV;
[0049] In the case of a conflict, optimize the preliminary optimal path according to the terrain information, obstacle information, and the battery status information of the drone, and use the optimized preliminary optimal path as the final search path of the drone;
[0050] In the case of no conflict, determine the preliminary optimal path of the drone as the final search path of the drone;
[0051] The leader uses the final search path as the negotiated search strategy and sends it to the corresponding drone.
[0052] In step S3, the information further includes the speed information of the drone; the leader writes the received preliminary optimal path, terrain information, obstacle information, and the battery status information and speed information of the drone as a log entry into the local log. After the log is submitted, the leader integrates the preliminary optimal path, terrain information, obstacle information, and the battery status information and speed information of the drone.
[0053] After the leader uses the final search path as the negotiated search strategy, the leader fine-tunes each of the final search paths to ensure that the search directions of each drone are consistent, and sets a unified search speed for the multiple drones according to the integrated speed information, terrain information, obstacle information, and the battery status information of the drone; the leader uses the fine-tuned final search path and the unified search speed as the negotiated search strategy and sends it to the corresponding drone.
[0054] The present invention also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the computer program.
[0055] In the present invention, each unmanned aerial vehicle (UAV) serves as an independent agent. According to the reinforcement learning algorithm, it determines the preliminary optimal search path of the three-dimensional environment model in which it is located, and shares the preliminary optimal search path, three-dimensional environment information, and its own state information with other UAVs through D2D communication. Based on the above information, multiple UAVs perform distributed collaborative decision-making through a consensus algorithm. That is, the present invention realizes, on the basis of a UAV making a path planning decision once based on the reinforcement learning algorithm, multiple UAVs performing a secondary distributed collaborative decision through the consensus algorithm. When performing the secondary distributed collaborative decision, it only needs to determine whether there is a conflict in the preliminary optimal search path, avoiding the problem of high algorithm complexity of finding the global optimal solution in the existing UAV distributed collaborative decision-making scheme, and at the same time improving the efficiency and accuracy of the UAV path planning decision. When performing reinforcement learning, a prioritized experience replay mechanism is adopted to improve the sampling efficiency, reduce the variance of parameter updates, accelerate the algorithm convergence, and make the algorithm more suitable for the scenario where the UAV has limited computing resources. In the present invention, a spectrum sensing module and a spectrum allocation module are provided in the UAV. The available communication spectrum of the UAV is determined through negotiation to ensure the accurate transmission of information such as the preliminary optimal search path of the UAV, and to avoid frequent adjustment of the collaborative decision-making parameters and conditions of the UAV caused by communication conflicts and other errors, thereby improving the robustness of the distributed collaborative decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0057] Figure 1 is a flowchart showing a method for UAV search path planning based on D2D communication according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0060] It should be understood that although terms such as first, second, and third may be used to describe... in the embodiments of the present invention, these... should not be limited to these terms. These terms are only used to distinguish... For example, without departing from the scope of the embodiments of the present invention, the first... may also be referred to as the second..., and similarly, the second... may also be referred to as the first....
[0061] It should be understood that the term "and / or" used herein is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.
[0062] Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "when...", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0063] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a commodity or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the commodity or device comprising the said element.
[0064] As Figure 1 shown, the present invention discloses a method for planning a search path of an unmanned aerial vehicle based on D2D communication, and the method includes:
[0065] Step S1: Each of the unmanned aerial vehicles collects surrounding environment information in real time through sensors, constructs a three-dimensional environment model corresponding to itself based on the surrounding environment information, the three-dimensional environment model at least includes terrain information and obstacle position information, and the three-dimensional environment model is dynamically updated according to the surrounding environment information collected by the unmanned aerial vehicle.
[0066] In step S1, the drone is equipped with a variety of high-precision and high-sensitivity sensors, such as lidar, cameras (including infrared cameras), inertial navigation systems (INS), etc. These sensors work together to collect various environmental information around the drone in real time and accurately. This information includes, but is not limited to, terrain undulations, building distributions, vegetation cover, weather conditions (such as wind speed, wind direction, visibility, etc.), and the positions and attributes of possible obstacles (such as trees, power poles, other flying objects, etc.).
[0067] Based on this rich and detailed environmental information collected, each drone independently constructs a three-dimensional environmental model corresponding to its current position and perspective. This three-dimensional environmental model is a digital virtual representation that contains at least two key elements: terrain information and obstacle position information.
[0068] Among them, the terrain information describes the surface characteristics of the search area, including the height and low undulations of the ground, slope changes, soil types, etc. This information is crucial for the drone to plan appropriate flight altitudes, speeds, and paths. For example, in mountainous or hilly areas, the drone needs to avoid flying directly over steep slopes or deep valleys to prevent collisions or excessive energy consumption. The obstacle position information details the positions, sizes, shapes, and dynamic changes (such as moving speeds, directions, etc.) of all fixed or moving obstacles that may affect the drone's flight within the search area. This information is crucial for the drone to avoid obstacles and maintain safe flight. For example, in an urban environment, the drone needs to be able to accurately identify and bypass obstacles such as buildings, wires, and trees; during flight, it also needs to continuously monitor and respond to other flying objects such as birds and drones that may appear.
[0069] Among them, the three-dimensional environmental model is not static, but can be dynamically updated according to the surrounding environmental information continuously collected by the drone. This means that when the drone encounters new terrain features or obstacles during the search process, it can immediately incorporate this information into the model, thereby updating and optimizing its environmental perception in real time.
[0070] This ability to dynamically update is crucial for the drone to maintain efficient and safe search in a complex and changing environment. It allows the drone to quickly adapt to environmental changes, timely adjust the search strategy, avoid unnecessary collisions and energy waste, thereby improving the overall search efficiency and success rate.
[0071] Step S2: Based on the three-dimensional environmental model, each drone uses an improved deep deterministic policy gradient model for preliminary path planning to determine the preliminary optimal path of the drone from its current position to the target area in the current three-dimensional environmental model.
[0072] The present invention utilizes the three-dimensional environment models constructed by each unmanned aerial vehicle (UAV), in combination with an improved deep deterministic policy gradient model, to perform preliminary path planning. This step aims to determine a preliminary optimal path for the UAV from its current position to the target area. By using this deep deterministic policy gradient model, the topographic information and obstacle position information in the three-dimensional environment model can be fully utilized to accurately evaluate the feasibility, safety, and efficiency of different paths. This enables the UAV to find a path that not only avoids obstacles but also complies with flight restrictions in complex and changing environments such as mountainous areas, urban building complexes, or the ocean. Moreover, the above-mentioned deep deterministic policy gradient model can screen out the optimal solution from numerous possible paths in a short time, greatly reducing the computational time and resource consumption for path planning, which is crucial for the rapid response and efficient execution of the UAV in search tasks.
[0073] In step S2, the improved deep deterministic policy gradient model is obtained through pre-training, and the training includes the following steps:
[0074] S21: Define the state space S and action space A of the UAV. The state space S = {P, V, G}, where P = (x, y, z) represents the position information of the UAV in three-dimensional space, and x, y, and z respectively represent the positions of the UAV on the three axes; represents the speed information of the UAV in three-dimensional space, , and respectively represent the speeds of the UAV on the three axes; , represents the attitude information of the UAV, , and respectively represent the roll angle, pitch angle, and yaw angle of the UAV; The action space A represents all possible actions taken by the UAV during flight. The action vector is an element in the action space A, , where represents the speed change situation of the UAV in three-dimensional space, , , respectively represent the speed changes of the UAV on the three axes, , and respectively represent the roll angle change rate, pitch angle change rate, and yaw angle change rate of the UAV;
[0075] S22: Initialize the Actor network and the Critic network. The Actor network is , and the Critic network is , where s represents the state of the UAV, , and represent the weight parameters of the Actor network and the Critic network respectively; the experience replay pool D, which is used to store tuples of states, actions, rewards, and next states; the target network parameters and , as and 's copies, are used to stabilize the training process; set the reward function to , where is the immediate reward of state ;
[0076] S23: Store the tuple in the experience replay pool D;
[0077] S234: Sample a batch of tuples from the experience replay pool D;
[0078] S235: Calculate the target Q value:
[0079] , ;
[0080] where is the target Q value corresponding to this batch of tuples; is the discount factor, and are the target networks;
[0081] S236: Update the Critic network using the mean squared error loss function:
[0082] ;
[0083] Update the Actor network using the policy gradient method:
[0084] ;
[0085] where and are both learning rates;
[0086] S237: After each time step, update the parameters of the target network using the soft update rule:
[0087] ;
[0088] ;
[0089] where is the soft update coefficient, ; Indicates assignment;
[0090] Repeat the above steps until the Actor network converges.
[0091] In step S234, sampling a batch of tuples from the experience replay pool D , specifically includes:
[0092] For each stored experience tuple , calculate its priority ;
[0093] ; ; ;
[0094] where, is a correction parameter;
[0095] Store the priority together with the experience tuple in the experience replay pool;
[0096] Sample a batch of tuples from the experience replay pool according to the priority probability distribution, and the priority probability distribution is calculated by the following formula:
[0097] ;
[0098] where, ext represents a parameter for adjusting the influence degree of the priority, and m and n represent the mth and nth experience tuples.
[0099] In the embodiments of the present invention, the deep deterministic policy gradient model is improved, that is, priority sampling is used to replace random sampling, and a priority experience replay mechanism is adopted to improve the sampling efficiency, reduce the variance of parameter update, accelerate the algorithm convergence, and make the algorithm more suitable for the scenario where the computing resources of the UAV are limited.
[0100] S24: Use the trained Actor network , and start generating a continuous search path from the starting state .
[0101] Step S3: The UAV performs information sharing through D2D communication, and the information includes the preliminary optimal path, terrain information, obstacle information, and the battery state information of the UAV; the UAV includes a D2D communication unit, and the D2D communication unit includes a spectrum sensing module and a spectrum allocation module.
[0102] The spectrum sensing module is used to monitor and analyze the wireless communication environment of the drone to determine the spectrum sensing result, which includes spectrum occupancy, signal strength, and interference level. The spectrum sensing module can monitor and analyze the wireless communication environment of the drone in real time. By accurately sensing the spectrum occupancy, the drone can make more efficient use of idle spectrum resources, thereby improving the overall utilization rate of the spectrum. This helps alleviate the problem of spectrum resource shortage, especially in areas or frequency bands with limited spectrum resources. Moreover, the spectrum sensing module can identify potential interference sources and signal strength changes, allowing the drone to select the best spectrum resources during communication. This helps reduce communication interruptions and errors, improving the reliability and stability of communication.
[0103] The spectrum allocation module is used to determine the communication spectrum of the drone according to the spectrum sensing result. The spectrum allocation module is used to determine the communication spectrum of the drone according to the spectrum sensing result, specifically including:
[0104] S31: The spectrum allocation module exchanges the spectrum sensing result with other drones through a wireless communication link, thereby obtaining multiple spectrum sensing results;
[0105] S32: Based on the multiple spectrum sensing results, construct a spectrum similarity matrix, where each element in the spectrum similarity matrix represents the similarity between two spectrum resources;
[0106] S33: According to the spectral clustering algorithm and the spectrum similarity matrix, cluster the spectrum resources to obtain a spectrum clustering result;
[0107] S34: Based on the spectrum clustering result, determine the communication spectrum of the drone itself according to the pre-set communication priority;
[0108] S35: Send the communication spectrum to other drones through the wireless communication link for negotiation. In the case of reaching an agreement, the spectrum allocation module determines the communication spectrum as the final available communication spectrum of the drone; if the negotiation fails, update the spectrum sensing result according to the communication spectrum, and return to step S31.
[0109] In a multi-drone communication scenario, the spectrum allocation module can ensure the fair allocation of spectrum resources. By considering the communication requirements and spectrum sensing results of each drone, the module can avoid some drones over-occupying spectrum resources while other drones face resource shortages. The spectrum allocation module can flexibly adjust the allocation of spectrum resources according to the dynamic positions and communication requirements of the drones. This flexibility helps the drones maintain efficient communication capabilities in complex environments, especially in search tasks that require quick response and high coordination.
[0110] In an embodiment of the present invention, a spectrum sensing module and a spectrum allocation module are provided in the unmanned aerial vehicle (UAV). The available communication spectrum of the UAV is determined through negotiation, ensuring the accurate transmission of information such as the initially optimal search path of the UAV, and avoiding frequent adjustment of the collaborative decision-making parameters and conditions of the UAV due to errors such as communication conflicts, thereby improving the robustness of subsequent distributed collaborative decision-making of the UAV.
[0111] Step S4: Based on the shared information, the UAVs perform distributed collaborative decision-making based on a consensus algorithm to ensure that multiple UAVs maintain the same search direction and search speed during the search process, while avoiding collisions and duplicate searches. According to the result of the distributed collaborative decision-making, the search strategy of each UAV is determined, where the search strategy includes the final search path.
[0112] In step S4, the consensus algorithm is the Raft algorithm. The UAVs perform distributed collaborative decision-making based on the consensus algorithm to ensure that multiple UAVs maintain the same search direction and search speed during the search process, while avoiding collisions and duplicate searches. According to the result of the distributed collaborative decision-making, the search strategy of each UAV is determined, including:
[0113] The leader writes the received initially optimal path, terrain information, obstacle information, and the battery status information of the UAV as log entries into the local log; where the leader is determined by voting among the multiple UAVs.
[0114] The voting process is specifically as follows:
[0115] Initially, the UAV nodes are all in the Follower state, waiting for leader election. When there is no leader or the leader fails, the Follower node will start an election timer. After the timer times out, this Follower node will convert to the Candidate state and increment the current term number (Term). The Candidate node sends a RequestVote RPC request for voting to other nodes in the cluster. Other nodes decide whether to vote based on the term number and log information in the received RPC. The Candidate node that obtains more than half of the node votes becomes the leader and starts sending heartbeat messages to maintain its status.
[0116] After the log is committed, the leader integrates the initially optimal paths, terrain information, obstacle information, and the battery status information of each UAV, and determines whether there are conflicts in the initially optimal paths of each UAV.
[0117] In the event of a conflict, optimize the preliminary optimal path according to the terrain information, obstacle information, and the battery status information of the drone, and use the optimized preliminary optimal path as the final search path of the drone.
[0118] Once a path conflict is detected, the leader will immediately activate the path optimization strategy to adjust the preliminary optimal path. Combining the integrated terrain information, obstacle information, and the battery status information of the drone, the leader adjusts the flight altitude and flight route of the initial optimal search path of the conflicting drone so that the adjusted and optimized search path no longer has conflicts.
[0119] In the case of no conflict, determine the preliminary optimal path of the drone as the final search path of the drone.
[0120] The leader uses the final search path as the negotiated search strategy and sends it to the corresponding drone.
[0121] To ensure that the drones have consistent search speeds and search directions, thus collaborating more effectively, avoiding duplicate searches and missed areas, and improving search efficiency, in step S3, when the drones share information through D2D communication, the information also includes the speed information of the drones; the leader writes the received preliminary optimal path, terrain information, obstacle information, the battery status information of the drone, and speed information as log entries into the local log. After the log is submitted, the leader integrates the preliminary optimal path, terrain information, obstacle information, the battery status information of the drone, and speed information.
[0122] After the leader uses the final search path as the negotiated search strategy, the leader fine-tunes each of the final search paths to ensure that the search directions of each drone are consistent, and sets a unified search speed for the multiple drones according to the integrated speed information, terrain information, obstacle information, and the battery status information of the drone; the leader uses the fine-tuned final search path and the unified search speed as the negotiated search strategy and sends it to the corresponding drone.
[0123] Step S5: The drones search according to the search strategy, update the three-dimensional environment model, and return to step S2 to re-plan the search path until the multiple drones complete the search task.
[0124] Since the 3D environmental model is dynamically updated, the improved Deep Deterministic Policy Gradient model also has the ability to dynamically adjust the path. When the drone encounters new obstacles or environmental changes during flight, it can immediately replan the path according to the updated environmental model to ensure that the initial optimal path of the drone in the current 3D environmental model is always optimal.
[0125] The above describes the preferred embodiments of the present invention, aiming to make the spirit of the present invention clearer and easier to understand, rather than to limit the present invention. Any modifications, substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope defined by the appended claims of the present invention.
Claims
1. A method for planning the search path of drones based on D2D communication, wherein the number of the drones is greater than or equal to two, characterized in that, Each drone acts as an independent agent, and the method includes the following steps: S1: Each drone collects real-time surrounding environment information through sensors, constructs a three-dimensional environment model corresponding to itself based on the surrounding environment information. The three-dimensional environment model at least includes terrain information and obstacle position information, and the three-dimensional environment model is dynamically updated according to the surrounding environment information collected by the drone; S2: Based on the three-dimensional environment model, each drone uses an improved deep deterministic policy gradient model for preliminary path planning to determine the preliminary optimal path of the drone from the current position to the target area in the current three-dimensional environment model; S3: The drones share information through D2D communication. The information includes the preliminary optimal path, terrain information, obstacle information, and the battery status information of the drones. The drone includes a D2D communication unit, and the D2D communication unit includes a spectrum sensing module and a spectrum allocation module; S4: Based on the shared information, the drones perform distributed collaborative decision-making based on a consensus algorithm to ensure that multiple drones maintain the same search direction and search speed during the search process, while avoiding collisions and duplicate searches. According to the distributed collaborative decision-making results, determine the search strategy for each drone, where the search strategy includes the final search path; S5: The drones search according to the search strategy, update the three-dimensional environment model, return to step S2, and re-perform search path planning until multiple drones complete the search task.
2. The method for planning a search path of an unmanned aerial vehicle based on D2D communication according to claim 1, wherein In step S2, the improved deep deterministic policy gradient model is obtained through pre-training, and the training includes the following steps: S21: Define the state space S and action space A of the drone. The state space S = {P, V, G}, where P = (x, y, z) represents the position information of the drone in three-dimensional space, and x, y, and z respectively represent the positions of the drone on three axes; represents the speed information of the drone in three-dimensional space, , and respectively represent the speeds of the drone on three axes; , represents the attitude information of the drone, , and respectively represent the roll angle, pitch angle, and yaw angle of the drone; The action space A represents all possible actions taken by the drone during flight. The action vector is an element in the action space A, , where represents the speed change of the drone in three-dimensional space, , , respectively represent the speed changes of the drone on three axes, , and respectively represent the roll angle change rate, pitch angle change rate, and yaw angle change rate of the drone; S22: Initialize the Actor network and the Critic network, where the Actor network is , and the Critic network is , where s represents the state of the UAV, , and represent the weight parameters of the Actor network and the Critic network respectively; an experience replay pool D for storing tuples of state, action, reward, and next state; target network parameters and , as copies of and for stabilizing the training process; set the reward function to , where is the immediate reward of state ; S23: Perform model training until the Actor network converges; S24: Use the trained Actor network , starting from the starting state to generate a continuous search path.
3. The method for planning a search path of a drone based on D2D communication according to claim 2, wherein In step S23, the performing model training until the Actor network converges specifically includes: For each training episode and each time step t, perform the following steps: S231: Select an action using the current Actor network and add exploration noise: , where is exploration noise; S232: Execute an action , observe the next state from the environment and the reward ; S233: Store the tuple in the experience replay pool D; S234: Sample a batch of tuples from the experience replay pool D ; S235: Calculate the target Q value: , ; Among them, is the target Q value corresponding to this batch of tuples; is the discount factor, and is the target network; S236: Update the Critic network using the mean squared error loss function: ; Update the Actor network using the policy gradient method: ; Among them, and are both learning rates; S237: After each time step, update the parameters of the target network using the soft update rule: ; ; Among them, is the soft update coefficient, ; represents assignment; Repeat the above steps until the Actor network converges.
4. The method for planning a search path of an unmanned aerial vehicle based on D2D communication according to claim 3, wherein In step S234, sampling a batch of tuples from the experience replay pool D , specifically including: For each stored experience tuple , calculate its priority ; ; ; ; Among them, is a calibration parameter; Store the priority along with the experience tuple in the experience replay pool; Sample a batch of tuples from the experience replay pool according to the priority probability distribution, and the priority probability distribution is calculated by the following formula: ; where ext represents a parameter for adjusting the influence degree of priority, and m and n represent the m-th and n-th experience tuples.
5. A method for planning a search path of an unmanned aerial vehicle based on D2D communication according to claim 1, characterized in that, The spectrum sensing module is used to monitor and analyze the wireless communication environment of the drone to determine the spectrum sensing result, and the spectrum sensing result includes spectrum occupancy, signal strength, and interference level; The spectrum allocation module is used to determine the communication spectrum of the drone according to the spectrum sensing result.
6. The method for planning a search path of a drone based on D2D communication according to claim 5, characterized in that, The spectrum allocation module is used to determine the communication spectrum of the drone according to the spectrum sensing result, specifically including: S31: The spectrum allocation module exchanges the spectrum sensing results with other UAVs via a wireless communication link, thereby obtaining multiple spectrum sensing results; S32: Construct a spectrum similarity matrix based on the multiple spectrum sensing results, where each element in the spectrum similarity matrix represents the similarity between two spectrum resources; S33: Cluster the spectrum resources according to the spectral clustering algorithm and the spectrum similarity matrix to obtain a spectrum clustering result; S34: Based on the spectrum clustering result, determine the communication spectrum of the UAV itself according to the pre-set communication priority; S35: Send the communication spectrum to other UAVs via the wireless communication link for negotiation. In the case of reaching an agreement, the spectrum allocation module determines the communication spectrum as the final available communication spectrum of the UAV; if the negotiation fails, update the spectrum sensing results according to the communication spectrum, and return to step S31.
7. The method for planning a search path of an unmanned aerial vehicle based on D2D communication according to claim 1, wherein, In step S4, the consensus algorithm is the Raft algorithm. The UAVs perform distributed collaborative decision-making based on the consensus algorithm to ensure that multiple UAVs maintain the same search direction and search speed during the search process, while avoiding collisions and duplicate searches. According to the distributed collaborative decision-making results, determine the search strategy for each UAV, including: The leader writes the received preliminary optimal path, terrain information, obstacle information, and the battery status information of the UAV as log entries into the local log; where the leader is determined by voting among the multiple UAVs; After the log is committed, the leader integrates the preliminary optimal paths, terrain information, obstacle information, and the battery status information of each UAV, and determines whether there are conflicts in the preliminary optimal paths of each UAV; In the case of conflicts, optimize the preliminary optimal path according to the terrain information, obstacle information, and the battery status information, and use the optimized preliminary optimal path as the final search path of the UAV; In the case of no conflicts, determine the preliminary optimal path of the UAV as the final search path of the UAV; The leader sends the final search path as the negotiated search strategy to the corresponding UAV.
8. The method for planning a search path of an unmanned aerial vehicle based on D2D communication according to claim 7, wherein In step S3, the information further includes the speed information of the UAV; the leader writes the received preliminary optimal path, terrain information, obstacle information, and the battery status information and speed information of the UAV as log entries into the local log. After the log is committed, the leader integrates the preliminary optimal path, terrain information, obstacle information, and the battery status information and speed information of the UAV.
9. A method for planning a search path of an unmanned aerial vehicle based on D2D communication according to claim 8, characterized in that After the leader takes the final search path as the negotiated search strategy, the leader fine-tunes each of the final search paths to ensure that the search directions of each drone are consistent, and sets a unified search speed for the multiple drones according to the integrated speed information, terrain information, obstacle information, and battery status information of the drones; the leader takes the fine-tuned final search path and the unified search speed as the negotiated search strategy and sends them to the corresponding drones.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Cooperative path planning method and system for 5G electric power inspection unmanned aerial vehicle
CN115951699A
Multi-unmanned aerial vehicle cooperative field source search trajectory planning method based on deep reinforcement learning
CN117193372A
Cited By
Task planning method for emergency communication unmanned aerial vehicle
CN122064127A