Ground vehicle base station path dynamic programming method and system
By constructing an air-ground collaborative system model and using reinforcement learning algorithms to plan the path of ground vehicle-mounted base stations, the problems of low communication rate and energy shortage in UAV collaborative missions were solved, achieving efficient communication and energy support for UAV swarms.
Patent Information
- Application Number
- CN202211623137.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-12-16
AI Technical Summary
When drones perform collaborative tasks, they face problems such as low communication speed and energy shortage. Existing technologies are unable to effectively provide high-quality communication and energy support for dynamically moving drone swarms.
An air-ground collaborative system model is constructed, and reinforcement learning algorithms are used to plan the movement path of the ground vehicle-mounted base station. The movement trajectory of the ground vehicle-mounted base station is optimized through Markov decision process to maximize the instantaneous communication transmission rate and energy reception power of the UAV.
It enables real-time communication support and energy supply for drone swarms, improves communication quality and endurance, and increases service utility by 5.83% to 42.62% compared to other algorithms.
Smart Images

Figure CN115951703B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) communication technology, and in particular to a method and system for dynamic path planning of ground vehicle-mounted base stations. Background Technology
[0002] Unmanned aerial vehicles (UAVs), as carriers of modern commercial, industrial, and military activities, are widely used in information warfare due to their advantages such as simple structure, long endurance, low cost, strong stealth, and high security. They perform tasks such as surveillance, reconnaissance, and acting as decoys to improve the command and control capabilities of troops and the ability of multi-service joint operations. However, due to the complexity of the environment and the diversity of missions, UAVs can hardly perform missions independently. Therefore, to meet the needs of difficult missions in complex scenarios, multiple UAVs of different specifications and types are often used in collaboration to complete missions. UAV collaborative reconnaissance, collaborative communication, and various UAV collaborative mission executions have been extensively researched and applied.
[0003] When drone swarms collaborate on missions, maintaining real-time communication among the drones is crucial for ensuring their coordination. Furthermore, in some scenarios, the swarm may need to communicate with ground-based user equipment. However, most existing research on drone systems focuses on the swarm receiving signals from satellite antennas or neglects the swarm's communication needs. This can lead to insufficient communication capabilities in practical applications due to factors such as non-line-of-sight links and electromagnetic interference. To address this issue, a relay drone approach has been proposed to provide communication services to the swarm. However, this approach also presents significant challenges to the communication capabilities of the relay drones in real-world applications.
[0004] In research on multi-UAV collaborative missions, energy consumption has become another key issue that has received widespread attention. For convenience and flexibility, UAVs are often not equipped with large energy storage devices. However, when performing complex missions, UAVs need to simultaneously perform functions such as communication and reconnaissance, which places higher demands on their endurance. Summary of the Invention
[0005] To address the issues of low communication speed and energy shortage faced by drones in collaborative missions, the present invention aims to provide a method and system for dynamic path planning for ground vehicle-mounted base stations.
[0006] To achieve the above objectives, the present invention provides the following solution:
[0007] In a first aspect, the present invention provides a method for dynamic path planning of a ground-based vehicle-mounted base station, comprising:
[0008] A ground-to-ground cooperative system model is constructed between a ground-based vehicle-mounted base station and an unmanned aerial vehicle (UAV) swarm, and an objective optimization function is determined based on this model. The air-to-ground cooperative system model is constructed according to the air-to-ground cooperative system. This system is a system in which the ground-based vehicle-mounted base station provides real-time communication support and energy supply services to the UAV swarm during mission execution. The objective optimization function is an optimization function designed to maximize the instantaneous communication transmission rate and instantaneous energy reception power of the UAVs while satisfying the minimum communication and energy requirements of all UAVs as much as possible, based on the design of the ground-based vehicle-mounted base station's movement trajectory. During mission execution, both the ground-based vehicle-mounted base station and the UAVs are in motion.
[0009] The objective optimization function is discretized, and a Markov decision process is constructed based on the discretized objective optimization function.
[0010] The Markov decision process is solved using an intensity learning algorithm to determine the movement path of the ground vehicle-mounted base station.
[0011] Secondly, the present invention provides a ground vehicle-mounted base station path dynamic planning system, comprising:
[0012] The objective optimization function determination module is used to construct an air-ground cooperative system model between a ground vehicle-mounted base station and an unmanned aerial vehicle (UAV) swarm, and to determine the objective optimization function based on the air-ground cooperative system model. The air-ground cooperative system is constructed based on an air-ground cooperative system. This system is a system in which the ground vehicle-mounted base station provides real-time communication support and energy supply services to the UAV swarm during mission execution. The objective optimization function is an optimization function designed to maximize the instantaneous communication transmission rate and instantaneous energy reception power of the UAVs while satisfying the minimum communication and energy requirements of all UAVs as much as possible, based on the design of the ground vehicle-mounted base station's movement trajectory. During mission execution, both the ground vehicle-mounted base station and the UAVs are in a mobile state.
[0013] A Markov decision process construction module is used to discretize the objective optimization function and construct a Markov decision process based on the discretized objective optimization function.
[0014] The mobile path determination module is used to solve the Markov decision process using an intensity learning algorithm to determine the mobile path of the ground vehicle-mounted base station.
[0015] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0016] To address the issues of low communication rates and energy shortages faced by UAVs in collaborative missions, this invention provides a reinforcement learning-based vehicle-based station dynamic path planning (VBSDPP) method and system. First, it presents an air-ground collaborative system where a ground-based station provides communication and power to multiple UAVs in the air. Specifically, it models the real-time communication support and energy supply provided by the ground-based station to the UAVs and proves that the mobile service process of the ground-based station is a Markov process. Then, it provides a reinforcement learning-based ground-based station dynamic path planning algorithm to optimize the ground-based station path. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the dynamic path planning method for ground vehicle-mounted base stations provided in an embodiment of the present invention.
[0019] Figure 2 This is a schematic diagram of the structure of the air-ground cooperative system model provided in an embodiment of the present invention;
[0020] Figure 3 An improved Q-table form diagram provided for embodiments of the present invention;
[0021] Figure 4 An initial distribution map of the moving area provided in an embodiment of the present invention;
[0022] Figure 5 This is a comparison chart of the service utility of ground vehicle-mounted base stations under different learning algorithms provided in embodiments of the present invention;
[0023] Figure 6 This is a comparison chart of the mobile paths of ground vehicle-mounted base stations under different learning algorithms provided in embodiments of the present invention.
[0024] Figure 7 This is a comparison chart of the service utility of ground vehicle-mounted base stations under different communication weights provided in an embodiment of the present invention;
[0025] Figure 8 This is a comparison chart of the service utility of ground vehicle-mounted base stations under different numbers of movement steps, provided in an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Previous studies on using drones as communication base stations to serve ground user equipment only considered the movement of communication base stations when ground user equipment is concentrated in certain distribution areas, without considering the movement of ground user equipment itself. This approach fails to effectively provide high-quality communication services to dynamically moving ground user equipment. In the scenario envisioned in this invention, the drone swarm is constantly in motion during mission execution. Therefore, it is essential to study how to plan the movement paths of ground-based vehicle-mounted base stations to better serve the dynamic drone swarm.
[0029] Example 1
[0030] Figure 1 This is a flowchart illustrating the dynamic path planning method for ground vehicle-mounted base stations provided in an embodiment of the present invention; as shown below. Figure 1 As shown in the figure, the ground vehicle-mounted base station path dynamic planning method provided in this embodiment of the invention includes the following steps.
[0031] Step 100: Construct an air-ground cooperative system model between the ground vehicle-mounted base station and the UAV swarm, and determine the objective optimization function based on the air-ground cooperative system model; the air-ground cooperative system model is constructed based on the air-ground cooperative system; the air-ground cooperative system is a system in which the ground vehicle-mounted base station provides real-time communication support services and energy supply services to the UAV swarm in the air during the execution of the mission; the objective optimization function is an optimization function that designs the movement trajectory of the ground vehicle-mounted base station to maximize the instantaneous communication transmission rate and instantaneous energy reception power of the UAVs while satisfying the minimum communication and energy requirements of all UAVs as much as possible; wherein, during the execution of the mission, both the ground vehicle-mounted base station and the UAVs are in a moving state.
[0032] As a preferred embodiment, step 100 of this invention specifically includes:
[0033] 1.1 Problems in constructing an air-ground cooperative system model
[0034] like Figure 2As shown, this embodiment of the invention considers an air-ground cooperative system, including a ground vehicle-mounted base station and k drones. The ground vehicle-mounted base station provides communication support and energy supply services to the drone swarm responsible for performing tasks, wherein both the drone swarm and the ground vehicle-mounted base station maintain a real-time mobile state.
[0035] A drone swarm can be represented by the set U = {1, 2, 3, ..., i, ..., k}. A ground-based vehicle-mounted base station simultaneously provides communication support and power supply services to all drones. The total mission duration is T0, and at time t, the location coordinates of the ground-based vehicle-mounted base station are (X... t ,Y t The three-dimensional coordinates of the kth UAV are (x, y, y). k,t ,y k,t ,h k,t At time t, the distance between the ground vehicle-mounted base station and the drone k is... Ground-based vehicle-mounted base stations and drone swarms move within a designated area S, i.e. (X t ,Y t ),(x k,t ,y k,t Given that )∈S, the ground-based vehicle-mounted base station should avoid obstacles when moving, and the set of obstacle coordinates within the specified area is S. block ,satisfy
[0036] During communication between a ground-based vehicle-mounted base station and an unmanned aerial vehicle (UAV), there are line-of-sight (LAS) link communication and non-LAS link communication. The probability of LAS link communication and the probability of non-LAS link communication can be calculated using formula (1):
[0037]
[0038] Among them, P LoS (θ k,t Let P be the probability of line-of-sight link communication between the ground vehicle-mounted base station and the UAV k at time t. NLoS (θ k,t Let θ be the probability of non-line-of-sight communication between the ground vehicle-mounted base station and the UAV k at time t. k,t Let be the elevation angle between the ground vehicle-mounted base station and the UAV k at time t. b1 and b2 are both environmental variables, and ζ is a constant value determined by both the UAV antenna and the environment.
[0039] At time t, the power gain of the communication channel between the ground vehicle-mounted base station and the UAV k is:
[0040]
[0041] Among them, g t(k) represents the power gain of the communication channel between the ground vehicle-mounted base station and the UAV k at time t, and d t Let be the distance between the ground vehicle-mounted base station and the UAV k at time t, α0 be the channel discount factor, and K0 be the channel gain coefficient. It means that f c Where c is the communication frequency, and P is the speed of light. LoS (θ k,t Let P be the probability of line-of-sight link communication between the ground vehicle-mounted base station and the UAV k at time t. NLoS (θ k,t Let μ be the probability of non-line-of-sight communication between the ground vehicle-mounted base station and the UAV k at time t. LoS Let μ be the fading variance under line-of-sight links. NLoS This represents the fading variance under non-line-of-sight links.
[0042] Considering the potential for mutual interference between multiple drones and the ground-based vehicle-mounted base station, the signal-to-interference-plus-noise ratio (SIR) between the ground-based vehicle-mounted base station and drone k at time t is:
[0043]
[0044] Among them, Γ t (k) represents the signal-to-interference-plus-noise ratio (SIR) between the ground vehicle-mounted base station and the UAV k at time t, p t (k) represents the communication transmission power of the ground vehicle-mounted base station to the UAV k at time t. k σ represents the mutual interference between drone k and other drones. k 2 The noise received by the drone k is denoted as: σ k 2 =BN0, where B is the channel bandwidth between the ground vehicle base station and the UAV k, and N0 is the noise figure.
[0045] The instantaneous communication transmission rate between the ground vehicle-mounted base station and the UAV k at time t is:
[0046]
[0047] Where, r t (k) represents the instantaneous communication transmission rate between the ground vehicle-mounted base station and the UAV k at time t.
[0048] The ground-based vehicle-mounted base station supplies energy to the UAV via wireless power transmission. At time t, the instantaneous power received by UAV k is:
[0049]
[0050] Among them, P t(k) represents the instantaneous energy received by UAV k at time t, β0 is the power gain per unit distance, and P T This refers to the constant radio frequency power of the ground vehicle-mounted base station.
[0051] 1.2 Dynamic Planning Problem for Ground Vehicle-Mounted Base Station Location
[0052] To ensure that the drone swarm can maintain uninterrupted communication and sufficient energy to successfully complete its mission, the optimization objective provided by this embodiment of the invention is to design the movement trajectory of the ground vehicle-mounted base station, maximizing the instantaneous communication transmission rate and instantaneous energy reception power of the drones while meeting the minimum communication and energy requirements of all drones as much as possible.
[0053] At time t, the sum of the communication rates from the ground vehicle-mounted base station to the UAV swarm can be expressed as: The sum of energy provided to the drone swarm can be expressed as:
[0054] Therefore, the trajectory optimization problem can be formulated as:
[0055]
[0056] Wherein, formula (6-a) is the objective function; the objective function is a function that designs the movement trajectory of the ground vehicle base station to maximize the instantaneous communication transmission rate and instantaneous energy receiving power of the UAV while satisfying the minimum communication and energy requirements of all UAVs as much as possible. R t Let r be the sum of the communication transmission rates of the k UAVs at time t. t (i) represents the instantaneous communication transmission rate of the i-th UAV at time t; E t Let P be the sum of the energy received by the k UAVs at time t. t (i) represents the instantaneous energy received power of the i-th UAV at time t; T0 represents the total mission duration; ω1 and ω2 are both weight values.
[0057] Equation (6-b) is the constraint condition for the minimum instantaneous communication transmission rate, r min This represents the lowest instantaneous communication transmission rate.
[0058] Equation (6-c) is the minimum instantaneous energy received power constraint condition, P min This represents the lowest instantaneous energy received power.
[0059] Formula (6-d) is the constraint condition for limiting the movement range of ground vehicle-mounted base stations; (X t ,Y t Let S be the location coordinates of the ground vehicle-mounted base station at time t.area S is the area within the specified region, excluding the area occupied by obstacles. block This is the set of coordinates of obstacles within a specified area.
[0060] Step 200: Discretize the objective optimization function and construct a Markov decision process based on the discretized objective optimization function.
[0061] In the field of UAV communication optimization and path planning, reinforcement learning has been widely used due to its efficient self-learning ability. This invention employs a reinforcement learning algorithm to plan the movement path of a ground-based vehicle-mounted base station in real time. Reinforcement learning is a method where an agent learns to achieve a goal. During the learning process, the agent is not told what actions to take, but rather discovers which actions bring better rewards through self-exploration. In each interaction, the agent obtains observational information about the state of its environment and then decides on the next action. The agent also perceives reward signals from the environment, a number indicating the quality of the current state. The agent's goal is to maximize the cumulative reward, or payoff. Tasks in reinforcement learning algorithms are typically described using Markov Decision Processes (MDPs). A Markov Decision Process (MDP) is simply a cyclical process in which an agent takes actions to change its state and obtain a reward, interacting with the environment. The basic framework of a Markov Decision Process is as follows:<S,A,R> At each moment, the state of the agent is s. t ∈S, in this state, choose an action a t ∈A(s). At this point, based on the environment, the intelligent entity receives an immediate feedback, which is represented by a reward value r. t ∈R(s t ,a t ,s t+1 (Here, t represents the current time, and t+1 represents the next time), the current state, the chosen action, and the state at the next time step are all considered together, and then a new state s is entered. t+1 This series of states and actions constitutes the agent's policy π. The goal of reinforcement learning is to optimize the agent's action choices to maximize the long-term cumulative reward value of the task. (Cumulative reward value) Here, γ is a parameter learned, γ∈(0,1), called the decay rate. The decay rate represents the impact of the agent's current action on the reward value at a future time step. The smaller γ is, the smaller the impact of the agent's action on the reward value at a future time step; conversely, the larger the γ is, the greater the impact of the agent's action choice on subsequent events. To maximize the cumulative reward value G... tTo maximize the value, the "value" of the "state-action" sequence under the current strategy is estimated using the action value function. The action value function is represented by formula (7), which specifically means the cumulative reward value of all decision sequences after choosing action a in state s.
[0062]
[0063] Q π (s,a) represents the maximum cumulative reward value that can be obtained by taking action a in state s according to the policy π.
[0064] The optimization result of reinforcement learning is to obtain an optimal policy, whose action-value function can be expressed as:
[0065]
[0066] Where q π This indicates the path planning and execution strategy for ground-based vehicle-mounted base stations.
[0067] Since P1 in this model scenario is a continuous optimization problem, it is discretized and transformed into an MDP process. The time T0 for the UAV swarm to perform its mission is discretized and represented as T0 = N × δ t , where in δ t Within a time slot, the sum of the communication rates between the ground-based vehicle-mounted base station and the UAV swarm, as well as the sum of the energy provided to the UAV swarm, can be considered constant. Therefore, t in P1 is not a continuous value at this time, but is represented as a discrete value t = nδ. t Problem P1 can be rewritten as P2:
[0068]
[0069] The MDP representation for problem P2 is as follows:
[0070] 1) State Space: Unlike traditional paths where only the location of the base station is represented as the state, the location of the UAV swarm in this model is constantly changing. The relative position of the ground vehicle-mounted base station and the UAV swarm is different at different times when they are in the same location, and therefore they are not in the same state. Therefore, the position of the ground vehicle-mounted base station and the position of the UAV swarm at the nth time slot are set as the state, and the state space constructed is as follows:
[0071]
[0072] 2) Action space: The actions performed by the ground vehicle base station in the nth time slot. The ground vehicle base station can perform 5 actions, including moving east, moving south, moving west, moving north, and staying in place.
[0073] 3) Reward Rules: The reward value for the movement of the ground vehicle-mounted base station mainly consists of three parts: first, the sum of the communication rate and energy supply of the ground vehicle-mounted base station to the UAV swarm; second, the penalty value when some UAVs cannot meet the minimum instantaneous communication transmission rate and instantaneous energy reception power when the ground vehicle-mounted base station moves to a certain position at time n; and third, when the ground vehicle-mounted base station moves out of the designated area or obstacle, the reward value is directly counted as 0. The specific expression is shown in the following formula:
[0074]
[0075] in, Let n be the location of the ground vehicle-mounted base station in the nth time slot. This represents the sum of the communication rates of the UAV swarm in the nth time slot. ω1 and ω2 represent the sum of energy received by each UAV in the nth time slot, respectively, and represent the weighting coefficients of the sum of communication rates and the sum of received energy. m represents the number of UAVs that cannot meet the minimum communication or energy supply requirements. ψ punishment This is the penalty coefficient.
[0076] Step 300: Solve the Markov decision process using an intensity learning algorithm to determine the movement path of the ground vehicle-mounted base station.
[0077] In the scenario proposed in this embodiment of the invention, since the action selection of the ground vehicle-mounted base station affects the current state and subsequent action selections, the VBSDPP algorithm provided in this embodiment of the invention optimizes the solution based on Q-learning and adopts an ε-greedy strategy, which is an improvement on the greedy strategy. This strategy seeks a better choice according to the probability ε when exploring the environment, and selects the optimal strategy according to the probability of 1-ε based on the action value function, where the action value function is defined as:
[0078] Q(s n ,a n )=Q(s n ,a n )+α[r n+1 +γmax a Q(s n+1 ,a n )-Q(s n ,a n )](12).
[0079] Where α represents the learning rate of Q-learning; Q(s n ,a n ) represents the state space s n Based on the executed strategy π, take action space a. n The maximum cumulative reward value that can be obtained, sn Let a represent the state space of the nth time slot. n Let r represent the action space of the nth time slot. n+1 Let Q(s) represent the reward value for the (n+1)th time slot. n+1 ,a n ) represents the state space s n+1 Based on the executed strategy π, take action space a. n The maximum cumulative reward value that can be obtained; γ represents the discount rate, with a value range of [0,1], max a Q(s n+1 ,a n ) indicates that in state s n+1 Next, iterate through all possible actions a. n To obtain the function Q(s) n+1 a n The maximum value of ), where a is. n It can be understood as a variable, max a Indicates changing a n Obtain the function Q(s) n+1 a n The maximum value of ).
[0080] Traditional Q-learning in path planning problems only considers the agent's position as the state, and the Q-table is built only regarding the agent's position and action selection. However, in the scenario of this invention, the position of the drone swarm changes over time, resulting in different reward values for the same action performed by the ground vehicle base station at the same location at different times. Therefore, the VBSDPP algorithm proposed in this invention introduces time series as a factor into the MDP state. The state change caused by the ground vehicle base station when making action selection is related to the drone swarm's position at that time. Since the drone's position at a certain moment is known, the drone's action selection is also time-dependent. The most direct manifestation of this in Q-learning is the construction of the Q-table, such as... Figure 3 As shown, the format has been improved from position × action to (position × time) × action, which is more in line with the scenario of the present invention.
[0081] The VBSDPP algorithm is as follows:
[0082] Step 1: Scene initialization, randomly assign obstacle positions and drone positions, set the initial position of the ground vehicle base station, and the number of steps N. step Number of iterations N ite Action value function Q(s) n ,a n ) = 0.
[0083] Step 2: n ite=n ite +1;
[0084] Step 3: Where ε0 is the initial exploration coefficient, used to initialize the positions of the UAV and the ground vehicle-mounted base station;
[0085] Step 4: n step =n step +1;
[0086] Step 5: Update action selection according to the ε-greedy policy, calculate the reward value and the new state value according to formula (11), and update it to the Q table through formula (12).
[0087] Step 6: Repeat steps 4 and 5 until n step =N step .
[0088] Step 7: Repeat steps 4, 5, and 6 until n ite =N ite .
[0089] In step 3, the exploration coefficient decreases as the number of iterations increases, which is beneficial for the agent's exploration in the early stages of training and speeds up the process of finding the optimal strategy.
[0090] During map initialization, the computational complexity of calculating the reward value at a certain location at each moment in the scene is O(k1N). 2 In the algorithm proposed in this embodiment, the computational complexity of updating the Q-table is O(k2M), and the computational complexity of action selection is O(k3M), where k1 is a constant for calculating the reward value complexity, determined by numerical precision, and k2 and k3 are constants for updating the Q-table and action selection complexity, respectively, also determined by numerical precision. Based on the above results, assuming the number of iterations is K, the computational complexity of the algorithm proposed in this embodiment is O(K(k1N)). 2 The computational complexity of the greedy selection algorithm is O(k1N +k2M+k3M). 2 +k3M)).
[0091] The proposed VBSDPP algorithm was verified using a simulation platform. To verify the superior performance of the proposed scheme, the greedy search scheme and the random scheme were used as benchmark schemes, and the schemes were verified under different scenario configurations. Table 1 shows the fixed parameter values in this scenario.
[0092] Table 1 Simulation Parameters
[0093]
[0094]
[0095] In addition to the fixed parameters in Table 1 above, the number of drones in a typical scenario is set to 10. The drone swarm needs to perform two tasks, in which the ground vehicle-mounted base station moves 40 steps in each task, for a total of N steps. step =80, set the communication weight ω1 = 0.12, and the number of iterations N ite The number is 10000. The mobile area of the ground vehicle base station is 3000m×3000m. The drone swarm also moves within this range, dividing the area into 20×20=400 states. The height of the drone swarm is uniformly 200 meters. Figure 4 The initial distribution map of the mobile area is shown, where the independent diamonds represent the locations of obstacles, and the different straight lines represent the movement trajectories of each drone.
[0096] Figure 5 In this scenario, the proposed algorithm is compared with greedy and random algorithms. Compared to the other two algorithms, the VBSDPP algorithm achieves effective convergence and, after convergence, provides better service utility, improving performance by approximately 5.83% compared to the greedy learning algorithm. From an overall learning perspective, the service utility of the ground vehicle-mounted base station continuously improves with increasing training iterations. In the early stages of training, service utility fluctuates due to two main reasons: firstly, the exploration coefficient is higher in the early stages, leading to a greater probability of exploring new actions; while in the later stages of training, the exploration coefficient decreases and the optimal movement route has been found; secondly, the proposed method uses a large Q-table space, requiring a significant number of training iterations to fill the Q-table.
[0097] Figure 6 This diagram illustrates the movement of mobile base stations after training with the VBSDPP algorithm and the greedy learning algorithm. Figure 6 It can be seen that the two algorithms have similar movement trajectories at the beginning and end, and both effectively avoid obstacles. The VBSDPP algorithm, however, chooses to move northeast earlier. Figure 4 It can be concluded that the VBSDPP algorithm is more effective in this scenario.
[0098] Figure 7 The paper compares the service utility of ground-based vehicle-mounted base stations under different communication weights. It can be seen that when the communication weight is high, the Q-learning-based algorithm has a more significant advantage over the greedy learning algorithm. Specifically, when ω1 is 0.10, 0.12, and 0.20, the advantages of the Q-learning algorithm are 0.96%, 5.83%, and 42.62%, respectively.
[0099] Figure 8 This demonstrates a comparison of the service utility of ground-based vehicle-mounted base stations under different numbers of steps taken. Figure 7 It can be seen that Q-learning-based algorithms have significant advantages over greedy learning algorithms in various situations, and have better performance.
[0100] This invention addresses the communication quality and energy shortage issues faced by UAVs in collaborative missions by proposing a Q-learning-based VBSDPP algorithm. This algorithm provides real-time communication support and energy supply to UAVs by deploying ground-based vehicle-mounted base stations along their movement paths. Simulation results demonstrate that, compared to other algorithms, the VBSDPP algorithm effectively achieves real-time dynamic support for UAV communication and energy. This research confirms that reinforcement learning can effectively solve dynamic environment optimization problems, deepening our understanding of dynamic path planning algorithms. Future research will further consider scenarios where multiple ground-based vehicle-mounted base stations simultaneously serve UAV swarms and explore implementing path planning algorithms in larger-scale scenarios.
[0101] Example 2
[0102] In order to implement the method corresponding to Embodiment 1 above and achieve the corresponding functions and technical effects, a ground vehicle-mounted base station path dynamic planning system is provided below.
[0103] This invention provides a ground-based vehicle-mounted base station path dynamic planning system, comprising:
[0104] The objective optimization function determination module is used to construct an air-ground cooperative system model between a ground vehicle-mounted base station and an unmanned aerial vehicle (UAV) swarm, and to determine the objective optimization function based on the air-ground cooperative system model. The air-ground cooperative system model is constructed based on an air-ground cooperative system. This air-ground cooperative system is a system in which the ground vehicle-mounted base station provides real-time communication support and energy supply services to the UAV swarm during mission execution. The objective optimization function is an optimization function designed to maximize the instantaneous communication transmission rate and instantaneous energy reception power of the UAVs while satisfying the minimum communication and energy requirements of all UAVs as much as possible, based on the design of the ground vehicle-mounted base station's movement trajectory. During mission execution, both the ground vehicle-mounted base station and the UAVs are in a mobile state.
[0105] The Markov decision process construction module is used to discretize the objective optimization function and construct a Markov decision process based on the discretized objective optimization function.
[0106] The mobile path determination module is used to solve the Markov decision process using an intensity learning algorithm to determine the mobile path of the ground vehicle-mounted base station.
[0107] To address the communication loss and energy shortage problems faced by unmanned aerial vehicle (UAV) swarms during missions, this paper proposes a scheme to deploy vehicle-mounted mobile base stations to provide communication services and energy supply for the UAV swarms. In this model, the UAV swarm remains in real-time motion, and the ground-based vehicle-mounted base station needs to adjust its position to achieve the best service effect. This paper aims to maximize the communication transmission rate and energy transmission efficiency of the base station. First, the problem is transformed into a Markov decision process, introducing time into the state. A reinforcement learning-based vehicle-base-station dynamic path-planning (VBSDPP) algorithm is proposed to maximize service utility by learning and planning the base station's movement path. Simulation results show that the proposed scheme can effectively achieve dynamic path planning, significantly improving the communication and energy supply service utility compared to existing schemes.
[0108] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0109] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A dynamic path planning method for ground vehicle-mounted base stations, characterized in that, include: A ground-to-ground cooperative system model is constructed between a ground vehicle-mounted base station and an unmanned aerial vehicle (UAV) swarm, and an objective optimization function is determined based on the air-to-ground cooperative system model. The air-ground cooperative system model is constructed based on the air-ground cooperative system; the air-ground cooperative system is a system in which the ground vehicle-mounted base station provides real-time communication support services and energy supply services to the UAV swarm in the air during mission execution; the objective optimization function is an optimization function that designs the movement trajectory of the ground vehicle-mounted base station to maximize the instantaneous communication transmission rate and instantaneous energy reception power of the UAVs while meeting the minimum communication and energy requirements of all UAVs as much as possible; wherein, during mission execution, both the ground vehicle-mounted base station and the UAVs are in a moving state; The objective optimization function is discretized, and a Markov decision process is constructed based on the discretized objective optimization function. The Markov decision process is solved using an intensity learning algorithm to determine the movement path of the ground vehicle-mounted base station, specifically including: Step 1: Scene initialization, randomly assign obstacle positions and drone positions, set the initial position of the ground vehicle base station, and the number of steps to move. Number of iterations Action value function ; Step 2: ; Step 3: ,in, Initialize the locations of the drone and the ground vehicle-mounted base station as the initial exploration coefficients; Step 4: ; Step 5: According to Strategy update action selection, based on formula Calculate the reward value and the new state value, and use the formula Update to table Q; Step 6: Repeat steps 4 and 5 until... ; Step 7: Repeat steps 4, 5, and 6 until... ; in, Let n be the location of the ground vehicle-mounted base station in the nth time slot. This represents the sum of the communication rates of the UAV swarm in the nth time slot. This represents the sum of energy received by each UAV in the nth time slot. , These represent the weighting coefficients for the sum of communication rates and the sum of received energy, respectively. This indicates the number of drones that cannot meet the minimum communication or energy supply requirements. The penalty coefficient is... This represents the learning rate in Q-learning; In the state space The following depends on the strategy implemented. Take action space The maximum cumulative reward value that can be obtained, This represents the state space of the nth time slot. This represents the action space of the nth time slot. This represents the reward value for the (n+1)th time slot. In the state space The following depends on the strategy implemented. Take action space The maximum cumulative reward value that can be obtained; This indicates the discount rate.
2. The method for dynamic path planning of a ground vehicle-mounted base station according to claim 1, characterized in that, The objective optimization function includes an objective function and corresponding constraints; the objective optimization function is: ; Wherein, formula (a) is the objective function; the objective function is a function that designs the movement trajectory of the ground vehicle base station to maximize the instantaneous communication transmission rate and instantaneous energy receiving power of the UAV while satisfying the minimum communication and energy requirements of all UAVs as much as possible. , In order to be in At that moment, k The sum of the communication transmission rates of the drones, In order to be in At time 1, the first i The instantaneous communication transmission rate of a drone; , In order to be in At that moment, k The sum of the energy receiving power of each drone, In order to be in At time 1, the first i The instantaneous power received by a drone; Total task duration; and All are weighted values; Formula (b) represents the minimum instantaneous communication transmission rate constraint. This represents the lowest instantaneous communication transmission rate. Formula (c) represents the minimum instantaneous energy received power constraint. This represents the lowest instantaneous energy received power. Formula (d) represents the constraint on the movement range of ground vehicle-mounted base stations; In order to be in Real-time location coordinates of the ground vehicle-mounted base station. The area within the designated region, excluding the area occupied by obstacles. This is the set of coordinates of obstacles within a specified area.
3. The method for dynamic path planning of a ground vehicle-mounted base station according to claim 2, characterized in that, The objective function after discretization is: ; in, ; .
4. The method for dynamic path planning of a ground vehicle-mounted base station according to claim 3, characterized in that, The Markov decision process includes: state space, action space, and reward rules; The state space includes the position of the ground vehicle-mounted base station in the nth time slot and the position of the UAV swarm in the nth time slot; The action space includes actions performed by the ground vehicle-mounted base station in the nth time slot; the actions include moving east, moving south, moving west, moving north, and staying in place; The reward rules include: first, the sum of the communication rate and energy supply of the ground vehicle base station to the UAV swarm; second, the penalty value when some UAVs cannot meet the minimum instantaneous communication transmission rate and instantaneous energy reception power when the ground vehicle base station moves to a certain position in the nth time slot; and third, when the ground vehicle base station moves out of the designated area or obstacle, the reward value is directly counted as 0.
5. A dynamic path planning system for a ground-based vehicle-mounted base station, characterized in that, include: The objective optimization function determination module is used to construct an air-ground cooperative system model between a ground vehicle-mounted base station and an unmanned aerial vehicle (UAV) swarm, and to determine the objective optimization function based on the air-ground cooperative system model. The air-ground cooperative system is constructed based on an air-ground cooperative system. This system is a system in which the ground vehicle-mounted base station provides real-time communication support and energy supply services to the UAV swarm during mission execution. The objective optimization function is an optimization function designed to maximize the instantaneous communication transmission rate and instantaneous energy reception power of the UAVs while satisfying the minimum communication and energy requirements of all UAVs as much as possible, based on the design of the ground vehicle-mounted base station's movement trajectory. During mission execution, both the ground vehicle-mounted base station and the UAVs are in a mobile state. A Markov decision process construction module is used to discretize the objective optimization function and construct a Markov decision process based on the discretized objective optimization function. The movement path determination module is used to solve the Markov decision process using an intensity learning algorithm to determine the movement path of the ground vehicle-mounted base station, specifically including: Step 1: Scene initialization, randomly assign obstacle positions and drone positions, set the initial position of the ground vehicle base station, and the number of steps to move. Number of iterations Action value function ; Step 2: ; Step 3: ,in, Initialize the locations of the drone and the ground vehicle-mounted base station as the initial exploration coefficients; Step 4: ; Step 5: According to Strategy update action selection, based on formula Calculate the reward value and the new state value, and use the formula Update to table Q; Step 6: Repeat steps 4 and 5 until... ; Step 7: Repeat steps 4, 5, and 6 until... ; in, Let n be the location of the ground vehicle-mounted base station in the nth time slot. This represents the sum of the communication rates of the UAV swarm in the nth time slot. This represents the sum of energy received by each UAV in the nth time slot. , These represent the weighting coefficients for the sum of communication rates and the sum of received energy, respectively. This indicates the number of drones that cannot meet the minimum communication or energy supply requirements. The penalty coefficient is... This represents the learning rate in Q-learning; In the state space The following depends on the strategy implemented. Take action space The maximum cumulative reward value that can be obtained, This represents the state space of the nth time slot. This represents the action space of the nth time slot. This represents the reward value for the (n+1)th time slot. In the state space The following depends on the strategy implemented. Take action space The maximum cumulative reward value that can be obtained; This indicates the discount rate.