A communication system and method for integrated distributed unmanned aerial vehicles
By optimizing the drone system through time-division multiple access communication and reinforcement learning, the resource sharing problem of perception and communication tasks in the multi-drone integrated system is solved, flexible resource allocation and dynamic adjustment are achieved in different scenarios, and communication performance and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202411526483.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-10-30
AI Technical Summary
In the existing technology of multi-UAV interawareness integrated system, when perception and communication tasks share resources in the same time period, they are prone to interfere with each other and are difficult to adapt to the asymmetric requirements in different scenarios, resulting in decreased communication performance. Factors such as UAV position configuration and bandwidth allocation also affect system performance.
By adopting time division multiple access communication technology and reinforcement learning methods, and adjusting the time frame ratio of perception and communication signals and the position of drones, a multi-agent decision-making model is designed to achieve flexible resource allocation and dynamic adjustment of perception and communication tasks, and optimize the position and bandwidth allocation of drones.
It improves communication performance and resource utilization efficiency, ensures the stability and reliability of user communication networks, maximizes the total communication rate of ground users, and realizes efficient, flexible and stable communication services for UAV systems.
Smart Images

Figure CN119485229B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication and computer technology, in particular to the field of multi-unmanned aerial vehicle (UAV) assisted ground user communication, and more particularly to a sensing-integrated distributed UAV communication system and method. BACKGROUND
[0002] Sensing-integration has become a key technology for future wireless networks. Unlike the separate radar sensing system and communication system, in the sensing-integrated system, communication and sensing share the same hardware infrastructure, providing communication services for communication users while extracting user positions, moving speeds and other information from the echo signal, and further utilizing the sensing information to assist communication resource scheduling. However, the sensing performance of the sensing-integrated system is closely related to the channel environment between the transmitter and the receiver, and a long distance and obstacles will cause serious echo signal loss, which will seriously degrade the sensing performance. Unmanned aerial vehicles (UAVs) have flexible 3D mobility and high-altitude flight characteristics, and can be quickly deployed in wireless communication scenarios such as disaster relief, emergency communication and military communication, to provide good line-of-sight links for ground users and significantly improve the coverage range and performance of the communication network. Therefore, UAVs are considered as a potential air sensing-integrated platform, which can provide more controllable and efficient communication and sensing services. However, the traditional work of multi-UAV communication networks focuses on the separate design of sensing or communication systems, such as optimizing technical solutions to improve sensing accuracy or maximize communication rate. Unlike the separately designed sensing or communication system, the design of the sensing-integrated system of the UAV is affected by various complex factors, including sensing-integrated waveform design, allocation of sensing and communication bandwidth resources, UAV trajectory optimization, etc., which brings new challenges to the multi-UAV scheduling under the sensing-integrated system.
[0003] The next generation of 6G mobile communication networks will go beyond traditional communication services and provide users with additional high-precision perception services, such as user positioning and navigation. However, in traditional wireless communication systems, communication and perception functions are often separated, and hardware composition and software design are independent, resulting in low resource utilization. Integrated communication and sensing technology, which shares common hardware and signals to simultaneously perform communication data transmission and sensing on the same frequency spectrum, is considered a promising solution to reduce hardware costs and improve spectral efficiency. Integrated communication and sensing platforms based on ground base stations are easily disturbed by obstacles, which affects the sensing performance. Unlike traditional ground base station cellular networks, which are fixed on the ground, integrated communication and sensing platforms based on unmanned aerial vehicles (UAVs) have additional design freedom, which can enhance the channel state of communication users and sensing targets by optimizing the position and flight trajectory of the UAVs, thereby improving the performance of integrated communication and sensing services. Existing multi-UAV deployment optimization schemes based on integrated communication and sensing can be roughly divided into two categories: static UAV deployment and mobile UAV trajectory design. When using static UAV deployment for wireless communication, optimization methods such as topology and clustering are used to maximize the coverage range of the UAVs by optimizing the position and power allocation of multiple UAVs. Mobile UAVs optimize flight trajectories to make the UAVs closer to mobile users and targets of interest, thereby improving the performance of integrated communication and sensing.
[0004] However, whether it is the deployment of static UAVs or the trajectory optimization of mobile UAVs, when performing sensing and communication tasks, it is often assumed that sensing and communication are performed at fixed periods, which may ignore the asymmetric sensing and communication needs in actual systems. For example, in scenarios where the environment state changes rapidly (high-speed vehicle driving), the system needs to perform sensing tasks at a higher frequency to obtain the position, speed, and other information of communication users in a timely manner, while for scenarios where the environment state changes slowly (crowded concert), the position of communication users remains basically unchanged, and the system only needs to provide stable communication connections without the need for high-frequency sensing of users. Existing UAV-based integrated communication and sensing technology solutions do not consider the impact of sensing frequency on the system.
[0005] In summary, existing technologies face the challenge of sharing limited resources between communication and sensing tasks within the same time period when providing communication services using a distributed multi-UAV system. Since communication and sensing share the same time-frequency resources, frequent sensing activities can weaken communication performance. In addition, the position configuration of UAVs, user association establishment, and bandwidth allocation directly affect the communication performance of the system, and the needs of these factors vary in different scenarios. Traditional optimization methods are difficult to solve this non-convex optimization problem.
[0006] It should be noted that the background technology is only used to introduce the related information of the application, so as to help understand the technical scheme of the application, but does not mean that the related information must be prior art. In the absence of evidence that the related information has been disclosed before the application date of the application, the related information should not be regarded as prior art. SUMMARY
[0007] Therefore, the purpose of the present application is to overcome the defects of the prior art, and to provide a sensing-integrated distributed unmanned aerial vehicle communication system and method.
[0008] The purpose of the present application is achieved by the following technical solutions:
[0009] According to the first aspect of the present application, a sensing-integrated distributed unmanned aerial vehicle communication system is provided for providing communication services for users in a target environment, the system comprising a plurality of distributed distributed unmanned aerial vehicles, wherein all unmanned aerial vehicles are configured to share spectrum, hardware, signal resources, and to provide communication services for users in the target environment in each sensing period by: each unmanned aerial vehicle sending a sensing-integrated signal to users in its coverage range for communication to obtain target environment information in its coverage range, wherein the sensing-integrated signal is inserted into the corresponding communication frame according to the sensing period; constructing the state of the system based on the target environment information obtained by all unmanned aerial vehicles and the position information of the unmanned aerial vehicles, and obtaining the scheduling strategy of the unmanned aerial vehicles based on the state of the system using a pre-trained decision model, the scheduling strategy indicating the position adjustment and communication state with the users that the unmanned aerial vehicles need to perform; the decision model is trained based on reinforcement learning with the state as input and the scheduling strategy as output; all unmanned aerial vehicles adjust their positions and communication states with ground users according to the scheduling strategy output by the decision model.
[0010] According to the first aspect of the present application, the sensing-integrated signal is allocated time division resources based on time division multiple access communication technology and the frequency of obtaining target environment information is changed by adjusting the proportion of time frames occupied by sensing signals and communication signals in the sensing-integrated signal in the same time period.
[0011] According to the first aspect of the present application, the decision model is an intelligent agent obtained based on reinforcement learning, with the position information of the unmanned aerial vehicles and the target environment information as the state, and the position adjustment of the unmanned aerial vehicles, the sensing period adjustment and the communication state adjustment with the ground users as the action.
[0012] According to the first aspect of the present application, the decision model is trained based on reinforcement learning through the following steps: step S1, constructing an action reward table for each UAV, wherein each action reward table is used to save the reward values that the corresponding UAV can obtain by performing different actions in different states; step S2, repeatedly performing the following steps according to a preset number of times; step S21, each UAV randomly explores an action as a first candidate action based on the current state of the system; each UAV obtains an action with the maximum reward value from the action reward table as a second candidate action based on the current state of the system; step S21, obtaining the probability of selecting the first candidate action and the probability of selecting the second candidate action in the current round, and each UAV selects one action as a subsequent action from the two candidate actions based on the probability; step S22, each UAV performs the corresponding subsequent action and updates the state, and obtains the reward corresponding to each UAV performing the subsequent action and updates the corresponding action reward table accordingly.
[0013] According to the first aspect of the present application, the subsequent action is a combination of the following four operations: adjusting the perception period, which is adjusting the proportion of time frames occupied by the communication signal and the perception signal in the integrated sensing signal in the same period to change the perception period for obtaining target environment information; adjusting the UAV position, which is adjusting the position of the UAV in the target environment; adjusting the service state, which is adjusting the communication service state provided by the UAV for each user; adjusting the bandwidth, which is adjusting the communication bandwidth allocated by each UAV for each user.
[0014] According to the first aspect of the present application, the probability of selecting the first candidate action and the probability of selecting the second candidate action are obtained in the following manner:
[0015]
[0016] wherein, represents the iteration round, represents the probability of selecting the first candidate action in the round, represents the initial probability of selecting the first candidate action, represents the minimum probability of selecting the first candidate action, is a decay parameter, and the probability of selecting the second candidate action in the round is .
[0017] According to a first aspect of the present invention, the action reward table for each drone is updated in the following manner: a first state and a second state are obtained, wherein the first state is the state of the system before the drone performs a subsequent action, and the second state is the state of the system after the drone performs the subsequent action; an immediate reward value obtained by the drone for performing the subsequent action is determined based on a preset reward function; and an expected reward that the drone can obtain for performing the subsequent action is determined in the following manner:
[0018]
[0019] in, It represents the expected reward that the drone can obtain after performing subsequent actions. Indicates the The subsequent actions of the drone, Indicates the The instant reward value obtained by the drone performing subsequent actions, Indicates the second state, Indicates the In the action reward table of the drone, the drone can obtain the action corresponding to the maximum reward value in the second state. Indicates the The maximum reward value that the drone can obtain in the second state in the action reward table of the drone, Represents the discount factor; based on the immediate reward and expected reward, the reward value obtained by the drone for executing subsequent actions is obtained in the following way:
[0020]
[0021] in, represents the reward value obtained by the drone performing subsequent actions in the first state, Indicates the reward value corresponding to the drone performing subsequent actions in the first state saved in the reward table. Represents a preset learning rate; based on the reward value obtained by the drone performing subsequent actions in the first state, the reward value stored in the reward table for the drone performing subsequent actions in the first state is replaced to update the drone's reward table.
[0022] According to the first aspect of the present invention, the preset reward function is configured as:
[0023]
[0024]
[0025] in, Indicates the Instant reward for drones, represents the total number of drones, represents the Euclidean distance between the first tier unmanned aerial vehicle and the second tier unmanned aerial vehicle, represents the reward priority calculation function between the first tier unmanned aerial vehicle and the second tier unmanned aerial vehicle, is the total communication data rate provided by the first tier unmanned aerial vehicle for users within the target environment, represents the total number of users in the target environment, represents the association relationship between the first tier unmanned aerial vehicle and the first user, when the first tier unmanned aerial vehicle provides communication services to the first user, , on the contrary, when the first tier unmanned aerial vehicle does not provide communication services to the first user, , represents the communication bandwidth allocated by the first tier unmanned aerial vehicle for the first user, represents the signal-to-interference-and-noise ratio between the first tier unmanned aerial vehicle and the first user.
[0026] According to a first aspect of the present application, the reward priority calculation function is configured to:
[0027]
[0028] wherein, , represents the serial number of the unmanned aerial vehicle, represents the communication radius size covered by the unmanned aerial vehicle to the ground, represents the Euclidean distance between the first tier unmanned aerial vehicle and the second tier unmanned aerial vehicle.
[0029] According to a second aspect of the present application, a method for providing communication services for users in a target environment is provided, the method comprising: obtaining a sensory-integrated distributed unmanned aerial vehicle communication system according to any one of the first aspect of the present application; deploying the obtained sensory-integrated distributed unmanned aerial vehicle communication system to the target environment to provide communication services for all users in the target environment.
[0030] Compared with the prior art, the present application has the following advantages:
[0031] The application provides a sensing and communication integrated distributed unmanned aerial vehicle communication system. The system realizes flexible resource allocation of sensing and communication tasks in different communication scenarios by designing time division sensing and communication integrated signals, ensures mutual non-interference between sensing tasks and communication tasks, and dynamically adjusts sensing frequency to adapt to environmental state changes, thereby improving communication performance and resource utilization efficiency. Through reinforcement learning, multiple unmanned aerial vehicles are trained as multiple agents to adjust the positions of the unmanned aerial vehicles and the communication state with the user in real time according to the state of the target environment, thereby ensuring the stability and reliability of the user communication network in the target environment. In addition, by designing a reward function for multiple unmanned aerial vehicle cooperation, information exchange and collaborative cooperation between unmanned aerial vehicles are promoted, the coverage range of the unmanned aerial vehicles on the ground is optimized, and the total communication rate of the ground users is maximized. Overall, the application provides a communication efficient, flexible deployment and stable operation solution for a multiple unmanned aerial vehicle communication system. BRIEF DESCRIPTION OF DRAWINGS
[0032] The embodiments of the application are further described below with reference to the drawings, in which:
[0033] Figure 1 FIG. 1 is a sensing and communication integrated distributed unmanned aerial vehicle communication system architecture diagram according to an embodiment of the application;
[0034] Figure 2 FIG. 2 is a time division sensing and communication integrated frame structure according to an embodiment of the application;
[0035] Figure 3 FIG. 3 is an iteration number-total communication data rate relationship diagram according to an embodiment of the application;
[0036] Figure 4 FIG. 4 is a total communication data rate comparison diagram of a comparative experiment according to an embodiment of the application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical scheme and advantages of the application clearer, the application is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.
[0038] As mentioned in the background section, existing technologies can obtain the state information such as the position and speed of users when providing communication services for multi-UAV systems based on time division sensing-integrated signals. However, in the time division sensing-integrated signal, the communication and sensing tasks compete for limited resources in the same time period, and the communication and sensing tasks share the same time-frequency resources. Frequent sensing will reduce the communication performance. The performance of the two is directly related to each other. In actual scenarios, the asymmetry of sensing and communication requirements, the different requirements for sensing periods in different scenarios, and the close coupling between the position deployment of UAVs, the establishment of an association relationship with users, and the bandwidth allocation scheme are directly related to the communication performance of the system. The traditional distributed UAV communication system is difficult to solve this non-convex problem.
[0039] To solve the above problems, the present application provides a sensing-integrated distributed UAV communication system. The system designs a time division sensing-integrated signal to flexibly allocate resources for sensing and communication tasks in different communication scenarios, ensures the mutual interference between sensing tasks and communication tasks, and dynamically adjusts the sensing frequency to adapt to the changes in the environment state, thereby improving the communication performance and resource utilization efficiency. Through reinforcement learning, multiple UAVs are trained as multiple agents to real-time regulate the position of the UAV and the communication state with the user according to the state of the target environment, thereby ensuring the stability and reliability of the user communication network in the target environment. In addition, by designing a reward function for multi-UAV cooperation, the information exchange and cooperative work between UAVs are promoted, the coverage range of the UAV on the ground is optimized, and the total communication rate of the ground users is maximized. Overall, the present application provides a communication-efficient, flexible deployment and stable operation solution for multi-UAV communication systems.
[0040] In summary, the sensing-integrated distributed UAV communication system is used to provide communication services for users in a target environment. The architecture of the system is as shown in Figure 1 The system includes multiple distributed UAVs. All UAVs are configured to share spectrum, hardware, and signal resources, and provide communication services for users in the target environment in each sensing period as follows: each UAV sends a sensing-integrated signal to the users in its coverage range for communication to obtain the target environment information in its coverage range, wherein the sensing-integrated signal is inserted into the corresponding communication frame according to the sensing period; the state of the system is constructed based on the target environment information obtained by all UAVs and the position information of the UAVs, and a pre-trained decision model is used to obtain the scheduling strategy of the UAVs based on the state. The scheduling strategy indicates the position adjustment and communication state of the UAVs that need to be performed. The decision model takes the state of the system as input and takes the scheduling strategy as output and is trained based on reinforcement learning. All UAVs adjust their positions and communication states with ground users according to the scheduling strategy output by the decision model.
[0041] In order to better understand the present application, the following will be described in detail in combination with specific embodiments.
[0042] According to one embodiment of the present application, the integrated sensing and communication signal is a time division integrated sensing and communication signal based on time division resource allocation of time division multiple access communication technology, which realizes the information of the position, distance and speed of the user in the target environment on the basis of communication data transmission, wherein the time division integrated sensing and communication signal can change the frequency of obtaining the target environment information by adjusting the time frame proportion of the sensing signal and the communication signal in the same period. The time frame structure of the time division integrated sensing and communication signal is shown in Figure 2 The task cycle of the unmanned aerial vehicle providing communication service for the user in the target environment is T, the whole task cycle is discretized into N time slots, each time slot is a wireless frame of 10 ms, it is assumed that the position of the unmanned aerial vehicle and the position of the user do not change during each wireless frame to facilitate the design of the frame structure. Each wireless frame can be further divided into subframes with a length of 1 ms, each subframe can be functionally divided into a communication subframe and a sensing subframe, the communication subframe is used for transmitting a communication signal, and the sensing subframe is used for transmitting a sensing signal. Considering the asymmetric sensing and communication demand in the scene, the communication often needs to be performed in real time, and the sensing of the position of the user only needs to be performed periodically, therefore, the determination of the sensing period also becomes a key factor affecting the communication rate. For example, if one sensing subframe is inserted in each wireless frame, the system performs communication transmission for 9 / 10 of the time in each time slot; if one sensing subframe is inserted in every two wireless frames, the system performs communication transmission for 19 / 20 of the time in each time slot, and the communication capacity is improved by 5.6%.
[0043] According to one embodiment of the present application, in the integrated sensing and communication distributed unmanned aerial vehicle communication system, the communication between unmanned aerial vehicles adopts D2D communication, the D2D communication is device-to-device communication, which is a technology allowing two peer user nodes to directly communicate. This communication mode does not need to pass through a base station or a core network, and therefore can significantly reduce network load and delay, and improve spectrum efficiency and user experience.
[0044] According to one embodiment of the present application, the target environment information obtained by the UAVs includes: the position of each user, the bandwidth required by each user, and the signal-to-interference-and-noise ratio (SINR) of each user to each UAV, wherein the bandwidth required by each user and the SINR of each user to each UAV are transmitted to the UAVs through the sense-and-communicate integrated signal, and the position of each user is obtained based on the propagation time of the sense-and-communicate integrated signal. In detail, the UAVs transmit the sense-and-communicate integrated signal to the ground, and receive the signal reflected from the ground users. The UAVs estimate the distance between the UAVs and the users according to the propagation time of the signal. It should be noted that the position of each user is a three-dimensional coordinate, and therefore at least three UAVs are required to estimate the distance to the same user in order to determine the three-dimensional coordinate of the user. The present application dynamically adjusts the position of the UAVs and efficiently allocates communication resources by using the estimated position information of the users, so as to improve the communication quality of the users.
[0045] According to one embodiment of the present application, the decision model is an agent obtained based on reinforcement learning, with the position information of the UAVs and the target environment information as the state, and the position adjustment of the UAVs, the sensing period adjustment, and the communication state adjustment of the ground users as the action.
[0046] According to one embodiment of the present application, the state of the decision model includes the position information of each UAV, the position information of each user, the bandwidth required by each user, and the SINR of each user to the UAVs. Assuming that the number of UAVs is , the number of users in the target environment is , and the state can be represented as:
[0047]
[0048] wherein, represents the position information (i.e. the three-dimensional coordinate position) of the th user sensed by the UAVs, represents the position information (i.e. the three-dimensional coordinate position) of the th UAV, wherein the flight height of each UAV is consistent, represents the bandwidth resource required by the th user, represents the SINR between the th UAV and the th user.
[0049] According to an embodiment of the present application, all actions that can be performed by each UAV in different states constitute an action space of a decision model, which determines the actions that can be selected by the UAV in each state. These actions are the ways in which the UAV interacts with the environment and determine how the UAV moves from one state to another. In the scenario of the present application, the actions performed by each UAV are combinations of the following four operations: adjusting the sensing period, which is adjusting the proportion of time frames occupied by the communication signals and the sensing signals in the integrated sensing and communication signals in the same period to change the sensing period for obtaining information about the target environment; adjusting the UAV position, which is adjusting the position of the UAV in the target environment. Adjusting the service state, which is adjusting the state of the communication service provided by the UAV to each user; adjusting the bandwidth, which is adjusting the communication bandwidth allocated by each UAV to each user. The present application takes into account the hardware and software limitations of a single UAV, and the number of users that can be served by each UAV at the same time is limited, and at the same time, each user communicates only with one UAV to establish a communication connection.
[0050] According to an embodiment of the present application, the decision model is obtained by training based on reinforcement learning through the following steps: step S1, constructing an action reward table for each UAV, wherein each action reward table is used to save the reward values corresponding to different actions performed by the corresponding UAV in different states; step S2, repeatedly performing the following steps according to a predetermined number of times; step S21, each UAV randomly explores an action based on the current state of the system to serve as a first candidate action; each UAV obtains an action with the largest reward value from the action reward table based on the current state of the system to serve as a second candidate action; step S21, obtaining the probability of selecting the first candidate action and the probability of selecting the second candidate action in the current round, and each UAV selects one action from the two candidate actions based on the probability as a subsequent action; step S22, each UAV performs the corresponding subsequent action and updates the state, and obtains the reward corresponding to the execution of the subsequent action by each UAV and updates the corresponding action reward table accordingly.
[0051] According to an embodiment of the present application, in the step S1, in order to maximize the total communication data rate of all users in the target environment, each UAV will establish an action reward table for storing the reward values corresponding to different actions performed by the UAV in different states, for example, the reward value corresponding to the action a taken by the UAV in the state s is Q, and the greater the value of Q indicates the higher the benefit of taking the action a in the current environment state s.
[0052] According to one embodiment of the present invention, in step S21, to prevent the scheduling policy output by the decision model from falling into a local optimum, the present invention uses an ϵ-greedy strategy to select actions. This strategy balances exploration and exploitation. This means that when selecting an action, each drone must not only select the optimal action based on its current knowledge, but also explore other possible actions with a certain probability to avoid falling into a local optimum. For example, setting the probability of selecting the first candidate action ϵ = 0.1 means that there is a 10% probability of randomly exploring an action based on the current state, and a 90% probability of selecting the action with the highest reward from the action reward table based on the current system state. Furthermore, the value of ϵ decreases with increasing training cycles. As training progresses, the agent's understanding of the environment deepens, and the action reward table contains more useful policies and knowledge. By gradually reducing ϵ, the decision model leverages its learned policies more effectively, selecting actions known to yield higher rewards based on the reward table, thereby improving overall performance.
[0053] According to one embodiment of the present invention, the probability of the first candidate action being selected and the probability of the second candidate action being selected are obtained in the following manner, that is, the probability ϵ of the first candidate action being selected in each round is configured as follows:
[0054]
[0055] in, represents the iteration round, Indicates the The probability that the first alternative action is selected in the round, represents the initial probability of selecting the first alternative action, represents the minimum probability of selecting the first alternative action, is the attenuation parameter, The probability of the second alternative action being selected is As the training progresses, the value of ϵ gradually decreases, that is, the probability of the first alternative action being selected gradually decreases. The decision model is more inclined to use the useful strategies and knowledge contained in the action reward table to select the optimal action for the drone.
[0056] According to one embodiment of the present invention, the action reward table for each drone is updated in the following manner: a first state and a second state are obtained, wherein the first state is the state of the system before the drone performs a subsequent action, and the second state is the state of the system after the drone performs the subsequent action; an immediate reward value obtained by the drone for performing the subsequent action is determined based on a preset reward function; and an expected reward that the drone can obtain for performing the subsequent action is determined in the following manner:
[0057]
[0058] in, represents an expected reward value that the UAV can obtain after performing a subsequent action, represents a first state, represents a subsequent action of the UAV, represents a second state, represents an immediate reward value obtained by the UAV after performing the subsequent action, represents a second state, represents a first state, represents an action in the action-reward table of the UAV corresponding to a maximum reward value that the UAV can obtain in the second state, represents a first state, represents a maximum reward value that the UAV can obtain in the second state, represents a discount factor; based on the obtained immediate reward and expected reward, the reward value obtained by the UAV after performing the subsequent action is obtained by:
[0059]
[0060] wherein, represents a reward value obtained by the UAV after performing the subsequent action in the first state, represents a reward value corresponding to the subsequent action of the UAV in the first state saved in the reward table, represents a preset learning rate; based on the reward value obtained by the UAV after performing the subsequent action in the first state, the reward value of the UAV performing the subsequent action in the first state saved in the reward table is replaced to update the reward table of the UAV. After each round of training, the reward value table is updated to adjust the rewards that the UAV can obtain by performing different actions in different states, so that the strategy of the decision model is more accurate.
[0061] According to an embodiment of the present application, the preset reward function determines the reward that the UAV can obtain after performing an action according to the total communication rate between the UAV and the user and the position information of the UAV. For a distributed UAV communication system, in order to improve the overall performance of the system, each UAV needs to cooperate through information exchange. In the present application, the UAVs share the reward values they obtain with other UAVs in the vicinity, and the definition of the vicinity is set by a communication distance threshold. Further, the rewards of each UAV are given priority based on the relative distance between the UAVs. The preset reward function is configured as:
[0062]
[0063]
[0064] wherein, represents a first state, represents an immediate reward of the UAV, represents the total number of drones, Indicates the drone and the The Euclidean distance between the two drones, Indicates the drone and the Reward priority calculation function between drones, It is The total communication data rate provided by the UAV to users in the target environment, Indicates the total number of users in the target environment, Indicates the The first drone and the The relationship between users, when A drone flew to When providing communication services to individual users, On the contrary, when The drone did not When providing communication services to individual users, , Indicates the The first drone The communication bandwidth allocated to each user, Indicates the The first drone and the The signal-to-interference-and-noise ratio between users.
[0065] According to one embodiment of the present invention, the reward priority calculation function obtains the priority of sharing the surrounding drone reward value based on the Euclidean distance evaluation between drones, and its configuration is configured as follows:
[0066]
[0067] in, 、 Indicates the serial number of the drone, Indicates the communication radius of the drone’s ground coverage. Indicates the The first drone and the The Euclidean distance between the drones, in particular, Indicates the reward priority of the drone to itself. The reward priority is configured as 1.
[0068] In order to more clearly illustrate the beneficial effects of the present invention, the inventors compared the communication rates of the distributed UAV communication system with synaesthesia integration based on reinforcement learning, the distributed UAV communication system based on random algorithm, and the distributed UAV communication system based on uniform algorithm proposed in the present invention through simulation tests to prove the beneficial effects of the present invention. Among them, the distributed UAV communication system based on random algorithm is that the UAVs are randomly distributed in the user activity area. Since there is no need to perceive the user's position, the synaesthesia integrated waveform is all used for communication transmission; the distributed UAV communication system based on uniform algorithm is that the UAVs are uniformly distributed in the user activity area. Since there is no need to perceive the user's position, the synaesthesia integrated waveform is all used for communication transmission.
[0069] The experimental environment for the simulation test described above was configured as follows: the user mobility range was a 1 km x 1 km area. A total of 100 communication users were randomly distributed within the area. User movement speeds were randomly initialized from {0 m / s, 1.5 m / s, 20 m / s}. When a user reached the area boundary, they moved in the opposite direction. Five drones provided communication services to ground users. The total system bandwidth was 20 MHz, with a bandwidth resolution of 200 kHz, meaning that the bandwidth allocated to each user was an integer multiple of 200 kHz. Table 1 summarizes the simulation parameters.
[0070] Table 1
[0071] System parameters Parameter settings Number of users 100 Number of UAVs 5 UAV flight height 350m User activity range 1000m*1000m Grid width 100m Total bandwidth of UAVs 20MHz Bandwidth resolution 200kHz System Los path loss 0.1dB System NLos path loss 21dB Discount factor 0.95 Learning rate 0.1 Maximum number of iterations 1000
[0072] The invention first verifies the convergence of the proposed technical solution. Figure 3 It can be seen from the figure that the distributed UAV communication system with integrated synaesthesia proposed in the present invention reaches convergence after 250 iterations in the process of training the decision model based on reinforcement learning, and the total communication data rate of users reaches a high value. In order to further verify the effectiveness of the proposed scheme, it is compared with the other two distributed UAV communication systems to observe the changes in the total communication data rate of communication users within 10 seconds. During this period, the users move dynamically in the activity area, and the UAVs make dynamic adjustments according to the corresponding algorithms. The simulation results are as follows: Figure 4 By calculating the average of the total user communication data rate within 10 seconds, we can obtain the results shown in Table 2 below:
[0073] Table 2
[0074] Algorithm Total communication data rate of users (average value within 10s) / bps The communication system proposed by the application 70802140.5119 Random algorithm 40923926.845 Uniform algorithm 59142798.859
[0075] It can be found from the data in the table that, compared with the distributed unmanned aerial vehicle communication system based on the random algorithm, the total user communication data rate of the distributed unmanned aerial vehicle communication system provided by the application is improved by 73.01%, and compared with the distributed unmanned aerial vehicle communication system based on the uniform algorithm, it is improved by 19.01%, so the multi-agent reinforcement learning algorithm is used to optimize the unmanned aerial vehicle group scheduling scheme, thereby improving the total communication data rate of the user, and the unmanned aerial vehicle group uses the time division sensing integrated signal to provide communication services for the ground user, and simultaneously senses the position, speed and other information of the ground user, and reasonably allocates communication resources based on the ground user state information to ensure the communication quality in the target environment.
[0076] It should be noted that although the above describes the steps in a specific order, it does not mean that the steps must be performed in the above specific order, in fact, some of the steps can be performed concurrently, or even in a changed order, as long as the required function can be achieved.
[0077] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present application.
[0078] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a magneto-optical or other optical medium, and / or any suitable combination of the foregoing. A non-transitory, or non-transmission, computer readable storage medium excludes wired, wireless, optical, acoustic or other communication links.
[0079] The embodiments of the application have been described above, the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical applications or technical improvements in the market of the embodiments, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A synaesthesia-integrated distributed UAV communication system for providing communication services to users in a target environment, characterized in that: The system includes multiple distributed drones, where all drones are configured to share spectrum, hardware, and signal resources and provide communication services to users in the target environment in each sensing cycle through the following methods: Each UAV sends a synaesthesia integrated signal to users within its coverage area to communicate and obtain target environment information within its coverage area, wherein the synaesthesia integrated signal is inserted into the corresponding communication frame according to the perception cycle; The system state is constructed based on the target environment information and the position information of all drones, and a pre-trained decision model is used to obtain the drone scheduling strategy based on the system state. The decision model takes the system state as input and the scheduling strategy composed of drone position adjustment, perception cycle adjustment and ground user communication state adjustment as output. The intelligent agent is obtained based on a preset reinforcement learning method. The preset reinforcement learning method includes the following steps: Step S1: Construct an action reward table for each drone, where each action reward table is used to store the reward value that the corresponding drone can obtain when performing different actions in different states; Step S2: Repeat the following steps according to a preset number of times; Step S21: Each UAV randomly explores an action based on the current state of the system as the first candidate action; each UAV obtains the action with the largest reward value from the action reward table based on the current state of the system as the second candidate action; Step S21: Obtain the probability of the first alternative action being selected and the probability of the second alternative action being selected in the current round. Each drone selects one action from its two alternative actions as the subsequent action based on the probabilities. Step S22: Each drone performs the corresponding subsequent action and updates the status, and obtains the reward corresponding to each subsequent action and updates the corresponding action reward table based on it; All drones adjust their positions and communication status with ground users according to the scheduling strategy output by the decision model.
2. The system according to claim 1, wherein: The synaesthesia integrated signal allocates time division resources based on time division multiple access communication technology and changes the frequency of acquiring target environment information by adjusting the time frame ratio of the perception signal and the communication signal in the synaesthesia integrated signal within the same time period.
3. The system according to claim 1, wherein: The subsequent action is a combination of the following four operations: Adjusting the perception cycle, which is to adjust the time frame ratio of the communication signal and the perception signal in the synaesthesia integrated signal within the same period to change the perception cycle for obtaining target environment information; Adjusting the drone's position, which is to adjust the drone's position within the target environment; Adjust the service status, which is to adjust the communication service status provided by the drone to each user; Adjust bandwidth, which is to adjust the communication bandwidth allocated by each drone to each user.
4. The system according to claim 1, wherein: The probability of the first alternative action being selected and the probability of the second alternative action being selected are obtained as follows: in, represents the iteration round, Indicates the The probability that the first alternative action is selected in the round, represents the initial probability of selecting the first alternative action, represents the minimum probability of selecting the first alternative action, is the attenuation parameter, The probability of the second alternative action being selected in the round is .
5. The system according to claim 1, wherein: Update the action reward table for each drone as follows: Acquire a first state and a second state, wherein the first state is the state of the system before the drone performs a subsequent action, and the second state is the state of the system after the drone performs the subsequent action; Determine the immediate reward value obtained by the drone for executing subsequent actions based on a preset reward function; The expected reward that the drone can obtain by performing subsequent actions is determined in the following way: in, It represents the expected reward that the drone can obtain after performing subsequent actions. Indicates the The subsequent actions of the drone, Indicates the The instant reward value obtained by the drone performing subsequent actions, Indicates the second state, Indicates the In the action reward table of the drone, the drone can obtain the action corresponding to the maximum reward value in the second state. Indicates the The maximum reward value that the drone can obtain in the second state in the action reward table of the drone, represents the discount factor; Based on the immediate reward and expected reward, the reward value obtained by the drone for subsequent actions is obtained in the following way: in, represents the reward value obtained by the drone performing subsequent actions in the first state, Indicates the reward value corresponding to the drone performing subsequent actions in the first state saved in the reward table. Indicates the preset learning rate; The reward value obtained by the drone performing the subsequent action in the first state is replaced with the reward value stored in the reward table so as to update the reward table of the drone.
6. The system according to claim 5, characterized in that The preset reward function is configured as: in, Indicates the Instant reward for drones, represents the total number of drones, Indicates the drone and the The Euclidean distance between the two drones, Indicates the drone and the Reward priority calculation function between drones, It is The total communication data rate provided by the UAV to users in the target environment, Indicates the total number of users in the target environment, Indicates the The first drone and the The relationship between users, when A drone to When providing communication services to individual users, On the contrary, when The drone did not When providing communication services to individual users, , Indicates the The first drone The communication bandwidth allocated to each user, Indicates the The first drone and the The signal-to-interference-and-noise ratio between users.
7. The system according to claim 6, characterized in that The reward priority calculation function is configured as follows: in, 、 Indicates the serial number of the drone, Indicates the communication radius of the drone’s ground coverage. Indicates the The first drone and the The Euclidean distance between the two drones.
8. A synaesthesia-integrated distributed UAV communication method for providing communication services to users in a target environment, characterized in that: The method comprises: Obtain the synaesthesia integrated distributed UAV communication system according to any one of claims 1 to 7; The acquired synaesthesia-integrated distributed UAV communication system is deployed to the target environment to provide communication services to all users in the target environment.
Citation Information
Patent Citations
Resource scheduling method for unmanned aerial vehicle assisted communication and inductance integrated system
CN116847460A
Unmanned aerial vehicle assisted common inductance calculation network fusion method
CN117241300A