Multi-unmanned aerial vehicle collaborative semantic communication resource allocation and trajectory optimization method
Through the multi-UAV collaborative semantic communication resource allocation and trajectory optimization method, the deep fusion problem between the UAV and the semantic communication system is solved, efficient semantic communication in resource-constrained scenarios is realized, the trajectory and resource allocation of the UAV are optimized, and the delay constraints of time-sensitive tasks are met.
Patent Information
- Application Number
- CN202510542541.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The deep integration of existing drones and semantic communication systems has problems such as resource limitations, strong coupling of dynamic channel characteristics, conflicts in semantic information transmission timeliness and task execution accuracy, and partial observability and environmental non-stationarity in multi-UAV collaborative scenarios, making it difficult for traditional optimization strategies to achieve global efficiency optimization.
Multi-UAV collaborative semantic communication resource allocation and trajectory optimization methods are adopted to build models by obtaining semantic data, delay constraints and performance indicator optimization are carried out, combined with reinforcement learning and pheromone mechanism, dynamic decision resource allocation and flight trajectory, and multi-agent reinforcement learning algorithms are used to solve sparse rewards and learning efficiency problems.
It realizes efficient semantic communication system application in resource-constrained scenarios, improves communication resource allocation capabilities, solves the calculation complexity and time problems of traditional optimization methods, quickly converge and optimizes the drone trajectory, and meets the delay constraints of time-sensitive tasks.
Smart Images

Figure CN120475445A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing technology, and specifically relates to a method for resource allocation and trajectory optimization of multi-UAV collaborative semantic communication. Background Art
[0002] Unlike traditional bit-centric communication systems, task-oriented semantic communication focuses on the "meaning" of information and its relevance to specific tasks, rather than simply data transmission. This communication model is focused on optimizing the execution of intelligent tasks. By receiving, understanding, and leveraging the sender's semantic information, it supports more intelligent and efficient task processing, avoids unnecessary data transmission, improves communication efficiency, and better meets the quality of service requirements of intelligent tasks.
[0003] Existing research on task-oriented semantic communication systems has largely focused on leveraging deep learning techniques to enhance the information recognition, extraction, and parsing capabilities of semantic codecs, while insufficient attention has been paid to the needs of resource-constrained devices. In scenarios such as the Industrial Internet of Things (IIoT) and smart wearables, low-power sensors and edge devices have limited computing and storage resources, making it difficult to independently deploy and run large-scale deep network models. Achieving efficient semantic communication under these conditions is a pressing challenge. Furthermore, traditional fixed network architectures struggle to meet the demands of disaster emergency communications or communications in remote areas. In this context, a flexible communication architecture utilizing drones as edge servers and micro-base stations has emerged as a potential solution. Drones can not only serve as edge computing devices, providing task-related semantic processing capabilities, but can also serve as temporary base stations to enhance communication coverage and provide computing and communication support for resource-constrained devices. This architecture has significant potential for application in complex scenarios.
[0004] However, the deep integration of drones and semantic communication systems still faces multiple challenges: First, the limited onboard resources of drones and the dynamic channel characteristics lead to a strong coupling between parameters such as semantic compression rate, flight trajectory, and task queue. Second, there is an inherent conflict between the timeliness requirements of semantic information transmission and the accuracy of task execution, and traditional resource allocation methods lack explicit modeling of semantic value metrics. Third, issues such as partial observability and environmental non-stationarity in multi-drone collaborative scenarios make it difficult for centralized optimization strategies to achieve global efficiency optimization. Existing research has mostly focused on optimizing semantic models in static networks, and a systematic solution has not yet been developed for the collaborative optimization of key elements such as drone kinematic constraints and multi-agent collaboration mechanisms. Summary of the Invention
[0005] In order to solve the problem of deep integration of UAVs and semantic communication systems, the present invention proposes a multi-UAV collaborative semantic communication resource allocation and trajectory optimization method.
[0006] The technical solution of the present invention is: a multi-UAV collaborative semantic communication resource allocation and trajectory optimization method comprises the following steps:
[0007] S1. Obtain semantic data and build a semantic communication model;
[0008] S2. Based on the semantic communication model, construct a delay constraint model and overall semantic performance indicators;
[0009] S3. Based on the delay constraint model and the overall semantic performance index, a nonlinear multi-constraint optimization model for the UAV is constructed;
[0010] S4. Determine the drone’s state information, observation information, action information, and reward function based on a nonlinear multi-constraint optimization model;
[0011] S5. Based on the drone’s state information, observation information, action information, and reward function, reinforcement learning is performed to complete the drone’s communication resource allocation and trajectory optimization.
[0012] Furthermore, S1 includes the following sub-steps:
[0013] S11, obtaining semantic data and extracting semantic features;
[0014] S12, compressing the semantic features to obtain compressed semantic features;
[0015] S13, performing channel coding on the compressed semantic features to obtain symbols suitable for channel transmission;
[0016] S14, transmitting symbols suitable for channel transmission to a signal receiver;
[0017] S15. Channel-decoding the received signal using the signal receiver to obtain restored semantics. S16. Inputting the restored semantics into a semantic decoder to obtain a classification result.
[0018] S17. Based on the classification results, a semantic communication model is constructed.
[0019] Furthermore, in S11, semantic features The expression is:
[0020]
[0021] Where I represents semantic data, S α (·) represents the semantic encoding network with parameter set α; in S12, the expression of the compressed semantic feature X is:
[0022]
[0023] Where Co (·) represents the semantic compression function, and o represents the semantic compression ratio;
[0024] In S13, the expression of the symbol M applicable to channel transmission is:
[0025] M=Q σ (X);
[0026] Where Q σ (·) represents the channel encoder network with parameter set σ;
[0027] In S14, the expression of the received signal Y of the signal receiver is:
[0028] Y=hM+n;
[0029] Where h represents the channel gain, n represents Gaussian white noise;
[0030] In S15, the expression of the restored semantics X′ is:
[0031]
[0032] Where, represents the channel decoder network with network parameters χ;
[0033] In S16, the expression of the classification result p is:
[0034]
[0035] Where, represents the semantic decoder with parameter set to β;
[0036] In S17, the expression of semantic communication model A is:
[0037]
[0038] Where α1 represents the first fitting variable, α2 represents the second fitting variable, and α3 represents the third fitting variable.
[0039] Furthermore, S2 includes the following sub-steps:
[0040] S21. Generate a semantic offloading task according to the semantic communication model and determine the semantic feature size of the semantic offloading task;
[0041] S22, determining the transmission delay according to the semantic feature size of the semantic offloading task;
[0042] S23, determining semantic performance indicators;
[0043] S24. Constructing a delay constraint model based on the transmission delay of the semantic offloading task;
[0044] S25. Determine the overall semantic performance index based on the semantic performance index.
[0045] Furthermore, in S21, the semantic feature size of the semantic offloading task completed by the n-th UAV to the m-th terminal device in the t-th time slot is The expression is:
[0046]
[0047] Where, Indicates the initial size of the semantic features to be transmitted before being compressed. It represents the semantic compression rate decision of all tasks that need to be transmitted by the n-th UAV in the process of scheduling the m-th terminal device in time slot t;
[0048] In S22, the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t is The expression is:
[0049]
[0050] Where, L m represents the semantic offloading task of the mth terminal device, represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, represents the transmission rate between the nth drone and the mth terminal device, P u Indicates the transmission power;
[0051] In S23, the semantic performance index of the nth UAV in time slot t The expression is:
[0052]
[0053] Where, represents the actual task execution accuracy of a single semantic task, and M represents the number of terminal devices;
[0054] In S24, the expression of the delay constraint model is:
[0055]
[0056] Where, T wait Indicates the delay of the last device starting to be served, T max represents the maximum delay constraint, T represents the maximum flight time slot, and N represents the number of UAVs;
[0057] In S25, the total semantic performance index S total The expression is:
[0058]
[0059] Furthermore, in S3, the expression of the nonlinear multi-constraint optimization model of the UAV is:
[0060]
[0061] Where S total represents the overall semantic performance index, represents the set of locations of all drones, represents the set of scheduling strategies between the UAV and the terminal device, represents semantic features, C1 represents the collision constraint of flight angle, C2 represents the collision constraint of flight speed, C3 represents the collision constraint of flight boundary, C4 represents the collision constraint of UAV, C5 represents the constraint that only N terminal devices can simultaneously accept the semantic offloading service provided by UAV in a time slot, C6 represents the constraint that UAV needs to provide services to all terminal devices, and C7 represents the constraint that UAV needs to complete the task within T max C8 means that the constraint semantic performance must be greater than the threshold. represents the horizontal direction of the nth UAV flying in time slot t, represents the flight speed of the nth UAV in time slot t, V max Indicates the maximum speed, C L represents the upper bound of the service area, represents the movement strategy of the nth UAV in time slot t, C U represents the lower bound of the service area, represents the moving strategy of the i-th iteration variable in time slot t, D min Indicates the minimum distance allowed during the flight of the drone. represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed. T represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t, wait Indicates the delay of the last device being served, T max represents the maximum delay constraint, represents the actual task execution accuracy of a single semantic task, A min represents the threshold, T represents the maximum flight time slot, N represents the number of drones, and M represents the number of terminal devices.
[0062] Furthermore, in S4, the state information of the drone adopts the global state; the global state s t The expression is:
[0063]
[0064] Where, represents the first intermediate variable, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed. represents the second intermediate variable, represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, represents the third intermediate variable, represents the mobility strategy of the nth UAV in time slot t, M represents the number of terminal devices, represents the pheromone collected by the first UAV in time slot t, represents the pheromone collected by the nth drone in time slot t, represents the pheromone collected by the nth UAV in the t-1 time slot, κ cov represents a constant, ω represents the semantic performance weight, It represents the semantic compression rate decision of all tasks that need to be transmitted by the n-th UAV in the process of scheduling the m-th terminal device in time slot t. represents the semantic performance index of the nth UAV in time slot t, represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t, κ dis It represents the penalty that the drone will receive every time it goes through a time slot, ρ c represents the out-of-bounds penalty, ρ ob represents a collision penalty;
[0065] In S4, the observation information of the UAV adopts the local observation state; the local observation state of the nth UAV in the t time slot is The expression is:
[0066]
[0067] In S4, the action information of the nth UAV in time slot t The expression is:
[0068]
[0069] Where, represents the flight speed of the nth UAV in time slot t, represents the horizontal direction of the nth UAV flying in time slot t;
[0070] In S4, the reward function of the nth drone in time slot t is The expression is:
[0071]
[0072] Where, T max represents the maximum delay constraint, N re T represents the reward obtained by the UAV when providing semantic offloading services to all terminal devices and meeting the delay constraint. wait Indicates the delay of the last device being served, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed in the lth time slot, N represents the number of drones, represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in the lth time slot, r tanh represents the normalized reward.
[0073] Furthermore, S5 includes the following sub-steps:
[0074] S51, forming a tuple according to the drone's state information, observation information, action information, and reward function;
[0075] S52, storing the tuple into the experience replay pool until the drone completes all semantic offloading tasks;
[0076] S53, when the UAV completes all semantic offloading tasks, it randomly selects several experience data from the experience replay pool to update the parameters of the Critic network;
[0077] S54. Use the updated Critic network to complete the communication resource allocation and trajectory optimization of the UAV.
[0078] Furthermore, in S51, the expression of the tuple is Where, represents the local observation state of the nth UAV in time slot t, represents the action information of the nth UAV in time slot t, represents the reward function of the nth drone in time slot t, Represents the local observation state of the nth UAV at time slot t+1.
[0079] The beneficial effects of the present invention are:
[0080] (1) The present invention utilizes drones as a supplement to the existing task-oriented semantic communication system architecture, and can well apply the task-oriented semantic communication system in scenarios with limited communication resources;
[0081] (2) The present invention effectively completes communication resource allocation through dynamic decision-making of semantic compression rate, further enhancing the practical application capability of task-oriented semantic communication systems;
[0082] (3) The present invention utilizes a multi-agent reinforcement learning algorithm, which can effectively address the computational complexity and solution time difficulties of traditional optimization methods.
[0083] (4) The present invention is based on the pheromone mechanism and can effectively solve the problems of sparse rewards and learning efficiency in multi-agent reinforcement learning algorithms. The algorithm can converge quickly and be applied. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 Flowchart of resource allocation and trajectory optimization method for cooperative semantic communication among multiple UAVs;
[0085] Figure 2 is a neural network structure diagram;
[0086] Figure 3 The first schematic diagram of the flight path trajectory and the dynamic adjustment process of the semantic compression rate under different numbers of drones;
[0087] Figure 4 The second schematic diagram is the flight path trajectory under different numbers of drones and the dynamic adjustment process of semantic compression rate;
[0088] Figure 5 The third schematic diagram shows the flight path trajectory under different numbers of drones and the dynamic adjustment process of the semantic compression rate;
[0089] Figure 6 The fourth schematic diagram is a flight path trajectory under different numbers of drones and a dynamic adjustment process of the semantic compression rate; DETAILED DESCRIPTION
[0090] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0091] like Figure 1 As shown, the present invention provides a multi-UAV collaborative semantic communication resource allocation and trajectory optimization method, comprising the following steps:
[0092] S1. Obtain semantic data and build a semantic communication model;
[0093] S2. Based on the semantic communication model, construct a delay constraint model and overall semantic performance indicators;
[0094] S3. Based on the delay constraint model and the overall semantic performance index, a nonlinear multi-constraint optimization model for the UAV is constructed;
[0095] S4. Determine the drone’s state information, observation information, action information, and reward function based on a nonlinear multi-constraint optimization model;
[0096] S5. Based on the drone’s state information, observation information, action information, and reward function, reinforcement learning is performed to complete the drone’s communication resource allocation and trajectory optimization.
[0097] The present invention aims at the delay constraint characteristics of time-sensitive tasks in practical application scenarios such as disaster emergency response, introduces a time constraint mechanism, and proposes a nonlinear optimization problem with spatiotemporal constraints, which constrains drones to complete area coverage and data collection within a time threshold, avoiding the loss of semantic timeliness due to task backlogs. The multi-constrained nonlinear optimization problem established above is solved using a multi-agent reinforcement learning algorithm to solve the problems of traditional optimization methods in computational efficiency and practical feasibility, and complete the resource allocation and trajectory coordination problems of TOSC (task-oriented semantic communication system) in multi-UAV collaborative scenarios. The algorithm does not need to rely on the precise location information of dynamic users, but drives drones to perform adaptive trajectory optimization and resource allocation for terminal devices based on real-time channel gain information. In order to solve the problems of sparse rewards and learning efficiency in the training process of the multi-agent reinforcement learning algorithm, the idea of ant colony algorithm is borrowed, and a reinforcement learning algorithm is proposed using the acquisition and volatilization mechanism of pheromones.
[0098] In this embodiment of the present invention, S1 includes the following sub-steps:
[0099] S11, obtaining semantic data and extracting semantic features;
[0100] S12, compressing the semantic features to obtain compressed semantic features;
[0101] S13, performing channel coding on the compressed semantic features to obtain symbols suitable for channel transmission;
[0102] S14, transmitting symbols suitable for channel transmission to a signal receiver;
[0103] S15. Channel-decoding the received signal by the signal receiver to obtain restored semantics;
[0104] S16, inputting the restored semantics into the semantic decoder to obtain a classification result;
[0105] S17. Based on the classification results, a semantic communication model is constructed.
[0106] In the embodiment of the present invention, in S11, the semantic feature The expression is:
[0107]
[0108] Where I represents semantic data, S α (·) represents the semantic encoding network with parameter set α;
[0109] In S12, the expression of the compressed semantic feature X is:
[0110]
[0111] Where C o (·) represents the semantic compression function, and o represents the semantic compression ratio;
[0112] In S13, the expression of the symbol M applicable to channel transmission is:
[0113] M=Q σ (X);
[0114] Where Q σ (·) represents the channel encoder network with parameter set σ;
[0115] In S14, the expression of the received signal Y of the signal receiver is:
[0116] Y=hM+n;
[0117] Where h represents the channel gain, n represents Gaussian white noise;
[0118] In S15, the expression of the restored semantics X′ is:
[0119]
[0120] Where, represents the channel decoder network with network parameters χ;
[0121] In S16, the expression of the classification result p is:
[0122]
[0123] Where, represents the semantic decoder with parameter set to β;
[0124] In S17, the expression of semantic communication model A is:
[0125]
[0126] Where α1 represents the first fitting variable, α2 represents the second fitting variable, and α3 represents the third fitting variable.
[0127] In this embodiment of the present invention, S2 includes the following sub-steps:
[0128] S21. Generate a semantic offloading task according to the semantic communication model and determine the semantic feature size of the semantic offloading task;
[0129] S22, determining the transmission delay according to the semantic feature size of the semantic offloading task;
[0130] S23, determining semantic performance indicators;
[0131] S24. Constructing a delay constraint model based on the transmission delay of the semantic offloading task;
[0132] S25. Determine the overall semantic performance index based on the semantic performance index.
[0133] In the embodiment of the present invention, in S21, the semantic feature size of the semantic offloading task completed by the n-th drone to the m-th terminal device in the t-th time slot is The expression is:
[0134]
[0135] Where, Indicates the initial size of the semantic features to be transmitted before being compressed. It represents the semantic compression rate decision of all tasks that need to be transmitted by the n-th UAV in the process of scheduling the m-th terminal device in time slot t;
[0136] In S22, the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t is The expression is:
[0137]
[0138] Where, L m represents the semantic offloading task of the mth terminal device, represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, represents the transmission rate between the nth drone and the mth terminal device, P u Indicates the transmission power;
[0139] In S23, the semantic performance index of the nth UAV in time slot t The expression is:
[0140]
[0141] Where, represents the actual task execution accuracy of a single semantic task, and M represents the number of terminal devices;
[0142] In S24, the expression of the delay constraint model is:
[0143]
[0144] Where, Twait Indicates the delay of the last device starting to be served, T max represents the maximum delay constraint, T represents the maximum flight time slot, and N represents the number of UAVs;
[0145] In S25, the total semantic performance index S total The expression is:
[0146]
[0147] In this embodiment of the present invention, in S3, the expression of the nonlinear multi-constraint optimization model of the UAV is:
[0148]
[0149]
[0150] Where S total represents the overall semantic performance index, represents the set of locations of all drones, represents the set of scheduling strategies between the UAV and the terminal device, represents semantic features, C1 represents the collision constraint of flight angle, C2 represents the collision constraint of flight speed, C3 represents the collision constraint of flight boundary, C4 represents the collision constraint of UAV, C5 represents the constraint that only N terminal devices can simultaneously accept the semantic offloading service provided by UAV in a time slot, C6 represents the constraint that UAV needs to provide services to all terminal devices, and C7 represents the constraint that UAV needs to complete the task within T max C8 means that the constraint semantic performance must be greater than the threshold. represents the horizontal direction of the m-th UAV flying in time slot t, represents the flight speed of the nth UAV in time slot t, V max Indicates the maximum speed, C L represents the upper bound of the service area, represents the movement strategy of the nth UAV in time slot t, C U represents the lower bound of the service area, represents the moving strategy of the i-th iteration variable in time slot t, D min Indicates the minimum distance allowed during the flight of the drone. represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed. T represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t, wait Indicates the delay of the last device being served, T maxrepresents the maximum delay constraint, represents the actual task execution accuracy of a single semantic task, A min represents the threshold, T represents the maximum flight time slot, N represents the number of drones, and M represents the number of terminal devices.
[0151] It is used to encourage drones to quickly establish connections with terminal devices to provide semantic offloading services. cov is a constant. is the semantic performance weight, which is used to balance the semantic performance dimension and encourage the UAV to choose a higher compression rate. Used to constrain the UAV to complete all tasks within the delay constraint. dis It means that every time the drone goes through a time slot, it will be affected by κ dis Penalty is used to force the drone to find the best path to traverse all users as much as possible. c Represents the out-of-bounds penalty. If the drone's flight trajectory does not satisfy C3, then ρ c >0, otherwise ρ c =0. ρ ob Represents the collision penalty, including collisions between drones and collisions between drones and obstacles. If the drone's flight trajectory does not meet C4 or C5, then ρ ob >0, otherwise ρ ob =0.
[0152] In the embodiment of the present invention, in S4, the state information of the drone adopts the global state; the global state s t The expression is:
[0153]
[0154] Where, represents the first intermediate variable, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed. represents the second intermediate variable, represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, represents the third intermediate variable, represents the mobility strategy of the nth UAV in time slot t, M represents the number of terminal devices, represents the pheromone collected by the first UAV in time slot t, represents the pheromone collected by the nth drone in time slot t, represents the pheromone collected by the nth UAV in the t-1 time slot, k cov represents a constant, ω represents the semantic performance weight, It represents the semantic compression rate decision of all tasks that need to be transmitted by the n-th UAV in the process of scheduling the m-th terminal device in time slot t. represents the semantic performance index of the nth UAV in time slot t, represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t, κ dis It represents the penalty that the drone will receive every time it goes through a time slot, ρ c represents the out-of-bounds penalty, ρ ob represents a collision penalty;
[0155] In S4, the observation information of the UAV adopts the local observation state; the local observation state of the nth UAV in the t time slot is The expression is:
[0156]
[0157] In S4, the action information of the nth UAV in time slot t The expression is:
[0158]
[0159] Where, represents the flight speed of the nth UAV in time slot t, Represents the horizontal direction of the n-th UAV flying in the t-time slot; in S4, in the model established by the present invention, it is assumed that each terminal device contains some pheromones, and these pheromones can also be represented as some special data that need to be collected. When the UAV traverses the terminal devices and provides semantic unloading services for them, it will collect pheromones, and the pheromones will also be transferred to the UAV at this time. At the same time, the pheromones on the UAV will continue to evaporate, and when the flight trajectory of the UAV violates the constraints, more pheromones will also evaporate. Pheromones represent the fusion information of the UAV and the environment to a certain extent. Therefore, the present invention uses the concentration of pheromones as a reference for rewards, thereby guiding the UAV to better explore the environment and avoid training falling into local optimal solutions. Moreover, due to the volatilization mechanism of pheromones, the original sparse rewards can be converted into dense rewards, so that the intelligent body can converge better. The reward function of the n-th UAV in the t-time slot The expression is:
[0160]
[0161] Where, T max represents the maximum delay constraint, N re T represents the reward obtained by the UAV when providing semantic offloading services to all terminal devices and meeting the delay constraint. wait Indicates the delay of the last device being served, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed in the lth time slot, N represents the number of drones, represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in the lth time slot, r tanh represents the normalized reward.
[0162] In this embodiment of the present invention, S5 includes the following sub-steps:
[0163] S51, forming a tuple according to the drone's state information, observation information, action information, and reward function;
[0164] S52, storing the tuple into the experience replay pool until the drone completes all semantic offloading tasks;
[0165] S53, when the UAV completes all semantic offloading tasks, it randomly selects several experience data from the experience replay pool to update the parameters of the Critic network;
[0166] S54. Use the updated Critic network to complete the communication resource allocation and trajectory optimization of the UAV.
[0167] In the embodiment of the present invention, in S51, the expression of the tuple is Where, represents the local observation state of the nth UAV in time slot t, represents the action information of the nth UAV in time slot t, represents the reward function of the nth drone in time slot t, Represents the local observation state of the nth UAV at time slot t+1.
[0168] Table 1 compares the semantic performance of various algorithms under different numbers of terminal devices. As the number of terminal devices increases from 10 to 30, MATD3-P's cumulative semantic performance continues to improve, outperforming MADDPG-P, MASAC-P, and MADQN-P. In low-load scenarios (M = 10), the performance differences among the algorithms are small (MATD3-P leads MASAC-P by only 1.85%), indicating that when resources are sufficient, various algorithms can approach suboptimal solutions using greedy strategies. In medium-to-high-load scenarios (M ≥ 15), MATD3-P's performance advantage significantly expands. For example, when M = 20, its semantic performance improves by 3.44% over MADDPG-P, and when M = 30, the gap is reduced to 2.37%.
[0169] Table 1
[0170] Number of terminal devices MATD3-P MADDPG-P MASAC-P MADQN-P 10 761.93 744.85 748.02 725.96 15 1145.26 1134.56 1137.17 None 20 1578.35 1525.76 1514.70 None 25 1898.97 1853.97 1838.44 None 30 2183.63 2132.95 2129.84 None
[0171] Table 2 compares the cumulative semantic performance of each algorithm under different numbers of drones. Experimental results show that increasing the number of drones significantly impacts the system's semantic performance and collaborative efficiency. MATD3-P demonstrates stronger scalability and resource allocation capabilities in multi-drone collaborative scenarios. When the number of drones increases from 1 to 4, MATD3-P's cumulative semantic performance improves from 1516.00 to 1664.39, a 9.8% increase, outperforming MADDPG-P and MASAC-P. In single-drone scenarios, MATD3-P achieves a 2.2% improvement in semantic performance over MADDPG-P through dynamic compression rate adjustment and pheromone-driven progressive path planning. MADQN-P, constrained by its discrete action space, is unable to generate continuous trajectories, resulting in failure in all tasks.
[0172] Table 2
[0173] Number of drones MATD3-P MADDPG-P MASAC-P MADQN-P 1 1516.00 1482.89 1462.38 None 2 1549.43 1501.58 1475.42 None 3 1578.35 1525.76 1514.70 None 4 1664.39 1560.00 1552.09 None
[0174] Figure 2 This is the neural network structure of the MATD3-P algorithm, which includes 2N action networks, corresponding to the drone's original action network and target action network, and 4N evaluation networks (2 original evaluation networks and 2 target evaluation networks). Since most of the state dimensions in existing POMDPs are related to the coverage identifiers of terminal devices, which correspond to the number of terminal devices M, while the drone's position has only two dimensions and the pheromone has only one, there is a dimensional imbalance problem. Therefore, an expansion network is established to expand the state dimensions. The low-dimensional state (drone position and pheromone) first passes through a dense network to expand the dimension to 2M. The propagated state is then spliced with the remaining states as the input to the Actor and Critic networks.
[0175] Figure 3-6 The flight path trajectories and the dynamic adjustment process of semantic compression ratio under different numbers of drones are demonstrated. Experimental results show that the proposed MATD3-P algorithm can reasonably plan drone flight paths and semantic compression ratio decisions for scenarios with different numbers of drones, providing stable and reliable semantic offloading services.
[0176] In a single drone (N=1) scenario, Figure 3 As shown in the figure, limited by the communication bandwidth and delay constraints, the UAV adopts a higher semantic compression rate in the initial stage to reduce the transmission delay and ensure that all terminal devices are covered within a limited time. Its flight trajectory presents a progressive regional coverage feature, advancing along the direction of the spatial distribution density gradient of the terminal devices to avoid delay constraint violations caused by path detours.
[0177] In the multi-UAV scenario (N≥2), e.g. Figure 4-6As shown in Figure 2, the algorithm uses a lower semantic compression rate at the beginning of the task to improve the semantic performance weight, thereby optimizing the overall semantic performance S. total As the remaining time slots decrease, the UAV cluster dynamically improves the compression rate to a higher level. By sacrificing local semantic performance in exchange for strict satisfaction of delay constraints, the trajectory planning of the UAV cluster presents collaborative partitioning characteristics. The UAVs autonomously divide the service area according to the real-time pheromone concentration distribution, reduce path overlap and improve equipment coverage, verifying the effectiveness of the pheromone mechanism in multi-agent collaboration.
[0178] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for resource allocation and trajectory optimization of multi-UAV cooperative semantic communication, characterized in that: The following steps are involved: S1. Obtain semantic data and build a semantic communication model; S2. Based on the semantic communication model, construct a delay constraint model and overall semantic performance indicators; S3. Based on the delay constraint model and the overall semantic performance index, a nonlinear multi-constraint optimization model for the UAV is constructed; S4. Determine the drone’s state information, observation information, action information, and reward function based on a nonlinear multi-constraint optimization model; S5. Based on the drone’s state information, observation information, action information, and reward function, reinforcement learning is performed to complete the drone’s communication resource allocation and trajectory optimization.
2. The multi-UAV collaborative semantic communication resource allocation and trajectory optimization method according to claim 1 is characterized in that: The S1 comprises the following sub-steps: S11, obtaining semantic data and extracting semantic features; S12, compressing the semantic features to obtain compressed semantic features; S13, performing channel coding on the compressed semantic features to obtain symbols suitable for channel transmission; S14, transmitting symbols suitable for channel transmission to a signal receiver; S15. Using the signal receiving party, channel decode the received signal to obtain restored semantics; S16, inputting the restored semantics into the semantic decoder to obtain a classification result; S17. Based on the classification results, a semantic communication model is constructed.
3. The multi-UAV collaborative semantic communication resource allocation and trajectory optimization method according to claim 2 is characterized in that: In S11, semantic features The expression is: Where I represents semantic data, S α (·) represents the semantic encoding network with parameter set α; In S12, the expression of the compressed semantic feature X is: Where C o (·) represents the semantic compression function, and o represents the semantic compression ratio; In S13, the expression of the symbol M applicable to channel transmission is: M=Q σ (X); Where Q σ (·) represents the channel encoder network with parameter set σ; In S14, the expression of the received signal Y of the signal receiver is: Y=hM+n; Where h represents the channel gain, n represents Gaussian white noise; In S15, the expression of the restored semantics X′ is: Where, represents the channel decoder network with network parameters χ; In S16, the classification result p is expressed as: Where, represents the semantic decoder with parameter set to β; In S17, the expression of the semantic communication model A is: Where α1 represents the first fitting variable, α2 represents the second fitting variable, and α3 represents the third fitting variable.
4. The multi-UAV collaborative semantic communication resource allocation and trajectory optimization method according to claim 1 is characterized in that: The S2 includes the following sub-steps: S21. Generate a semantic offloading task according to the semantic communication model and determine the semantic feature size of the semantic offloading task; S22, determining the transmission delay according to the semantic feature size of the semantic offloading task; S23, determining semantic performance indicators; S24. Constructing a delay constraint model based on the transmission delay of the semantic offloading task; S25. Determine the overall semantic performance index based on the semantic performance index.
5. The multi-UAV collaborative semantic communication resource allocation and trajectory optimization method according to claim 4 is characterized in that: In S21, the semantic feature size of the semantic offloading task completed by the n-th UAV to the m-th terminal device in the t-th time slot is The expression is: Where, Indicates the initial size of the semantic features to be transmitted before being compressed. It represents the semantic compression rate decision of all tasks that need to be transmitted by the n-th UAV in the process of scheduling the m-th terminal device in time slot t; In S22, the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in the t time slot is The expression is: Where, L m represents the semantic offloading task of the mth terminal device, represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, represents the transmission rate between the nth drone and the mth terminal device, P u Indicates the transmission power; In S23, the semantic performance index of the nth UAV in time slot t is The expression is: Where, represents the actual task execution accuracy of a single semantic task, and M represents the number of terminal devices; In S24, the expression of the delay constraint model is: Where, T wait Indicates the delay of the last device starting to be served, T max represents the maximum delay constraint, T represents the maximum flight time slot, and N represents the number of UAVs; In S25, the total semantic performance index S total The expression is:
6. The multi-UAV cooperative semantic communication resource allocation and trajectory optimization method according to claim 1 is characterized in that: In S3, the expression of the nonlinear multi-constraint optimization model of the UAV is: st Where S total represents the overall semantic performance index, represents the set of locations of all drones, represents the set of scheduling strategies between the UAV and the terminal device, represents semantic features, C1 represents the collision constraint of flight angle, C2 represents the collision constraint of flight speed, C3 represents the collision constraint of flight boundary, C4 represents the collision constraint of UAV, C5 represents the constraint that only N terminal devices can simultaneously accept the semantic offloading service provided by UAV in a time slot, C6 represents the constraint that UAV needs to provide services to all terminal devices, and C7 represents the constraint that UAV needs to complete the task within T max C8 means that the constraint semantic performance must be greater than the threshold. represents the horizontal direction of the nth UAV flying in time slot t, represents the flight speed of the nth UAV in time slot t, V max Indicates the maximum speed, C L represents the upper bound of the service area, represents the movement strategy of the nth UAV in time slot t, C U represents the lower bound of the service area, represents the moving strategy of the i-th iteration variable in time slot t, D min Indicates the minimum distance allowed during the flight of the drone. represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed. T represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t, wait Indicates the delay of the last device being served, T max represents the maximum delay constraint, represents the actual task execution accuracy of a single semantic task, A min represents the threshold, T represents the maximum flight time slot, N represents the number of drones, and M represents the number of terminal devices.
7. The multi-UAV collaborative semantic communication resource allocation and trajectory optimization method according to claim 1 is characterized in that: In S4, the state information of the drone adopts the global state; the global state s t The expression is: Where, represents the first intermediate variable, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed. represents the second intermediate variable, represents the scheduling strategy of the nth UAV and the mth terminal device in time slot t, represents the third intermediate variable, represents the mobility strategy of the nth UAV in time slot t, M represents the number of terminal devices, represents the pheromone collected by the first UAV in time slot t, represents the pheromone collected by the nth drone in time slot t, represents the pheromone collected by the nth UAV in the t-1 time slot, k cov represents a constant, ω represents the semantic performance weight, It represents the semantic compression rate decision of all tasks that need to be transmitted by the n-th UAV in the process of scheduling the m-th terminal device in time slot t. represents the semantic performance index of the nth UAV in time slot t, represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in time slot t, κ dis It represents the penalty that the drone will receive every time it goes through a time slot, ρ c represents the out-of-bounds penalty, ρ ob represents a collision penalty; In S4, the observation information of the UAV adopts the local observation state; the local observation state of the nth UAV in the t time slot is The expression is: In S4, the action information of the nth UAV in time slot t The expression is: Where, represents the flight speed of the nth UAV in time slot t, represents the horizontal direction of the nth UAV flying in time slot t; In S4, the reward function of the nth drone in time slot t is The expression is: Where, T max represents the maximum delay constraint, N re T represents the reward obtained by the UAV when providing semantic offloading services to all terminal devices and meeting the delay constraint. wait Indicates the delay of the last device being served, Indicates whether the semantic offloading task of the nth drone to the mth terminal device is completed in the lth time slot, N represents the number of drones, represents the transmission delay of the nth UAV completing the semantic offloading task to the mth terminal device in the lth time slot, r tanh represents the normalized reward.
8. The multi-UAV collaborative semantic communication resource allocation and trajectory optimization method according to claim 1 is characterized in that: The S5 comprises the following sub-steps: S51, forming a tuple according to the drone's state information, observation information, action information, and reward function; S52, storing the tuple into the experience replay pool until the drone completes all semantic offloading tasks; S53, when the UAV completes all semantic offloading tasks, it randomly selects several experience data from the experience replay pool to update the parameters of the Critic network; S54. Use the updated Critic network to complete the communication resource allocation and trajectory optimization of the UAV.
9. The multi-UAV cooperative semantic communication resource allocation and trajectory optimization method according to claim 8 is characterized in that: In S51, the expression of the tuple is Where, represents the local observation state of the nth UAV in time slot t, represents the action information of the nth UAV in time slot t, represents the reward function of the nth drone in time slot t, Represents the local observation state of the nth UAV at time slot t+1.
Citation Information
Patent Citations
Unmanned aerial vehicle trajectory optimization method and system in Internet of Things data collection
CN113382060A
Semantic communication framework and optimization method for unmanned aerial vehicle video target detection task
CN117354865A
Unmanned aerial vehicle edge calculation anti-interference method based on semantic communication and resource allocation
CN117896756A
Resource scheduling method based on unmanned aerial vehicle assisted semantic communication energy efficiency enhancement
CN118354460A
Cited By
A multi-unmanned aerial vehicle multi-modal task offloading and deployment joint optimization method and system for semantic communication
CN122420921A