Cloud-side multi-unmanned aerial vehicle collaborative resource optimization method assisted by large language model
By introducing large language models and deep reinforcement learning models into the cloud-edge collaborative resource optimization system, the problem of unstable decision-making quality in UAV swarm collaborative operations is solved, achieving efficient resource scheduling and trajectory optimization, and improving the system's real-time response and resource utilization capabilities.
Patent Information
- Application Number
- CN202511245624.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-09
AI Technical Summary
In large-scale drone swarm collaborative operations, existing technologies struggle to effectively model the spatiotemporal correlation of the global state, leading to unstable decision-making quality, conflicting responses and lags in resource scheduling and trajectory optimization, and an inability to meet the real-time response requirements of low-altitude scenarios.
A cloud-edge collaborative resource optimization method assisted by a large language model is adopted. This method generates a global policy by deploying a large language model in the cloud and deploys a deep reinforcement learning model on the edge node. Combined with a collaborative feedback mechanism and a dynamic knowledge flow collaborative optimization algorithm, it achieves collaborative optimization of resource allocation, task splitting and channel parameters.
It improves the quality and adaptability of the system's global policy generation, takes into account the system's real-time response capability and resource utilization, enhances the system's robustness and adaptability, and realizes efficient collaboration and resource optimization among multiple agents.
Smart Images

Figure CN121099339A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV communication technology and relates to a cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model. Background Technology
[0002] With the rapid development of the low-altitude economy, drone swarms are increasingly being used in smart logistics, disaster monitoring, border patrol, and other fields. Large-scale drone collaborative operations require real-time processing of massive amounts of sensing data, placing extremely high demands on network bandwidth and computing resources. However, the high-speed mobility of drones, limited onboard computing power, and dynamically changing wireless channel environments severely restrict the system's high-precision real-time inference capabilities.
[0003] To address these issues, edge-cloud collaborative inference architecture has gradually become the mainstream solution: the edge is responsible for lightweight perception and feature extraction, while the cloud undertakes high-complexity global aggregation and inference tasks. However, this architecture faces multiple dynamic challenges in actual deployment: Communication reliability challenges: Doppler shift and link interruption risks caused by the high-speed movement of drones, combined with the drastic fluctuations in signal strength caused by time-varying path loss in line-of-sight (LoS) channels, make key perception features transmitted to the cloud easily distorted or lost, directly affecting the data foundation for cloud inference. Resource competition and constraint challenges: Limited shared spectrum resources lead to bandwidth competition when multiple drones transmit uplinks; the edge computing power is insufficient to independently execute high-precision tasks, requiring reliance on cloud task offloading; at the same time, the energy and maneuverability resources of drones are also strictly limited.
[0004] Existing technologies primarily rely on deep reinforcement learning (DRL) for resource scheduling and trajectory optimization. However, in highly dynamic UAV swarm scenarios, traditional DRL methods have significant limitations: they depend on local observations, making it difficult to effectively model the spatiotemporal correlations of the global state, leading to unstable decision quality; historical experience retrieval is inefficient, conflict response is delayed, and it is difficult to adapt to rapid environmental changes. On the other hand, relying solely on large language models (LLM) for centralized decision-making faces bottlenecks of high communication latency and insufficient real-time adaptability, failing to meet the immediate response requirements of low-altitude scenarios. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A large language model-assisted cloud-edge multi-UAV collaborative resource optimization method includes the following steps:
[0008] S1: Construct an edge-cloud drone collaborative inference system, consisting of one cloud node and K edge nodes;
[0009] S2: Establish a joint optimization model that maximizes the discriminant gain;
[0010] S3: Deploy a large language model on a cloud node. The large language model generates a global strategy through a planner, a memory bank, and a reflective evaluator. The global strategy includes a bandwidth allocation scheme, a task splitting ratio, and a macro trajectory planning.
[0011] S4: Deploy a deep reinforcement learning model at each edge node. The deep reinforcement learning model performs real-time optimization based on the global policy and local observation data. The real-time optimization includes trajectory fine-tuning, dynamic bandwidth adjustment, and transmit power optimization.
[0012] S5: Establish a collaborative feedback mechanism. Edge nodes will feed back the execution results to cloud nodes. Cloud nodes will update the global strategy based on the feedback results. Edge nodes will adjust the real-time optimization parameters based on the updated global strategy.
[0013] S6: Employs an actor-critic resource allocation algorithm based on dynamic knowledge flow collaborative optimization. Through a distributed perception and centralized decision-making framework, it achieves collaborative optimization of resource allocation, task ratio allocation, and channel parameter adjustment.
[0014] Furthermore, the cloud node is a large drone, which provides bandwidth allocation, task scheduling, global decision-making for large drone trajectory planning, and generates macro-level instructions; the edge node is an edge drone, which is used to collect environmental data, perform feature extraction, and proportionally distribute tasks to the large drone.
[0015] Furthermore, the joint optimization model takes maximizing the overall discrimination gain of the system as its optimization objective, and the objective function is expressed as:
[0016]
[0017] Where k represents the index of the edge drone, n represents the discretized time slot index, m represents the index of the receiving feature dimension of the large drone, and c k,n Let q[n] represent the equivalent received signal strength of the edge drone k in the nth time interval, and let q[n] represent the trajectory of the large drone in the nth time interval. k B represents the local task allocation ratio of edge drone k. k Let μ[m] represent the uplink transmission bandwidth of edge drone k, and μ[m] represent the mean value of the m-th feature dimension. This represents the perceptual distortion variance of the m-th feature dimension. This represents the variance of sensing distortion for edge drone k. This represents the additive Gaussian noise at the large drone end; N is the total number of time intervals, and K is the total number of edge drones;
[0018] The constraints include:
[0019]
[0020] C5:q[0]=[x0,y0],
[0021]
[0022] Where C1 represents the nonnegativity of the received signal strength of the large UAV; C2 represents the classification task proportion range of the edge UAV k; and C3 represents the constraint of the uplink channel bandwidth on the feature stream of the edge UAV k. M represents the average size of the features extracted per frame. k S is the frequency at which the camera captures frames. k C1 represents the uplink transmission spectral efficiency of edge drone k; C2 represents the uplink transmission bandwidth of edge drone k not exceeding the total transmission bandwidth B. t C5 is the initial position of the large UAV, where q[0] is the initial position of the large UAV, and x0 and y0 are the initial horizontal and vertical coordinates of the large UAV, respectively; C6 is the flight speed constraint of the large UAV, where q[n-1] is the position of the large UAV in the (n-1)th time interval, τ is the time interval, and V m C7 and C8 represent the maximum horizontal speed of the large drone; C7 and C8 represent the peak transmission power P of the edge drone k. max,k and average transmission power Constrained by the following: L is the number of target categories, H1 is the flight altitude of the large UAV, H2 is the flight altitude of the edge UAV k, and l k [n] represents the position of the edge drone k within time interval n. Let z be the random variable corresponding to the edge drone k in time slot n. k The expectation of [n] squared.
[0023] Furthermore, the planner generates a global policy through a three-layer architecture: the semantic parsing layer encodes the system state into a natural language description D. t The tool call layer triggers resource conflict checks and bandwidth query API tools; the strategy generation layer outputs structured bandwidth allocation, trajectory planning, and task diversion instructions.
[0024] Furthermore, the memory bank is used to store historical decision tuples. The i-th memory entry is represented as m. i =(s i ,a i ,c i ,e i metai ), N m Let s be the total number of entries to remember, where s i a represents the state of the scenario at the moment of decision-making. i c represents the decision action generated by the LLM planner in this state. i A natural language description of the reasoning process, tools invoked, and key considerations of the LLM planner when generating actions. i Meta indicates outcome evaluation. i Meta-information includes storage time, last access time, and number of times it was retrieved; memory entries include state information, executed actions, reasoning process, and result evaluation. Relevant historical records are retrieved through weighted cosine similarity and dynamically maintained using a dual-modal mechanism of conventional incremental updates and reflection-driven updates.
[0025] Furthermore, the reflective evaluator calculates the decision quality score (DQS) based on discriminant gain and constraint penalty, denoted as DQS. (κ) =G t ·exp(-γ·∑p t ), where κ is the iteration round, G t ∑p is the discrimination gain, γ is the penalty factor, and ∑p t It is the sum of penalties; when the score is lower than the preset threshold, iterative correction of the global policy is triggered, and the corrected policy is stored in the memory bank to achieve closed-loop optimization.
[0026] Furthermore, the deep reinforcement learning model adopts the Actor-Critic framework, which integrates global policies and local observation data through a multi-head self-attention mechanism. The local observation data includes the real-time channel quality, location information, and resource usage status of edge nodes.
[0027] Furthermore, in step S5, the execution results of the collaborative feedback mechanism include channel state reports, resource utilization data, and constraint satisfaction status. The collaborative feedback mechanism performs an update once every preset time interval, with the time interval ranging from 0.1 to 1 second.
[0028] Furthermore, the dynamic knowledge flow collaborative optimization actor-critic resource allocation algorithm adopts the Actor-Critic framework, with the state value function V π and action value function Q π The estimation is performed recursively; for policy π(a|s), in state s t The state value function and action value function at a given point are defined as follows:
[0029]
[0030] Where r(s) t ,a t) is the immediate reward function, a t ~π(·|s t ) is in state s t Next, sample the action according to strategy π; Q π (s t ,a t ) is the action value function. Indicates the current state s t Perform action a t After that, the next state s t+1 Based on the expected value that strategy π can bring;
[0031] The V and Q functions are learned iteratively by minimizing the mean square Bellman error (MSBE), and the policy π is optimized by maximizing the Q value, where MSBE is defined as:
[0032]
[0033] in For empirical tuples (s) t ,a t ,r t ,s t+1 It follows the expected distribution B of the empirical replay buffer. r is the predicted value of the Q-network. t V is the immediate reward for the current step. g (s t+1 The state value of the value network;
[0034] By introducing the Kullback-Leibler divergence constraint, the learning objective of the agent is formalized into a constrained optimization problem:
[0035]
[0036]
[0037] Where s represents the state. For the state s t From the distribution Sampling, Action From strategy π S In s t The expected value of subsequent random variables is calculated for all cases of "action distribution sampling". For state s t Next action Expected cumulative rewards To minimize the negative Q-function value, π is an estimator of the KL divergence. S (s t ) represents the strategy π to be optimized. SIn state s t The action distribution under π T (s t (referencing strategy π) T In state s t The action distribution is given by σ, which is the constraint threshold for KL divergence.
[0038] The beneficial effects of this invention are as follows: Firstly, the introduction of a large language model to assist global decision-making, through the collaborative work of the planner, memory bank, and reflective evaluator, improves the quality and adaptability of global policy generation, effectively addressing resource optimization needs in complex dynamic environments. Secondly, the adoption of a layered architecture of cloud and edge nodes, with the cloud responsible for global policy planning and edge nodes responsible for real-time optimization execution, balances the system's global optimization capabilities and real-time response capabilities, improving overall system performance. Thirdly, the establishment of an effective collaborative feedback mechanism enables information exchange and dynamic policy adjustment between the cloud and edge nodes, allowing the system to continuously optimize based on actual execution, thus improving robustness and adaptability. Fourthly, the use of a dynamic knowledge flow collaborative optimization actor-critic resource allocation algorithm achieves efficient collaboration and optimized resource allocation among multiple agents, improving resource utilization and task processing efficiency. Fifthly, the storage and utilization of historical decision-making experience in the memory bank, combined with the reflective evaluation mechanism, enables iterative policy optimization, allowing the system to continuously accumulate experience and improve decision-making capabilities, thereby enhancing the system's intelligence level.
[0039] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0041] Figure 1 A schematic diagram of the cloud-edge multi-UAV collaborative resource optimization method assisted by large language models;
[0042] Figure 2 A scenario diagram for an edge-cloud drone collaborative inference system;
[0043] Figure 3 This is a schematic diagram of the edge-cloud drone collaborative inference system architecture. Detailed Implementation
[0044] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0045] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0046] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0047] Example 1:
[0048] like Figure 1 As shown, this invention provides a multi-agent resource collaborative optimization method assisted by a large language model, comprising:
[0049] Step S1: Construct an edge-cloud drone collaborative inference system, comprising one cloud node and K edge nodes, such as... Figure 2 As shown. Specifically includes:
[0050] The cloud nodes are large drones, responsible for providing global decisions such as bandwidth allocation, task scheduling, and large drone trajectory planning, and generating macro-level instructions;
[0051] The horizontal position of the large UAV at time t (projected onto the same plane as the edge UAV) is: x(t) and y(t) are the horizontal and vertical coordinates at time t, respectively. Let T be a set of 1x2 matrices over the real number field, and let T be the total time.
[0052] Large drones must meet the following constraints:
[0053] q[0]=[x0,y0],
[0054]
[0055] q[0] represents the initial position of the large UAV; x0 and y0 represent the initial horizontal and vertical coordinates of the large UAV, respectively; q[n-1] represents the position of the large UAV in the (n-1)th time interval. V represents the total number of sequences of large drone trajectories; τ represents the time interval. m This represents the maximum horizontal speed of a large drone.
[0056] Edge nodes are edge drones, responsible for collecting environmental data, performing feature extraction, and distributing tasks to larger drones proportionally.
[0057] The position of the kth edge drone in time gap n is defined as... Where x k [n]、y k [n] represents the edge drone k, Let N be a set of 1x2 matrices over the real number field, and let N be a sequence of predetermined routes for edge drone k of length (N+1).
[0058] Step S2 involves establishing a joint optimization model that maximizes the discriminant gain to optimize the overall system performance. This includes:
[0059] The overall discriminant gain is defined as the average of all pairwise discriminant gains, and is defined as follows:
[0060]
[0061] Where c k,n Indicates the equivalent received signal strength. This represents the variance of k-sensing distortion in edge drones. This represents additive Gaussian noise at the large UAV end; L is the total number of target categories, l′ and l are different target categories, and G... l,l′ For pairwise discriminative gain, G l,l′ [m] represents the pairwise discriminant gain in feature dimension m, and G[m] represents the discriminant gain in feature dimension m. To represent the perceptual distortion variance, u[m] is the intrinsic discriminant on feature dimension m, and μ l [m] represents the ideal feature mean of the l-th type of target on the m-th feature dimension, μ l′ [m] represents the ideal feature mean of the l′-th target on the m-th feature dimension.
[0062] The objective function, which aims to maximize the overall discriminative gain of the system to effectively assess its inference accuracy, is expressed as follows:
[0063]
[0064] Where β k B represents the local task allocation ratio of edge drone k. k This represents the uplink transmission bandwidth of the edge drone k.
[0065] Multiple constraints are established, including the non-negativity of the received signal strength of the large UAV; the classification task ratio range of the edge UAV; the feature stream being constrained by the uplink channel bandwidth; the uplink transmission bandwidth of the edge UAV not exceeding the total transmission bandwidth; the initial position of the large UAV; the flight speed constraint of the large UAV; and the peak and average transmission power constraints of the edge UAV. These are shown below:
[0066]
[0067] C5:q[0]=[x0,y0],
[0068]
[0069] Where C1 represents the nonnegativity of the received signal strength of the large UAV; C2 represents the classification task proportion range of the edge UAV k; and C3 represents the constraint of the uplink channel bandwidth on the feature stream of the edge UAV k. M represents the average size of the features extracted per frame. k S is the frequency at which the camera captures frames. k C1 represents the uplink transmission spectral efficiency of edge drone k; C2 represents the uplink transmission bandwidth of edge drone k not exceeding the total transmission bandwidth B. t C5 is the initial position of the large UAV, where q[0] is the initial position of the large UAV, and x0 and y0 are the initial horizontal and vertical coordinates of the large UAV, respectively; C6 is the flight speed constraint of the large UAV, where q[n-1] is the position of the large UAV in the (n-1)th time interval, τ is the time interval, and V m C7 and C8 represent the maximum horizontal speed of the large drone; C7 and C8 represent the peak transmission power P of the edge drone k. max,k and average transmission power Constrained by the following: L is the number of target categories, H1 is the flight altitude of the large UAV, H2 is the flight altitude of the edge UAV k, and l k [n] represents the position of the edge drone k within time interval n. Let z be the random variable corresponding to the edge drone k in time slot n. k The expectation of [n] squared.
[0070] Step S3: Deploy a large language model on cloud nodes. This large language model generates a global strategy through a planner, a memory bank, and a reflective evaluator. The global strategy includes a bandwidth allocation scheme, task splitting ratio, and macro-trajectory planning, such as... Figure 3As shown. Includes:
[0071] The planner generates a global strategy through a three-tier architecture;
[0072] LLM encodes multidimensional network states (such as bandwidth allocation, task proportion allocation, and drone location) into natural language descriptions. t ,
[0073] The tool invocation layer triggers API tools such as resource conflict checks and bandwidth queries. Based on the Toolformer's self-supervised mechanism, the LLM dynamically triggers tool API calls according to semantic descriptions. The API call decision is as follows:
[0074]
[0075] Where P M (<API>|D) t ) represents the probability of LLM predicting the next-token after the API call is triggered, τ s This indicates the sampling threshold.
[0076] The formula for generating API call instructions is:
[0077]
[0078] This represents the final output result, LLM. gen (·) represents the process of performing generation operations on a large language model, D t P represents the data / task-related input to the large language model. k API for policies, configurations, and alerts for edge drone k ζ (·) represents the ζ-th API, and params represents the API generated by the preceding large language model.
[0079] The policy generation layer outputs structured bandwidth allocation, trajectory planning, and task diversion instructions. The LLM generates hybrid policies based on templates and parses the natural language actions of the final decision into executable parameters.
[0080]
[0081] Where u t For the final executable instructions, g(·) is the instruction parser, and a t The natural language action generated for LLM, where Δβ is the adjustment amount to the task splitting ratio vector, β k ,k∈K represents the local task processing ratio of edge drone k, ΔB t B is the adjustment amount for the bandwidth allocation vector. k,k∈K is the bandwidth allocated to edge drone k, Δc k c is the adjustment amount for the equivalent received signal strength. k ,k∈K represents the transmitted signal strength of the edge drone k, Δq t This is the adjustment amount for the position of the large drone at the next moment.
[0082] The memory is used to store historical decision tuples. The i-th memory entry is represented as m. i =(s i ,a i ,c i ,e i meta i ), N m Let s be the total number of entries to remember, where s i a represents the state of the scenario at the moment of decision-making. i c represents the decision action generated by the LLM planner in this state. i A natural language description of the reasoning process, tools invoked, and key considerations of the LLM planner when generating actions. i Meta indicates outcome evaluation. i It represents metadata such as storage time, last access time, and number of times it was retrieved. Memory entries include state information, executed actions, reasoning processes, and result evaluations. Relevant historical records are retrieved through weighted cosine similarity, and the system is dynamically maintained using a dual-modal mechanism of conventional incremental updates and reflection-driven updates.
[0083] Relevant historical records are retrieved using weighted cosine similarity, with the similarity calculated using the weighted cosine similarity metric.
[0084]
[0085] Where D is the dimension of the state vector, d is the dimension index, and w d Let s be the weight of the d-th dimension feature. t,d For the current state s t The eigenvalues of the vector in the d-th dimension, s i,d Let s be the historical state vector i The eigenvalues in the d-th dimension, ||s t || w Let s be the current state vector. t The weighted norm, ||s i || w Let s be the historical state vector i The weighted norm.
[0086] A dual-modal mechanism of conventional incremental updates and reflection-driven updates is used for dynamic maintenance, with conventional incremental updates being executed automatically in each decision cycle:
[0087]
[0088] in For the updated memory at time step t+1, For the updated memory at time step t, s t For the scenario state at the moment of decision-making, a t For decision-making actions, c t For the decision context, e t For outcome evaluation, mata t This is metadata.
[0089] Reflection-driven updates occur after the evaluator identifies high-risk events:
[0090]
[0091] Where s risk The state of a high-risk event, a′ t For the corrected action, ∑p′ t The total penalty for violating the revised constraints, "Highrisk" t "This is a high-risk label."
[0092] The reflective evaluator calculates the decision quality score (DQS) based on discriminant gain and constraint penalty, and synthesizes the decision quality score function as follows:
[0093] DQS (κ) =G t ·exp(-γ·∑p t )
[0094] Where κ is the iteration number, and G t ∑p is the discrimination gain, γ is the penalty factor, and ∑p t It is the sum of punishments.
[0095] When the score is below a preset threshold, the decision is marked as a high-risk event.
[0096]
[0097] HighRisk t This is a high-risk indicator. This is an indicator function; the condition within the parentheses outputs 1 if true and 0 if false. θ risk A preset threshold is set for high risk.
[0098] Trigger iterative correction of the global strategy. The revised strategy is stored in the memory bank to achieve closed-loop optimization, where For memory bank, The optimal action to be adopted in the end. p Diagnosis is a comprehensive penalty assessment following the execution of the optimal action. final The final diagnostic report includes an analysis of the root causes of this high-risk event by the LLM faculty and a description of possible solutions.
[0099] Step S4, deploy deep reinforcement learning models at edge nodes, such as Figure 3 As shown. Includes:
[0100] The deep reinforcement learning model performs real-time optimization based on global policy and local observation data. The real-time optimization includes trajectory fine-tuning, dynamic bandwidth adjustment, and transmit power optimization.
[0101] The deep reinforcement learning module adopts the Actor-Critic framework, which integrates global policies and local observation data through a multi-head self-attention mechanism. The local observation data includes the real-time channel quality, location information, and resource usage status of edge nodes.
[0102] Step S5, establish a collaborative feedback mechanism, including:
[0103] The edge nodes will send the execution results back to the cloud nodes;
[0104] The cloud nodes update the global strategy based on the feedback results;
[0105] Edge nodes adjust real-time optimization parameters based on the updated global policy, and the results include channel status reports, resource utilization data, and constraint satisfaction status.
[0106] Step S6: The dynamic knowledge flow collaborative optimization actor-critic resource allocation algorithm achieves collaborative optimization of resource allocation, task ratio allocation, and channel parameter adjustment through a distributed perception and centralized decision-making framework. This algorithm employs a centralized training and distributed execution architecture, where each edge node acts as an independent agent, observing local resource status and associated terminal device information. It generates resource allocation strategies through an Actor-Critic network, and the Critic network evaluates the strategy value based on the global state to optimize the training process. Specifically, this includes:
[0107] The Actor-Critic framework is adopted, where the state-value function V π and action value function Q π The estimation is performed recursively. For policy π(a|s), in state s t The state value function and action value function at a given point are defined as follows:
[0108]
[0109] Where r(s) t ,at ) is the immediate reward function, a t ~π(·|s t ) is in state s t Next, perform the sampling action according to strategy π. π (s t ,a t ) is the action value function. Indicates the current state s t Perform action a t After that, the next state s t+1 The expected value that strategy π can bring.
[0110] In this invention, the V function and Q function are learned iteratively by minimizing the mean square Bellman error (MSBE), and the policy π is optimized by maximizing the Q value, where MSBE is defined as:
[0111]
[0112] in For empirical tuples (s) t ,a t ,r t ,s t+1 It follows the expected distribution B of the empirical replay buffer. r is the predicted value of the Q-network. t V is the immediate reward for the current step. g (s t+1 ) represents the state value of the value network.
[0113] The Kullback-Leibler (KL) divergence constraint is introduced, which formalizes the learning objective of the agent into a constrained optimization problem:
[0114]
[0115] Where s represents the state. For the state s t From the distribution Sampling, Action From strategy π S In s t The expected value of subsequent random variables is calculated for all cases of "action distribution sampling". For state s t Next action Expected cumulative rewards To minimize the negative Q-function value, π is an estimator of the KL divergence. S (s t ) represents the strategy π to be optimized. S In state st The action distribution under π T (s t (referencing strategy π) T In state s t The action distribution is given by σ, which is the constraint threshold for KL divergence.
[0116] Example 2:
[0117] An electronic device, comprising a memory and a processor;
[0118] The memory is used to store computer programs;
[0119] The processor is configured to implement the method described in Embodiment 1 when executing the computer program.
[0120] Example 3:
[0121] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0122] Example 4:
[0123] A computer program product includes a computer program that, when executed by a processor, implements the method described in Example 1.
[0124] In the above embodiments, the reference to "this embodiment" in the specification indicates that a specific feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple appearances of "this embodiment" do not necessarily refer to the same embodiment.
[0125] In the above embodiments, although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. The embodiments of the invention are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims.
[0126] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0127] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal performs the steps of the above method.
[0128] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0129] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0130] This invention can be used in a wide range of general-purpose or special-purpose computing system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.
[0131] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model, characterized in that: Includes the following steps: S1: Construct an edge-cloud drone collaborative inference system, consisting of one cloud node and K edge nodes; S2: Establish a joint optimization model that maximizes the discriminant gain; S3: Deploy a large language model on a cloud node. The large language model generates a global strategy through a planner, a memory bank, and a reflective evaluator. The global strategy includes a bandwidth allocation scheme, a task splitting ratio, and a macro trajectory planning. S4: Deploy a deep reinforcement learning model at each edge node. The deep reinforcement learning model performs real-time optimization based on the global policy and local observation data. The real-time optimization includes trajectory fine-tuning, dynamic bandwidth adjustment, and transmit power optimization. S5: Establish a collaborative feedback mechanism. Edge nodes will feed back the execution results to cloud nodes. Cloud nodes will update the global strategy based on the feedback results. Edge nodes will adjust the real-time optimization parameters based on the updated global strategy. S6: Employs an actor-critic resource allocation algorithm based on dynamic knowledge flow collaborative optimization. Through a distributed perception and centralized decision-making framework, it achieves collaborative optimization of resource allocation, task ratio allocation, and channel parameter adjustment.
2. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The cloud node is a large drone, which provides bandwidth allocation, task scheduling, global decision-making for large drone trajectory planning, and generates macro commands; the edge node is an edge drone, which is used to collect environmental data, perform feature extraction, and proportionally distribute tasks to large drones.
3. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The joint optimization model takes maximizing the overall discriminant gain of the system as its optimization objective, and the objective function is expressed as: Where k represents the index of the edge drone, n represents the discretized time slot index, m represents the index of the receiving feature dimension of the large drone, and c k,n Let q[n] represent the equivalent received signal strength of the edge drone k in the nth time interval, and let q[n] represent the trajectory of the large drone in the nth time interval. k B represents the local task allocation ratio of edge drone k. k Let μ[m] represent the uplink transmission bandwidth of edge drone k, and μ[m] represent the mean value of the m-th feature dimension. This represents the perceptual distortion variance of the m-th feature dimension. This represents the variance of the sensing distortion of the edge drone k. This represents the additive Gaussian noise at the large drone end; N is the total number of time intervals, and K is the total number of edge drones; The constraints include: Where C1 represents the nonnegativity of the received signal strength of the large UAV; C2 represents the classification task proportion range of the edge UAV k; and C3 represents the constraint of the uplink channel bandwidth on the feature stream of the edge UAV k. M is the average size of the features extracted for each frame. k S is the frequency at which the camera captures frames. k C1 represents the uplink transmission spectral efficiency of edge drone k; C2 represents the uplink transmission bandwidth of edge drone k not exceeding the total transmission bandwidth B. t C5 is the initial position of the large UAV, where q[0] is the initial position of the large UAV, and x0 and y0 are the initial horizontal and vertical coordinates of the large UAV, respectively; C6 is the flight speed constraint of the large UAV, where q[n-1] is the position of the large UAV in the (n-1)th time interval, τ is the time interval, and V m C7 and C8 represent the maximum horizontal speed of the large drone; C7 and C8 represent the peak transmission power P of the edge drone k. max,k and average transmission power Constrained by the following: L is the number of target categories, H1 is the flight altitude of the large UAV, H2 is the flight altitude of the edge UAV k, and l k [n] represents the position of the edge drone k within time interval n. Let z be the random variable corresponding to the edge drone k in time slot n. k The expectation of [n] squared.
4. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The planner generates a global policy through a three-layer architecture: the semantic parsing layer encodes the system state into a natural language description D. t The tool call layer triggers resource conflict checks and bandwidth query API tools; the strategy generation layer outputs structured bandwidth allocation, trajectory planning, and task diversion instructions.
5. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The memory bank is used to store historical decision tuples. The i-th memory entry is represented as m. i =(s i ,a i ,c i ,e i meta i ), N m Let s be the total number of entries to remember, where s i a represents the state of the scenario at the moment of decision-making. i c represents the decision action generated by the LLM planner in this state. i A natural language description of the reasoning process, tools invoked, and key considerations of the LLM planner when generating actions. i Meta indicates outcome evaluation. i Meta-information includes storage time, last access time, and number of times it was retrieved; memory entries include state information, executed actions, reasoning process, and result evaluation. Relevant historical records are retrieved through weighted cosine similarity and dynamically maintained using a dual-modal mechanism of conventional incremental updates and reflection-driven updates.
6. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The reflective evaluator calculates the decision quality score (DQS) based on discriminant gain and constraint penalty, denoted as DQS. (κ) =G t ·exp(-γ·∑p t ), where κ is the iteration round, G t ∑p is the discrimination gain, γ is the penalty factor, and ∑p t It is the sum of punishments; When the score is lower than the preset threshold, the global strategy is iteratively corrected, and the corrected strategy is stored in the memory to achieve closed-loop optimization.
7. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The deep reinforcement learning model adopts the Actor-Critic framework, which integrates global policy and local observation data through a multi-head self-attention mechanism. The local observation data includes the real-time channel quality, location information and resource usage status of edge nodes.
8. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: In step S5, the execution results of the collaborative feedback mechanism include channel state reports, resource utilization data, and constraint satisfaction status. The collaborative feedback mechanism performs an update once every preset time interval, with the time interval ranging from 0.1 to 1 second.
9. The cloud-edge multi-UAV collaborative resource optimization method assisted by a large language model according to claim 1, characterized in that: The dynamic knowledge flow collaborative optimization actor-critic resource allocation algorithm adopts the Actor-Critic framework, with a state value function V. π and action value function Q π The estimation is performed recursively; for policy π(a|s), in state s t The state value function and action value function at a given point are defined as follows: Where r(s) t ,a t ) represents the instant reward function, a t ~π(·|s t ) indicates that in state s t Next, sample the action according to strategy π; Q π (s t ,a t ) is the action value function. Indicates the current state s t Perform action a t After that, the next state s t+1 Based on the expected value that strategy π can bring; The V and Q functions are learned iteratively by minimizing the mean square Bellman error (MSBE), and the policy π is optimized by maximizing the Q value, where MSBE is defined as: in For empirical tuples (s) t ,a t ,r t ,s t+1 The expected distribution of the empirical replay buffer is B. r is the predicted value of the Q-network. t V is the immediate reward for the current step. g (s t+1 The state value of the value network; By introducing the Kullback-Leibler divergence constraint, the learning objective of the agent is formalized into a constrained optimization problem: Where s represents the state. For "state s" t From the distribution Sampling, Action From strategy π S In s t The expected value of subsequent random variables is calculated for all cases of "action distribution sampling". For state s t Next action Expected cumulative rewards To minimize the negative Q-function value, π is an estimator of the KL divergence. S (s t ) represents the strategy π to be optimized. S In state s t The action distribution under π T (s t (referencing strategy π) T In state s t The action distribution is given by σ, which is the constraint threshold for KL divergence.
Citation Information
Cited By
Energy-saving optimization method and device based on large language model reasoning, equipment and medium
CN122331740A