A cloud-edge-end model inference joint optimization method under air-ground cooperation

By employing a joint optimization method for cloud-edge-device model inference under air-ground collaboration, the selection of UAVs, language models, and trajectory planning are optimized, thus solving the trade-off between inference quality and latency in UAV inference services and achieving system latency reduction and resource efficiency improvement.

CN121706997BActive Publication Date: 2026-04-14NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing research on drone computation offloading or inference services has not fully considered the trade-off between inference quality and inference latency during large-scale model inference. This makes it difficult to achieve optimal global performance in scenarios with highly dynamic changes in user requests and diverse task types, and can easily lead to local congestion or wasted computing power.

Method used

A joint optimization method for cloud-edge-device model inference under air-ground collaboration is proposed. By jointly optimizing the decision variables of UAV selection, language model selection, inference task offloading, and UAV trajectory planning, a multi-agent reinforcement learning framework is constructed. Combined with the successive convex approximation method, the total execution latency and inference accuracy of the system are optimized.

Benefits of technology

While ensuring inference accuracy, it significantly reduces task execution latency, thereby improving the overall service quality and resource utilization efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706997B_ABST
    Figure CN121706997B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of cloud edge-end collaborative reasoning, and discloses a cloud edge-end model reasoning joint optimization method under air-ground collaboration. A network architecture modeling module, a reasoning performance modeling module, a delay modeling module and a joint optimization module are designed. By constructing an air-ground collaborative reasoning system, a thinking chain prompting mechanism is introduced to model the reasoning accuracy, and the unmanned aerial vehicle selection, language model selection, reasoning task offloading decision and unmanned aerial vehicle trajectory are jointly optimized to minimize the total system cost. The continuous convex approximation method is used to optimize the unmanned aerial vehicle trajectory, and the multi-agent reinforcement learning method is combined for distributed decision and centralized training, so that the low-delay and high-precision collaborative reasoning service is realized. The method disclosed by the application can effectively coordinate the communication, calculation and reasoning resources in the cloud edge-end collaborative scene with dynamic user request and various task types, significantly improve the overall service quality and resource utilization efficiency of the system, and is superior to other existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud-edge collaborative reasoning technology, and in particular to a joint optimization method for cloud-edge-end model reasoning under air-ground collaboration. Background Technology

[0002] With the widespread application of Large Language Models (LLMs) in fields such as intelligent question answering, decision support, and automated reasoning, inference services place extremely high demands on computing resources, communication bandwidth, and response latency. However, limited by the large model size, high computational complexity of inference, and limited computing power of terminal devices, most existing inference services rely on centralized cloud computing centers, leading to communication link congestion and excessive end-to-end latency, making it difficult to meet the application requirements of low latency and high reliability. To alleviate these problems, edge computing reduces inference latency to some extent by offloading some computing power to edge nodes closer to the user. However, ground edge nodes still have significant limitations in terms of coverage, deployment flexibility, and adaptability to emergency scenarios. In recent years, the rapid development of UAV technology has provided a new technical means for cloud-edge-device collaborative inference. UAVs have advantages such as flexible deployment, high mobility, and good line-of-sight communication conditions, and can serve as aerial edge computing nodes to provide on-demand inference services to ground users and as an important relay connecting terminals and the cloud. However, existing research on UAV-based computation offloading or inference services largely focuses on communication coverage or simple task offloading strategies, failing to adequately consider the trade-off between inference quality and latency during large-scale model inference. In particular, the performance of large-scale model inference is highly dependent on contextual information and intermediate inference processes, such as the number of inference terms introduced by Chain-of-Thought (CoT) hints. Under the constraint of limited UAV computing resources, how to rationally allocate models, manage inference context, and coordinate cloud computing power to complete inference tasks is a key challenge in cloud-edge-device collaborative inference scenarios. Furthermore, in scenarios with highly dynamic user requests, diverse task types, and sensitivity to inference accuracy, relying solely on fixed rules or local optimization strategies is insufficient to achieve optimal global performance, easily leading to problems such as local congestion, wasted computing power, or degraded inference quality. Therefore, a comprehensive optimization method that jointly considers communication, computation, inference quality, and UAV scheduling constraints in cloud-edge-device collaborative inference scenarios is urgently needed to effectively improve the overall system performance. Summary of the Invention

[0003] To address the aforementioned issues, this invention proposes a joint optimization method for cloud-edge-device model inference under air-ground collaboration. By jointly optimizing the decision variables for UAV selection, language model selection, inference task unloading, and UAV trajectory planning, the method effectively reduces the total execution latency of the task while ensuring inference quality.

[0004] The technical solution of the present invention is as follows: a cloud-edge-device model joint optimization method under air-ground collaboration, comprising a network architecture modeling module, an inference performance modeling module, a latency modeling module, and a joint optimization decision module;

[0005] The network architecture modeling module is used to construct an air-ground collaborative inference system consisting of multiple ground users, multiple drones and the cloud. Multiple drones are introduced to cover the target area, providing air edge inference services for ground users and connecting to the cloud through a backhaul link. The running time of the air-ground collaborative inference system is discretized into multiple time slots of equal length, and the position and speed of the drones and the communication relationship between ground users and drones are characterized in each time slot.

[0006] The inference performance modeling module is used to maintain the accumulated and newly added context words in continuous time slots for the air edge inference task requested by ground users, and to calculate the impact of the number of context words and the length of the thought chain formed by the context words on the inference accuracy.

[0007] The latency modeling module calculates the communication transmission latency and inference calculation latency of the air edge inference task in two modes: local execution on the UAV and execution in the cloud. The total execution latency of the task is obtained by summing them up.

[0008] The joint optimization decision module, based on the constructed air-ground collaborative reasoning system, constructs a joint optimization problem that includes decision variables for UAV selection, language model selection, reasoning task unloading, and UAV trajectory planning, taking into account the total task execution delay, reasoning accuracy, and constraints. The decision variables are uniformly incorporated into the same optimization framework, and the joint optimization problem is solved by combining the successive convex approximation method with the multi-agent reinforcement learning method.

[0009] In the air-ground collaborative reasoning system, the UAV is limited to flying at a fixed altitude and its horizontal speed does not exceed the maximum allowable speed; the uplink transmission rate between the ground user and the UAV is calculated based on the wireless channel model, and the backhaul rate between the UAV and the cloud is set.

[0010] In the air-ground collaborative reasoning system, the cloud and drones deploy large-scale language models that support the thought chain prompting mechanism for reasoning. The reasoning process takes into account the length of the thought chain and the complexity of the reasoning.

[0011] The aerial edge inference task requested by ground users is represented as a task tuple containing the original data size, the number of effective lexical units, and the minimum inference accuracy requirement. Decision variables for ground user and UAV selection, language model selection, and inference task offloading are defined to constrain each ground user to select at most one UAV and at most one large-scale language model in a single time slot, and to ensure that the offloading consistency relationship holds.

[0012] The joint optimization problem in the joint optimization decision-making module is as follows:

[0013]

[0014] In the joint optimization problem, the UAV selects decision variables. , Indicates ground users In the time slot Should we choose a drone? The decision variables, among which Indicates ground users In the time slot Assigning aerial edge inference tasks to drones Otherwise, it is 0; the inference task unloads decision variables. , Indicates ground users The In time slots, the aerial edge reasoning task via drone Uninstalled to the cloud Each inference instance is executed. Indicates no uninstallation; language model selects decision variables. , Indicates ground users The Aerial edge reasoning tasks in drones Model inference is performed locally, and must satisfy the following conditions: The relationship between them; Represents the drone's position variable. and They represent drones Initial position and drone The final position is determined by trajectory constraints, which ensure smooth movement throughout the entire time slot. Indicates drone In the time slot The horizontal speed variable is subject to the maximum permissible speed. constraint; Indicates in time slot Inside, drones For ground users The airborne edge reasoning task was performed in the first The computational load of inference on a large-scale language model is measured by the number of lexical units, including historically accumulated context lexical units and newly introduced lexical units in the current time slot. This indicates the maximum number of tokens that the GPU on the drone can process in a single time slot;

[0015] The objective function of the joint optimization problem consists of two parts: This represents the weighted sum of the total execution delays of the tasks, where For ground users No. In time slots, the aerial edge reasoning task Total execution latency, This is the time delay weighting factor; This represents the weighted sum of inference performance error cost under task allocation, where This indicates the performance degradation of airborne edge inference tasks when performed on drones or in the cloud. Precision weighting factor;

[0016] Constraint C1 limits the horizontal speed of the UAV at arrive Between; Constraint C2 restricts each ground user to selecting at most one drone in a single time slot, and Constraint C3 restricts each ground user's airborne edge inference task to executing at most one large-scale language model within a single time slot. ,and Constraint C4 guarantees that the drone is selected only when the ground user chooses it. Its aerial edge inference tasks are either executed locally on the drone or offloaded to the cloud, i.e. Constraint C5 ensures that after the aerial edge inference task is associated with the UAV, inference is performed on the local model or offloaded to the cloud for inference. Constraint C6 limits the maximum number of tasks that the UAV can process in a single time slot. Constraint C7 specifies that the initial position of the UAV is The final location of the drone was .

[0017] The inference performance modeling module specifically includes: dynamically updating the accumulated context words in continuous time slots, combining them with the newly introduced context words in the current time slot to form the inference context, forming a complete thought chain input, and calculating the impact of the thought chain on the inference accuracy; using the cloud inference accuracy as a reference benchmark, defining the difference between the UAV local inference accuracy and the cloud inference accuracy as the inference performance error cost.

[0018] The solution process of the joint optimization decision module includes a UAV trajectory optimization sub-process and a UAV resource allocation optimization sub-process;

[0019] A successive convex approximation method is used to handle the UAV trajectory optimization subprocess.

[0020] A multi-agent reinforcement learning method is used to handle the sub-process of drone resource allocation optimization.

[0021] The specific sub-process for optimizing the drone trajectory is as follows:

[0022] Given the decision variables for selecting drones Language model selection of decision variables Unloading decision variables from reasoning tasks Under these conditions, the UAV trajectory optimization sub-process involves adjusting the UAV position variables... The optimization problem is modeled as a non-convex optimization problem, and a first-order Taylor expansion is performed on the trajectory constraint function to construct a locally convex approximation subproblem.

[0023]

[0024] Solve the local convex approximation subproblem to obtain the current optimal UAV trajectory, and determine whether the difference between two adjacent trajectory update results is less than a preset convergence threshold. When the convergence condition is met, the current trajectory is output as the final optimization result. At the same time, during the entire iteration process, the UAV flight speed is constrained to meet constraint condition C1, and the UAV start and end position constraint condition C7.

[0025] The specific sub-process for optimizing UAV resource allocation is as follows:

[0026] Under the condition of fixed drone trajectory, the decision variables for drone selection, language model selection, and inference task unloading are jointly optimized and modeled as a partially observable Markov decision process under a multi-agent reinforcement learning framework, thereby constructing a sub-problem of drone resource allocation optimization:

[0027]

[0028] Specifically, each drone is modeled as an independent intelligent agent, whose actions are as follows: Within each discrete time slot, each agent makes decisions based on its local observation sequence, which includes the user task status, UAV resource status, and communication link status.

[0029] The multi-agent reinforcement learning method is a multi-agent transform reinforcement learning method used to jointly represent the decision variables for UAV selection, the decision variables for language model selection, and the decision variables for inference task unloading.

[0030] The multi-agent reinforcement learning method uses a Transformer network with an encoder-decoder structure for policy modeling.

[0031] First, the encoder embeds and extracts features from the local observation sequences acquired in continuous time slots to capture the temporal correlations and cross-agent dependencies in the multi-agent decision-making process; based on the encoded representation, the decoder outputs the UAV selection decision variables, the language model selection decision variables, and the inference task unloading decision variables in an autoregressive manner, thereby forming a joint action.

[0032] During the policy execution phase, each agent interacts with the environment based on the joint actions generated by the current policy, receives immediate rewards, and updates the state of the air-ground collaborative inference system. The reward function consists of a weighted sum of the total task execution latency and the inference performance error cost. in It is a penalty factor; during the training process, the state, action and reward data of the air-ground collaborative reasoning system generated by interaction are stored in the experience replay buffer, and small batch training samples are formed through random sampling, which in turn drive the parameter learning of the encoder and decoder networks and perform policy iterative optimization.

[0033] The beneficial effects of this invention are as follows: This invention proposes a joint optimization method for cloud-edge-device model inference under air-ground collaboration. By jointly optimizing the decision variables for UAV selection, language model selection, inference task offloading, and UAV trajectory planning, the method significantly reduces task execution latency while ensuring inference accuracy, thereby improving the overall service quality and resource utilization efficiency of the system. Attached Figure Description

[0034] Figure 1 This is a diagram illustrating the overall framework of the cloud-edge-device model inference joint optimization method under air-ground collaboration.

[0035] Figure 2 This is a schematic diagram of the network architecture modeling module.

[0036] Figure 3 This is a schematic diagram of the inference performance modeling module.

[0037] Figure 4 This is a schematic diagram comparing the training performance of MAT-RA (the method of this invention) and the baseline algorithm.

[0038] Figure 5 This is a schematic diagram comparing the inference accuracy of MAT-RA (the method of this invention) and the baseline algorithm.

[0039] Figure 6 This is a schematic diagram comparing the latency performance of MAT-RA (the method of this invention) and the baseline algorithm.

[0040] Figure 7 It is a preference and Schematic diagram of MAT-RA (the method of this invention) training performance under different values.

[0041] Figure 8 It is a preference and A schematic diagram showing the inference accuracy of MAT-RA (the method of this invention) under different numerical values.

[0042] Figure 9 It is a preference and Schematic diagram of the delay performance of MAT-RA (the method of this invention) under different values. Detailed Implementation

[0043] The core idea of ​​this invention is to treat drones as aerial edge inference nodes with limited computing power, introduce a performance modeling method driven by inference context, comprehensively consider the latency and accuracy differences of the three-layer collaborative inference path of ground user-drone-cloud, and construct a unified joint optimization framework under the time-slotted system model.

[0044] A joint optimization method for cloud-edge-device model inference under air-ground collaboration includes a network architecture modeling module, an inference performance modeling module, a latency modeling module, and a joint optimization decision module;

[0045] The network architecture modeling module is used to construct an air-ground collaborative inference system, which is a cloud-drone-ground user collaborative inference architecture. It introduces multiple drones to cover the target area, provides air edge inference services for ground users, and connects to the cloud through a backhaul link. The running time of the air-ground collaborative inference system is discretized into multiple equal-length time slots, and the position and speed of the drone and the communication relationship between the ground user and the drone are characterized in each time slot.

[0046] The reasoning performance modeling module, based on the prompting mechanism of the thought chain, constructs a relationship model between reasoning accuracy and the number of context words, which is used to evaluate the reasoning performance of different models on different tasks.

[0047] The latency modeling module is used to calculate the communication latency and computation latency of the air edge inference task in two modes: UAV local inference and cloud inference. It also introduces inference performance error cost to quantify the accuracy decline caused by the large-scale language model and context limitations.

[0048] The joint optimization decision module, under the premise of satisfying the constraints of UAV flight speed, computing power and consistency, jointly optimizes the decision variables for UAV selection, language model selection, inference task offloading and UAV trajectory planning, so as to minimize the overall cost.

[0049] The network architecture modeling module is specifically as follows:

[0050] Consider an air-to-ground collaborative inference system consisting of multiple ground users, multiple drones, and a cloud. Ground users are located at fixed altitudes, and drones fly at fixed altitudes with their horizontal speed limited by a maximum speed. The uplink transmission rate between ground users and drones is determined by channel gain, bandwidth, noise power, and user transmit power. Drones connect to the cloud, and the drone-to-cloud transmission rate is a fixed value. Each ground user can request multiple types of air-to-edge inference services, and each air-to-edge inference task is characterized by data size, the number of effective context terms, and a minimum inference accuracy requirement.

[0051] The inference performance modeling module is specifically as follows:

[0052] Based on the Chain-of-Thought (CoT) hint mechanism, inference performance is modeled as a function of the cumulative number of context terms and the number of newly added context terms. A sigmoid function is used to simulate the accuracy improvement brought by CoT, and zero-sample accuracy and CoT sensitivity coefficient are introduced to construct inference accuracy models for the task on different models.

[0053] The delay modeling module is specifically as follows:

[0054] The total mission execution latency includes communication transmission latency and inference computation latency. Communication transmission latency includes the uplink transmission time from the ground user to the UAV, and the additional transmission time from the UAV to the cloud during mission unloading. Inference computation latency is modeled based on the mission execution location (local, UAV, or cloud) and the computational capabilities of the selected large-scale language model.

[0055] The joint optimization decision-making module is as follows:

[0056] The decision variables for drone selection, language model selection, inference task unloading, and drone trajectory planning are jointly modeled as a joint optimization problem with the objective of minimizing the total system cost. The total system cost includes the weighted sum of total task execution latency and inference accuracy error. The joint optimization problem is as follows:

[0057]

[0058] The objective function minimizes the total system cost, which is determined by the total execution time latency of the tasks. And a weighted penalty term reflecting the decrease in reasoning performance under task allocation. The constraints are as follows: Constraint C1 limits the maximum horizontal speed of the UAV. Constraints C2 and C3 respectively ensure that each ground user can select at most one UAV and one large-scale language model for their aerial edge inference task. Constraint C4 ensures that any UAV... Offloading aerial edge inference tasks to the cloud first involves drones Correlation. Constraint C5 ensures that if an aerial edge inference task is associated with a drone... If associated, it must either be executed locally on the drone using a large-scale language model or offloaded to the cloud. Constraint C6 limits the maximum processing workload for each drone. Finally, constraint C7 specifies the initial and final positions of the drones.

[0059] The successive convex approximation (SCA) method is used to optimize the UAV trajectory. First, decision variables are selected based on the initial UAV parameters. Language model selection of decision variables Unloading decision variables from reasoning tasks The non-convex primal problem of UAV trajectory optimization is expanded using a first-order Taylor series at the initial trajectory point to obtain a locally convex approximation subproblem. Then, in each iteration, other variables are fixed, and this convex approximation subproblem is solved to obtain the current optimal trajectory. Next, it is determined whether the difference between the obtained trajectory and the previous trajectory is less than a preset convergence threshold. If the convergence condition is not met, the current optimal trajectory is used as a new reference point, and the convex approximation is performed again and solved. If the convergence condition is met, the current trajectory is output as the final optimized solution. Throughout the entire iteration process, the UAV trajectory is ensured to meet constraints such as maximum speed, start and end point positions, and obstacle avoidance.

[0060] A multi-agent reinforcement learning approach is employed to solve for the optimal decision variables based on the state of the air-ground collaborative reasoning system in a given time slot, aiming to minimize the current system cost. After obtaining the UAV trajectory, the system cost is solved using the same optimization problem as an evaluation metric. The rewards for the agents are defined. ,in This represents the maximum number of tokens that the GPU on the drone can process in a single time slot; the drone selection decision variable for the next time slot. Language model selection of decision variables Unloading decision variables from reasoning tasks Each vector is encoded as a one-hot vector and used as the drone's action. The drone continuously provides services to the user during flight and ensures system-level constraints at all times, including drone speed limits, user-drone association constraints, task-model matching constraints, GPU load limits, and drone trajectory start and end position constraints.

[0061] The specific embodiments of the present invention will be described in detail below.

[0062] Considering that current research has neglected in-depth exploration of UAV-assisted large model task execution in dynamic user distribution and cloud-edge-device collaborative reasoning scenarios, and addressing the difficulties in deploying UAVs and formulating inference offloading decisions in such scenarios, this invention designs a cloud-edge-device large model inference method in air-ground collaborative scenarios based on inference context modeling and joint optimization.

[0063] This invention is primarily implemented based on four modules. First, a network architecture modeling module is used to perceive and model the location of the terminal ground user, the task arrival status, the UAV's operational status, and the backhaul link information to obtain the overall operational status of the air-ground collaborative inference system in the current time slot. Second, an inference performance modeling module is used to parse the inference task requests and their task parameters generated in each time slot, maintain the inference context lexical units in continuous time slots, and characterize the impact of the thought chain formed by the context on the inference accuracy. Third, a latency modeling module is used to model and evaluate the communication transmission latency and inference computation latency generated by the user's inference tasks, providing a quantitative basis for joint optimization decision-making. Finally, a joint optimization decision-making module is used to formulate an execution strategy for each inference task based on a comprehensive consideration of the air-ground collaborative inference system status, computing resources, and inference context, jointly determining the UAV selection decision variables, language model selection decision variables, inference task unloading decision variables, and UAV trajectory planning decision variables.

[0064] During the operation of the air-ground collaborative inference system, at the beginning of each time slot, the ground user terminal equipment generates an airborne edge inference task and uploads it to the UAVs within its coverage area. The network architecture modeling module collects information such as UAV location, user distribution, and communication link status, and encapsulates it into a system state vector. Subsequently, the joint optimization decision module combines the UAV's remaining computing resources and inference context information to determine the task execution location and language model selection, and calls the latency modeling module to calculate the corresponding communication transmission latency and inference computation latency. At the same time, it combines the inference performance modeling module to evaluate the inference performance error cost under different decisions.

[0065] The inference performance modeling module dynamically maintains the accumulated context words in consecutive time slots and combines them with newly added context words in the current time slot to form a complete inference context, thus creating a thought chain input to improve the accuracy of local UAV inference. For airborne edge inference tasks offloaded to the cloud, the inference performance modeling module decides whether to synchronize context information based on the air-ground collaborative inference system strategy to achieve a trade-off between communication overhead and inference performance. After each time slot, the air-ground collaborative inference system calculates the comprehensive system cost of all tasks based on the evaluation results of the latency modeling module and the inference performance modeling module, and updates the task allocation decisions and UAV operation strategies in subsequent time slots with the goal of minimizing this cost through the joint optimization decision module.

[0066] This invention considers a cloud-edge-device model inference service architecture under air-ground collaboration, which enables ground users to perform inference tasks in a remote environment, wherein the ground users are... Index. (Page number missing) The location of each ground user is fixed as follows A group of drones It can simultaneously cover the area and provide aerial edge inference services to ground users. To simplify the solution space, the mission duration... Discretized into There are several equal-length time slots, with a length of [missing information]. ,Right now Each time slot is composed of Index. Let the three-dimensional matrix... Indicates the first A drone in a time slot The location. Due to airspace control and energy efficiency constraints, it is assumed that the UAV operates at a fixed altitude, and its horizontal flight speed is limited by the maximum permissible speed.

[0067] (1)

[0068] in Indicates drone In the time slot horizontal flight speed, This is the maximum speed limit. The ground user and the first The channel power gain between drones can be expressed as:

[0069] (2)

[0070] in It is a time slot The decline factor is related to the influence of meteorological conditions. The channel power represents the reference distance. Therefore, the... The ground user and the first The uplink transmission rate between drones can be expressed as:

[0071] (3)

[0072] in , and These are bandwidth and Gaussian white noise, respectively. Indicates the first The transmit power of each ground user. Furthermore, to gain additional computing power, the drone connects to a cloud computing center via the LEO constellation network. For simplified optimization, this invention assumes a fixed drone-to-cloud transmission rate. .

[0073] This invention assumes that each ground user can request Types of aerial edge reasoning services, indexed as Define a task tuple. To characterize ground users In the time slot Requested aerial edge reasoning task ,in This refers to the size of the original data to be uploaded. Indicates the number of valid information tokens. This represents the minimum required inference accuracy. Furthermore, each drone carries and maintains this accuracy within its GPU. A large-scale language model, indexed as This invention defines binary variables. For ground users In the time slot The choice of drones, among which Indicates ground users Choose a drone And send a task request to it, otherwise According to the proposed design, each ground user can only select one drone within a time slot:

[0074] (4)

[0075] In real-world scenarios, a single large-scale language model can simultaneously serve multiple ground users' in-flight edge inference tasks. This requires a comprehensive understanding of the relationship between large-scale language models and in-flight edge inference tasks. To model the relationship between them, this invention defines a binary variable. ,in Indicates task In the time slot For drones Executed and by large-scale language models Perform reasoning, otherwise For the decision variable of the reasoning task, this invention defines a binary variable. ,in Represents the task of airborne edge reasoning In the time slot via drone Uninstall and run it in the cloud, otherwise... According to the proposed design, each user's task must satisfy the following consistency constraint:

[0076] (5)

[0077] Constraint (5) ensures that in the time slot For drones Execution of aerial edge reasoning mission The reasoning can either be performed locally by a large-scale language model on the drone, or it can be offloaded to the cloud for execution.

[0078] Quality of service is defined by the performance improvements in inference brought about by thought chain (CoT) hints. Two types of contextual lexics contribute to the inference process: accumulated contextual lexics, which are contextual lexics generated from previous thought chain inference steps, reflecting the historical inference context available in large-scale language models; and newly added contextual lexics, which are effective contextual lexics newly provided by the ground user in the current step. Let... Represents the task of airborne edge reasoning In the time slot For drones Using large-scale language models The number of context terms used during local inference.

[0079] (6)

[0080] in Represents the task of airborne edge reasoning In the time slot For drones Using large-scale language models The cumulative number of context lexical units during local inference.

[0081] (7)

[0082] In equation (7), ,in and They represent time slots respectively The number of newly generated and temporarily discarded context terms.

[0083] Zero-shot accuracy of a large-scale language model is used as a baseline, and task-specific contextual lexical cues are incorporated. Specifically, large-scale language models... In drones task of reasoning at the edge of the sky The reasoning performance is denoted as ;

[0084] (8)

[0085] In equation (8), Defined as a sigmoid function, it is used to simulate the diminishing marginal returns of accuracy gains from contextual terms, while ensuring that accuracy remains within the range of [0,1]. Furthermore, Representing large-scale language models In-flight edge reasoning task The zero-sample accuracy, while Represents the task of airborne edge reasoning The contextual lexical sensitivity coefficient.

[0086] Because the cloud possesses large-scale language models and abundant computing resources, the accuracy of the cloud is denoted as... Therefore, the task of airborne edge reasoning will be... Offloading to the cloud can achieve higher accuracy, i.e. Therefore, if the airborne edge reasoning task... In drones The accuracy error cost of execution is expressed as follows:

[0087] (9)

[0088] The total GPU inference load for each drone is limited by its processing power.

[0089] (10)

[0090] in, It is the maximum number of context terms that a GPU can process in a single time slot.

[0091] If the task is to perform edge reasoning in the air In drones The communication transmission delay is given by the following formula:

[0092] (11)

[0093] when At that time, the aerial edge reasoning mission In drones For local inference, transmission latency only includes the transmission time from the user to the drone. Conversely, when At that time, the aerial edge reasoning mission The drone was offloaded to the cloud, which introduced additional transmission latency from the drone to the cloud.

[0094] When the edge of the air reasoning task In drones When the local inference is executed, its inference computation delay is given by the following formula:

[0095] (12)

[0096] in Representing large-scale language models The computational workload per context word (unit: FLOPs / word). Indicates drone The computing power (unit: FLOPs / s). If the airborne edge inference task... From drones When offloaded to the cloud, its inference computation latency is given by the following formula:

[0097] (13)

[0098] in Indicates cloud-based edge inference task The number of historical experience context lexicons, and These represent the computational cost per context word in the cloud (unit: FLOPs / word) and the cloud computing capability (unit: FLOPs / word), respectively. Therefore, in drones... Aerial edge reasoning mission performed at the location The total inference computation delay is:

[0099] (14)

[0100] So, the aerial edge reasoning task The total execution latency of the task is expressed as:

[0101] (15)

[0102] To reduce the task execution latency and inference performance error costs of user inference tasks, this invention constructs a joint optimization problem that simultaneously considers the UAV selection decision variables. Language model selection of decision variables Inference task unload decision variables and decision variables for drone trajectory planning and their corresponding flight speed Accordingly, this joint optimization problem can be modeled as:

[0103]

[0104] The objective function minimizes the total system cost, which is determined by the total task execution delay. And a weighted penalty term reflecting the decrease in reasoning performance under task allocation. composition. This is a weighting factor: larger values ​​emphasize inference quality, while smaller values ​​focus on reducing latency; its value can be adjusted according to specific application requirements. Constraint C1 limits the maximum horizontal speed of the drone. Constraints C2 and C3 ensure that each ground user can select at most one drone and one large-scale language model for its aerial edge inference task, respectively. Constraint C4 ensures that any... Offloading aerial edge inference tasks to the cloud first involves drones Correlation. Constraint C5 ensures that if an aerial edge inference task is associated with a drone... If associated, it must either be executed locally on the drone using a large-scale language model or offloaded to the cloud. Constraint C6 limits the maximum processing workload for each drone. Finally, constraint C7 specifies the initial and final positions of the drones.

[0105] Given the decision variables for selecting drones Language model selection of decision variables Unloading decision variables from reasoning tasks Under these conditions, the UAV trajectory optimization sub-process involves adjusting the UAV position variables... The optimization problem is modeled as a non-convex optimization problem, and a first-order Taylor expansion is performed on the trajectory constraint function to construct a locally convex approximation subproblem.

[0106]

[0107] Given a fixed drone trajectory, the problem of jointly optimizing the decision variables for drone selection, language model selection, and inference task unloading is modeled as a partially observable Markov decision process in multi-agent reinforcement learning, denoted as: ,in It is a collection of intelligent agents. This is the current state of the air-ground collaborative reasoning system. It is an intelligent agent The space of motion It is an intelligent agent The observed values, It is an intelligent agent The reward This represents the discount factor. The definition of this concept in this invention will be explained in detail below:

[0108] Collection of intelligent agents :

[0109] In the air-ground collaborative reasoning system, each UAV corresponds to an intelligent agent, which is responsible for learning whether to select processing, model allocation and unloading strategies for the received tasks, and taking the optimal action according to the algorithm strategy.

[0110] Observation :

[0111]

[0112] Each agent constructs local observations based on locally available information, including task characteristics, the resource status of the UAV, and relevant network link parameters.

[0113] Action space :

[0114]

[0115] In the time slot When the drone receives the mission At that time, its actions are a three-way decision combination, including ground users. In the time slot Should we choose a drone? Decision variables Ground users The In time slots, the aerial edge reasoning task via drone Uninstalled to the cloud Execute one inference instance and representing ground users The Is airborne edge reasoning a task applicable to drones? Local model inference .

[0116] The reward function consists of the total task execution latency and the inference performance error cost. At the same time, it applies a soft penalty to behaviors that violate the constraints of the drone's GPU processing power, thereby guiding the agent to achieve efficient inference services while meeting resource constraints.

[0117]

[0118] Discount factor This is used to balance immediate rewards with long-term benefits, ensuring the stability and convergence of the learning process;

[0119] Finally, the environment is set up and the reinforcement learning algorithm is trained. The training process includes initializing system parameters and the environment, interacting with the environment to obtain training sample data, and randomly extracting data for policy updates. Finally, simulations are performed to evaluate the algorithm's performance, with the system cost serving as the evaluation metric.

[0120] The hardware and software configuration environment of this invention is as follows: the operating system is Ubuntu 22.04, the deep learning framework is PyTorch, the CPU is i5-10700F, the memory is 16G, and the graphics card is NVIDIA GeForce GTX 1070.

[0121] Step 1: Implement the content of each module.

[0122] This invention focuses on the joint optimization objective in a cloud-edge-device collaborative large-scale model inference scenario, and implements the system functions in a modular way, mainly including a network architecture modeling module, an inference performance modeling module, a latency modeling module, and a joint optimization decision module.

[0123] During the system initialization phase, the network architecture module is implemented first. This module is used to build a collaborative reasoning environment between users, drones, and the cloud, and continuously collects the system's operating status within discrete time slots. Specifically, this module maintains user location, task arrival status, drone location and remaining computing resources, as well as communication link parameters between users and drones and between drones and the cloud. It then encapsulates all of this information into a unified network state representation, which serves as the input for subsequent joint decision-making algorithms.

[0124] Subsequently, an inference performance modeling module was implemented to model the accumulation of contextual information and changes in inference performance during large-scale model inference. This module, based on a thought chain hint mechanism, maintains the cumulative number of contextual terms in consecutive time slots and, combined with newly added valid terms, calculates the inference accuracy of different models on different inference tasks. Simultaneously, this module introduces inference performance error costs based on the difference in model scale between cloud and drone models, quantifying the accuracy loss caused by performing inference tasks under resource-constrained conditions.

[0125] Based on this, a latency modeling module is implemented to calculate the communication transmission latency and inference calculation latency of the air edge inference task in two modes: local execution on the UAV and execution in the cloud. The total execution latency of the task is obtained by superimposing them.

[0126] Finally, a joint optimization decision module is implemented. Based on the constructed air-ground collaborative reasoning system, and taking into account the total execution delay of the task, the reasoning accuracy, and resource constraints, a joint optimization problem is constructed, which includes decision variables for UAV selection, language model selection, reasoning task unloading, and UAV trajectory planning. The decision variables are uniformly incorporated into the same optimization framework, and the joint optimization problem is solved by combining the successive convex approximation method with the multi-agent reinforcement learning method.

[0127] Once the environment is set up, training can begin using a multi-agent reinforcement learning algorithm. This invention models the multi-agent reinforcement learning training problem as a sequential decision problem based on the multi-agent advantage decomposition theorem, and introduces a transformer architecture for solving it. Figure 1 The right side shows the architecture of the algorithm, which includes an encoder and a decoder. The agent's decisions are made sequentially one after another, which can significantly reduce the time complexity.

[0128] Step 2: Train the algorithm model.

[0129] After implementing all the content proposed in this invention, we can begin training the algorithm model. As mentioned earlier, this invention requires training a multi-agent reinforcement learning model. Furthermore, considering the strong heterogeneous dependencies between UAV decisions due to constraint C5 in the optimization objective, we propose a multi-agent transform reinforcement learning algorithm to jointly represent the UAV selection decision variables, the language model selection decision variables, and the inference task unloading decision variables. The specific process is shown in the algorithm below:

[0130] Initialization: Parameters of the encoder and decoder networks, replay buffer ;

[0131] Input: Number of training sets and the length of each episode ;

[0132] ; ;

[0133] Obtain observation sequences from the environment ;

[0134] By inputting observation data into the encoder, a representation sequence is generated. ;

[0135] will sequence Input to decoder;

[0136] ;

[0137] Input sequence Get the next action from the decoder ;

[0138] Perform joint actions in the environment Receive reward and current traffic data;

[0139] Will insert And store it in historical traffic data;

[0140] from Randomly select small batches of sample data to train the encoder and decoder networks;

[0141] Train the model and update the network parameters;

[0142] Step 3: Perform simulation to verify the algorithm performance.

[0143] First, the system parameters used in this invention are introduced. The network considered in this invention consists of a single edge server, multiple unmanned aerial vehicles (UAVs), and a large number of terminal devices, and a simulation was performed based on a Shanghai Telecom base station dataset. According to the data in the aforementioned dataset, the location coordinates of all devices are expressed in latitude and longitude values. Unless otherwise specified, the following parameters are used to simulate the system: The edge server is located at (31.24, 121.46), and the number and location of the terminal devices dynamically change based on the actual data in the dataset. Each terminal device has the same computing power (1 GHz) and transmit power (0.1 W). Furthermore, the length of each time slot is one second. The default number of UAVs in the scenario is 4. Each UAV has a computing power of 10 GHz, a transmit power of 0.2 W, a coverage range limited to 50, and a maximum task queue length of 20. Finally, the background noise and channel path loss index are -100 dBm / Hz and -4, respectively.

[0144] Secondly, this invention selects MH, RANDOM and MO as baseline algorithms for comparative experiments, and names the algorithm proposed in this invention as MAT-RA, while the algorithm without adding a trajectory optimization module is denoted as MO. Figure 4 , Figure 5 and Figure 6 The performance comparison between MAT-RA and the baseline algorithm is shown under the same experimental conditions. Figure 4 The different learning capabilities of each algorithm during the training phase are described. The reward trajectory shows that, compared to the benchmark algorithm, MAT-RA exhibits faster convergence speed and more stable convergence state, which is beneficial for solving complex scenarios. Furthermore, this invention evaluates the actual performance of the algorithms, such as... Figure 5As shown in Figure 6, the MAT-RA algorithm achieved the highest system utility and a stable system utility range, consistent with its performance during the training phase, demonstrating its high efficiency. Figure 7 , Figure 8 and Figure 9 Then we will further test the algorithm's performance under different preferences.

Claims

1. A joint optimization method for cloud-edge-device model inference under air-ground collaboration, characterized in that, It includes a network architecture modeling module, an inference performance modeling module, a latency modeling module, and a joint optimization decision module; The network architecture modeling module is used to construct an air-ground collaborative inference system consisting of multiple ground users, multiple drones and the cloud. Multiple drones are introduced to cover the target area, providing air edge inference services for ground users and connecting to the cloud through a backhaul link. The running time of the air-ground collaborative inference system is discretized into multiple time slots of equal length, and the position and speed of the drones and the communication relationship between ground users and drones are characterized in each time slot. The inference performance modeling module is used to maintain the accumulated and newly added context words in continuous time slots for the air edge inference task requested by ground users, and to calculate the impact of the number of context words and the length of the thought chain formed by the context words on the inference accuracy. The latency modeling module calculates the communication transmission latency and inference calculation latency of the air edge inference task in two modes: local execution on the UAV and execution in the cloud. The total execution latency of the task is obtained by summing them up. The joint optimization decision module, based on the constructed air-ground collaborative reasoning system, constructs a joint optimization problem that includes decision variables for UAV selection, language model selection, reasoning task unloading, and UAV trajectory planning, taking into account the total task execution delay, reasoning accuracy, and constraints. The decision variables are uniformly incorporated into the same optimization framework, and the joint optimization problem is solved by combining the successive convex approximation method with the multi-agent reinforcement learning method.

2. The cloud-edge-device model inference joint optimization method under air-ground collaboration as described in claim 1, characterized in that, In the air-ground collaborative reasoning system, the UAV is limited to flying at a fixed altitude and its horizontal speed does not exceed the maximum allowable speed; the uplink transmission rate between the ground user and the UAV is calculated based on the wireless channel model, and the backhaul rate between the UAV and the cloud is set. In the air-ground collaborative reasoning system, the cloud and drones deploy large-scale language models that support the thought chain prompting mechanism for reasoning. The reasoning process takes into account the length of the thought chain and the complexity of the reasoning. The aerial edge inference task requested by ground users is represented as a task tuple containing the original data size, the number of effective lexical units, and the minimum inference accuracy requirement. Decision variables for ground user and UAV selection, language model selection, and inference task offloading are defined to constrain each ground user to select at most one UAV and at most one large-scale language model in a single time slot, and to ensure that the offloading consistency relationship holds.

3. The cloud-edge-device model inference joint optimization method under air-ground collaboration as described in claim 2, characterized in that, The joint optimization problem in the joint optimization decision-making module is as follows: In the joint optimization problem, the UAV selects decision variables. , Indicates ground users In the time slot Should we choose a drone? The decision variables, among which Indicates ground users In the time slot Assigning aerial edge inference tasks to drones Otherwise, it is 0; the inference task unloads decision variables. , Indicates ground users The In time slots, the aerial edge reasoning task via drone Uninstalled to the cloud Each inference instance is executed. Indicates no uninstallation; language model selects decision variables. , Indicates ground users The Aerial edge reasoning tasks in drones Model inference is performed locally, and must satisfy the following conditions: The relationship between them; Represents the drone's position variable. and They represent drones Initial position and drone The final position is determined by trajectory constraints, which ensure smooth movement throughout the entire time slot. Indicates drone In the time slot The horizontal speed variable is subject to the maximum permissible speed. constraint; Indicates in time slot Inside, drones For ground users The airborne edge reasoning task was performed in the first The computational load of inference on a large-scale language model is measured by the number of lexical units, including historically accumulated context lexical units and newly introduced lexical units in the current time slot. This indicates the maximum number of tokens that the GPU on the drone can process in a single time slot; The objective function of the joint optimization problem consists of two parts: This represents the weighted sum of the total execution delays of the tasks, where For ground users No. In time slots, the aerial edge reasoning task Total execution latency, This is the time delay weighting factor; This represents the weighted sum of inference performance error cost under task allocation, where This indicates the performance degradation of airborne edge inference tasks when performed on drones or in the cloud. Precision weighting factor; Constraint C1 limits the horizontal speed of the UAV at arrive Between; Constraint C2 restricts each ground user to selecting at most one drone in a single time slot, and Constraint C3 restricts each ground user's airborne edge inference task to executing at most one large-scale language model within a single time slot. ,and ; Constraint C4 guarantees that the drone is selected only when the ground user selects it. Its aerial edge inference tasks are either executed locally on the drone or offloaded to the cloud, i.e. Constraint C5 ensures that after the aerial edge inference task is associated with the UAV, inference is performed on the local model or offloaded to the cloud for inference. Constraint C6 limits the maximum number of tasks that the UAV can process in a single time slot. Constraint C7 specifies that the initial position of the UAV is The final location of the drone was .

4. The cloud-edge-device model inference joint optimization method under air-ground collaboration as described in claim 3, characterized in that, The inference performance modeling module specifically includes: dynamically updating the accumulated context words in continuous time slots, combining them with the newly introduced context words in the current time slot to form the inference context, forming a complete thought chain input, and calculating the impact of the thought chain on the inference accuracy; using the cloud inference accuracy as a reference benchmark, defining the difference between the UAV local inference accuracy and the cloud inference accuracy as the inference performance error cost.

5. The cloud-edge-device model inference joint optimization method under air-ground collaboration as described in claim 4, characterized in that, The solution process of the joint optimization decision module includes a UAV trajectory optimization sub-process and a UAV resource allocation optimization sub-process; A successive convex approximation method is used to handle the UAV trajectory optimization subprocess. A multi-agent reinforcement learning method is used to handle the sub-process of drone resource allocation optimization.

6. The cloud-edge-device model inference joint optimization method under air-ground collaboration as described in claim 5, characterized in that, The specific sub-process for optimizing the drone trajectory is as follows: Given the decision variables for selecting drones Language model selection of decision variables Unloading decision variables from reasoning tasks Under these conditions, the UAV trajectory optimization sub-process involves adjusting the UAV position variables... The optimization problem is modeled as a non-convex optimization problem, and a first-order Taylor expansion is performed on the trajectory constraint function to construct a locally convex approximation subproblem. Solve the local convex approximation subproblem to obtain the current optimal UAV trajectory, and determine whether the difference between two adjacent trajectory update results is less than a preset convergence threshold. When the convergence condition is met, the current trajectory is output as the final optimization result. At the same time, during the entire iteration process, the UAV flight speed is constrained to meet constraint condition C1, and the UAV start and end position constraint condition C7.

7. The cloud-edge-device model inference joint optimization method under air-ground collaboration as described in claim 5, characterized in that, The specific sub-process for optimizing UAV resource allocation is as follows: Under the condition of fixed drone trajectory, the decision variables for drone selection, language model selection, and inference task unloading are jointly optimized and modeled as a partially observable Markov decision process under a multi-agent reinforcement learning framework, thereby constructing a sub-problem of drone resource allocation optimization: Specifically, each drone is modeled as an independent intelligent agent, whose actions are as follows: Within each discrete time slot, each agent makes decisions based on its local observation sequence, which includes the user task status, UAV resource status, and communication link status. The multi-agent reinforcement learning method is a multi-agent transform reinforcement learning method used to jointly represent the decision variables for UAV selection, the decision variables for language model selection, and the decision variables for inference task unloading. The multi-agent reinforcement learning method uses a Transformer network with an encoder-decoder structure for policy modeling. First, the encoder embeds and extracts features from the local observation sequences acquired in consecutive time slots to capture the temporal correlations and cross-agent dependencies in the multi-agent decision-making process. Based on the encoded representation, the decoder sequentially outputs the drone selection decision variables, the language model selection decision variables, and the inference task unloading decision variables in an autoregressive manner, thereby forming a joint action. During the policy execution phase, each agent interacts with the environment based on the joint actions generated by the current policy, receives immediate rewards, and updates the state of the air-ground collaborative inference system. The reward function consists of a weighted sum of the total task execution latency and the inference performance error cost. in It is a penalty factor; during the training process, the state, action and reward data of the air-ground collaborative reasoning system generated by interaction are stored in the experience replay buffer, and small batch training samples are formed through random sampling, which in turn drive the parameter learning of the encoder and decoder networks and perform policy iterative optimization.

Citation Information

Patent Citations

  • Dynamic frequency and deep learning model unloading joint adjustment method and system based on deep reinforcement learning

    CN115827239A

  • Deep neural network task-oriented collaborative reasoning and unmanned aerial vehicle trajectory optimization algorithm

    CN120257797A