An AI model placement, inference and training task scheduling and resource allocation joint optimization method in edge intelligent network
Patent Information
- Application Number
- CN202610975924.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]然而,现有研究在边缘智能网络中往往将推理任务和训练任务分开考虑,缺乏对二者的联合优化
[0107]本发明采用以上技术方案与现有技术相比,具有以下有益效果:
Smart Images

Figure CN122802973A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of edge computing and artificial intelligence technology, specifically to a joint optimization method for AI model placement, inference and training task scheduling, and resource allocation in an edge intelligent network. Background Technology
[0002] In recent years, artificial intelligence technology has developed rapidly, and user devices have increasingly demanded AI services. However, the limited computing and storage resources of user devices make it difficult to support local AI model inference, while cloud-based solutions suffer from high latency and network congestion. Edge intelligence, as an emerging paradigm that integrates AI with edge computing, provides a low-latency, high-efficiency solution to these challenges by extending AI capabilities to the network edge.
[0003] In edge intelligent networks, AI model inference and AI model training coexist and complement each other. On the one hand, AI models perform inference to meet the diverse service needs of user devices; on the other hand, AI models adapt to dynamic environments and changes in user behavior through continuous training, thereby improving inference accuracy and adaptability. Inference tasks are typically characterized by low latency and diversity, while training tasks are characterized by high computational load and long processing time, and rely on training data collected from distributed user devices.
[0004] However, existing research in edge intelligent networks often considers inference and training tasks separately, lacking joint optimization of the two. Furthermore, some studies on edge node training tasks neglect the distributed nature of training data, typically assuming it is given and ignoring optimization of data transmission from user devices to edge nodes. Only a few studies have explored the joint optimization of AI model inference and training in edge intelligent networks, but these studies usually focus on single AI models, making them difficult to apply to the joint placement and scheduling of multiple heterogeneous AI models. Simultaneously, updating AI model placement decisions usually involves additional system overhead (such as model download and initialization), making frequent updates unsuitable, while task scheduling and resource allocation require dynamic decisions in each time slot. This difference in the update frequency of different variables further exacerbates the complexity of joint optimization. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] To overcome the shortcomings of the prior art, this invention targets edge intelligent networks and comprehensively considers factors such as inference task arrival rate, training data size, network transmission rate, and edge server computing and storage capabilities to jointly optimize AI model placement, inference and training task scheduling, and resource allocation decisions. The proposed method is integrated into the container (or virtual machine) deployment module and resource configuration module of the server management system. Combined with network status and model performance perception acquisition model, it realizes optimal AI model placement, dynamic scheduling of inference and training tasks, and dynamic optimization of edge server resource allocation across multiple edge servers. By comprehensively considering inference task processing and training task execution, it ensures the efficient operation of the edge intelligent system and improves the quality of AI services.
[0007] (II) Technical Solution
[0008] To address the aforementioned technical problems, this invention proposes a joint optimization method for AI model placement, inference and training task scheduling, and resource allocation in edge intelligent networks, comprising the following steps:
[0009] S1: Construct an edge intelligent network system model, an AI model placement model, and an inference training task scheduling model;
[0010] S2: Construct a model of inference task processing latency and energy consumption, and train the model of inference task processing latency and energy consumption, thereby establishing a system overhead optimization problem;
[0011] S3: Decompose the system overhead optimization problem established in S2, solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm, solve the optimal edge server transmission rate allocation decision based on the Karush-Kuhn-Tucker condition, and solve the optimal edge server computing resource allocation decision based on the convex optimization method.
[0012] S4: Construct the AI model placement and inference training task scheduling subproblem into a dual-timescale Markov decision process, define the system's state, action and reward, and design an adaptive two-layer deep reinforcement learning framework.
[0013] S5: Based on deep reinforcement learning technology, train and apply joint AI model placement, inference and training task scheduling and resource allocation strategies.
[0014] Furthermore, step S1 specifically includes:
[0015] (1) Constructing an edge intelligent network system model. The edge computing network consists of... The system consists of several edge servers. Edge servers are deployed at base stations, which are connected to each other via a wired access network. Each edge server has multiple user devices within its coverage area. User devices generate various types of AI inference tasks and offload them to the edge servers for processing; simultaneously, user devices upload data required for AI model training to the edge servers where the corresponding AI models are deployed. The set of edge servers is represented as... Edge server Total computing resources are The total storage resources are represented by This indicates that the cloud server's index is 0, and its total computing resources are... The total storage resources are represented by Indicates. Define the time interval as... .
[0016] (2) The set of user equipment is represented as The set of user devices within the coverage area of the edge server consists of... The set of AI model types is represented as follows. Each type of reasoning task can only be handled by the corresponding type of AI model. The reasoning task is represented as ,in Indicates the first Data size for reasoning tasks Indicates processing the first Computational load required for class-based reasoning tasks. User equipment. The first The number of reasoning tasks is represented as .
[0017] (3) AI model placement model. Definition To deploy the first Storage resources required for AI-like models. Introducing binary variables. ,in Used to indicate whether it is on the edge server The first was deployed on AI-like models. Specifically, if the first... AI-like models are deployed on edge servers Up, settings The placement of AI models must meet the storage capacity constraints of edge servers. Using sets This indicates where all AI models place variables.
[0018] (4) Inference task scheduling model. Introducing binary variables. ,in Used to represent user equipment The first Scheduling decisions for reasoning tasks. Specifically, if the user device... The first Inference tasks are scheduled to edge servers. ,set up Inference task scheduling must satisfy: This means that inference tasks can only be scheduled to edge servers where the corresponding AI models have been deployed; and This means that each inference task must be scheduled to exactly one edge server. Using a set... This represents the scheduling variables for all inference tasks.
[0019] (5) Train the task scheduling model. Introduce binary variables. ,in Used to represent edge servers The first Scheduling decisions for training tasks. Specifically, if the edge server... The first The training task was scheduled to the edge server. ,set up Training task scheduling must meet the following requirements: ,and Using sets This represents the scheduling variables for all training tasks.
[0020] Furthermore, step S2 specifically includes:
[0021] (1) Construct a model of inference task processing latency and energy consumption.
[0022] Since no AI model is deployed on the user device, the inference tasks generated by the user device can only be offloaded to edge servers with the corresponding AI models for processing. Inference task processing includes four stages: the user device uploads the inference task to its affiliated edge server; the affiliated edge server further transmits the inference task to the target edge server with the corresponding AI model deployed; the target edge server performs the inference computation; and the inference result is returned to the user device along the reverse path. Because the size of the inference result data is much smaller than the size of the inference task data, the latency of the inference result return process is negligible.
[0023] The wireless access system for each edge server is defined based on Orthogonal Frequency Division Multiple Access (OFDMA), where multiple user devices associated with the same edge server transmit through orthogonal channels, thus avoiding inter-channel interference. Bandwidth is determined by... This indicates that the channel gain is determined by... This indicates that the noise power spectral density is derived from... This indicates that the average channel interference power is determined by... Indicate. Assume user equipment Maximum transmit power is Introducing variables Indicates user equipment Offload its inference tasks to the edge server's transmit power. According to Shannon's theorem, the user equipment... The transmission rate to its affiliated edge server is:
[0024]
[0025] User equipment The first The upload latency from the inference task to its associated edge server is:
[0026]
[0027] User equipment Uninstall the The energy consumption of reasoning tasks is:
[0028]
[0029] If the corresponding AI model is not deployed on the auxiliary edge server, the inference task needs to be further offloaded to another edge server where the corresponding AI model is deployed. Define the edge server. and edge servers The total transmission rate between them is Introducing variables Represents edge server and Allocation between transmission user equipment The Transmission rate for inference tasks. User equipment. The first The transmission latency of reasoning tasks between edge servers is:
[0030]
[0031] The corresponding transmission power consumption is:
[0032]
[0033] in For edge servers The transmission power.
[0034] The computation latency of the inference task on the target edge server is:
[0035]
[0036] The corresponding calculated energy consumption is:
[0037]
[0038] in, To calculate the energy factor.
[0039] (2) Construct a training task processing latency and energy consumption model.
[0040] When the inference accuracy of an AI model deployed on an edge server consistently falls below a preset threshold, a corresponding model training task is triggered. Edge server The first The number of training tasks per class is represented as ,in . No. The training task consists of tuples Depiction, among which Indicates the mini-batch size. Indicates the number of training rounds. This indicates the computational load for each batch of data within a single round. This represents the set of user devices that provide training data.
[0041] The training task processing includes five stages: model parameters are transferred from the initiating edge server to the target edge server; training data is uploaded from the user device to its affiliated edge server; training data is transferred from the affiliated edge server to the target edge server; the target edge server performs training calculations; and the updated model parameters are synchronized to all edge servers that have deployed the AI model.
[0042] The model parameter transmission delay is:
[0043]
[0044] in yes The data size of hyperparameters in AI-like models The transmission rate allocated to the training data.
[0045] The corresponding transmission power consumption is:
[0046]
[0047] The training data upload latency is:
[0048]
[0049] The corresponding upload power consumption is:
[0050]
[0051] in For user equipment The uploaded number The size of the training data.
[0052] The training data transmission latency is:
[0053]
[0054] The corresponding transmission power consumption is:
[0055]
[0056] The training computation latency is:
[0057]
[0058] in For the first Total data size for the training tasks For edge servers Computational resources allocated to the training task.
[0059] The corresponding calculated energy consumption is:
[0060]
[0061] The model parameter synchronization delay is:
[0062]
[0063] The corresponding synchronization energy consumption is:
[0064]
[0065] (3) Construct an optimization problem.
[0066] Total system latency This includes the latency of each stage of the inference task and the sum of the latency of each stage of the training task. Total system energy consumption. This includes the sum of user equipment power consumption and edge server power consumption. They are expressed as follows:
[0067]
[0068]
[0069] The total system overhead is defined as the weighted sum of total latency and total energy consumption: ,in and The weighting coefficients for latency and energy consumption are respectively, satisfying... .
[0070] Combining latency and energy consumption models, this paper addresses the optimization problems of joint AI model placement, inference and training task scheduling, and resource allocation in edge intelligent networks. :
[0071]
[0072] in, This indicates where the AI model places the variable set. This represents the scheduling variable for the inference task. Represents the training task scheduling variable. Represents the set of user equipment transmit power allocation variables. This represents the set of variables for allocating transmission rates between edge servers. This represents the set of variables for allocating computing resources to the edge server.
[0073] Furthermore, step S3 specifically includes:
[0074] (1) Optimization problem The problem is decomposed into three numerical optimization subproblems and one deep reinforcement learning subproblem, which are the user equipment transmit power allocation subproblems. Edge server transmission rate allocation sub-problem The sub-problem of edge server computing resource allocation And the sub-problem of AI model placement and inference training task scheduling. .
[0075] (2) Solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm. The user equipment transmit power variable... Decouple from the original problem and construct a subproblem of transmit power allocation. This subproblem has a linear numerator and a convex denominator, and the constraints are linear. Therefore, the Dinkelbach fractional programming algorithm can be used to efficiently solve the optimal user equipment transmit power allocation decision. .
[0076] (3) Solve the optimal edge server transmission rate allocation decision based on the Karush-Kuhn-Tucker conditions. Place the decision in the given AI model. Reasoning task scheduling decision Training task scheduling decision Under the condition of allocating variables to the transmission rate and Decouple from the original problem and construct a transmission rate allocation subproblem. This subproblem is a convex optimization problem, which can be solved using the KKT conditions to determine the optimal edge server transmission rate allocation decision. .
[0077] (4) Solve the optimal edge server computing resource allocation decision based on the convex optimization method. Given the relevant decision variables, the computing resource allocation variables are... and Decouple from the original problem and construct a subproblem of computing resource allocation. The objective function of this subproblem can be expressed as: and The convex function of the sum of terms can be used to solve the optimal edge server computing resource allocation decision using a convex optimization framework (such as CVXPY). .
[0078] Furthermore, step S4 specifically includes:
[0079] (1) Construct a dual-timescale Markov decision process for AI model placement and inference training task scheduling. Define the long update period as... Each long update cycle is composed of It consists of short update cycles. An adaptive two-layer deep reinforcement learning architecture is constructed, with the upper policy network responsible for AI model placement decisions and the lower policy network responsible for inference and training task scheduling decisions. The upper policy network adaptively determines whether to update the AI model placement policy and its update frequency, while the lower policy network updates task scheduling decisions in each time slot.
[0080] (2) The Markov decision process placed in the upper-level AI model is defined as follows:
[0081] Upper-level actions It consists of two parts: an update flag And AI model placement decision .when If the AI model's placement decision is updated, it is updated; otherwise, the placement decision from the previous time slot remains unchanged.
[0082] upper state This includes: the number of inference tasks generated by each user device. The number of training tasks generated by each edge server AI model placement strategy update interval And the AI model placement decision from the previous time slot. .
[0083] Upper-level rewards Defined as the negative value of the total system overhead: .
[0084] Introducing penalty items To reduce the update frequency of AI model placement strategies: ,in This represents the penalty coefficient. The larger the update interval, the smaller the penalty; the greater the difference between the old and new placement strategies, the greater the penalty.
[0085] (3) The Markov decision process for lower-level inference and training task scheduling is defined as follows:
[0086] Lower-level actions Includes: all inference task scheduling decisions and all training task scheduling decisions The lower-level task scheduling decision is updated in each time slot.
[0087] Lower state Including: Current AI model placement strategy The number of inference tasks generated by each user device The number of training tasks generated by each edge server Channel conditions Total transmission rate between edge servers and total computing resources of edge servers .
[0088] Lower-level rewards Defined as the negative value of the total system overhead: .
[0089] Furthermore, step S5 specifically includes:
[0090] (1) Design an adaptive two-layer deep reinforcement learning algorithm based on TD3 (Twin Delayed Deep Deterministic Policy Gradient) and construct six neural networks for the upper policy network: main Actor network Target Actor Network Two main Critic networks and Two target Critic networks ¹and Six neural networks that construct the lower-level policy network: the main Actor network. Target Actor Network Two main Critic networks and Two target Critic networks and .
[0091] (2) Train an adaptive two-layer deep reinforcement learning algorithm. In each time slot, the upper-layer policy network adjusts the algorithm according to the current environmental state. Output Action The lower-level policy network determines whether to update the AI model placement strategy and make new placement decisions based on the current environmental state. (Including the current AI model placement decision) Output action The scheduling decisions for inference and training tasks are determined. The AI model placement and task scheduling decisions are input into a numerical optimization subroutine, which solves for optimal power allocation, transmission rate allocation, and computational resource allocation based on the Dinkelbach algorithm, KKT conditions, and convex optimization methods, respectively. After the environment executes all decisions, the total system overhead is calculated, rewards and penalties are generated, and experience tuples are constructed and stored in the upper-level experience replay pool. and lower-level experience replay pool Once the experience replay pool is full, a batch of experiences is randomly sampled from it. The Actor network parameters are updated using the policy gradient method, the Critic network parameters are updated by minimizing the temporal difference error loss function, and the target network parameters are updated using a soft update algorithm. The specific steps of the training process are as follows:
[0092] Step 1: Randomly initialize the neural network parameters of the upper and lower policy networks, and initialize time slots. Set the interval to 1, and initialize the update interval. =1;
[0093] Step 2: Determine the output of the upper-layer policy network If the value is greater than 0.5, update the AI model's placement decision based on the output of the upper-level policy network. and reset If it is 1; otherwise, keep the AI model placement decision the same as the previous time slot, and at the same time... ;
[0094] Step 3: The lower-level policy network determines the state based on... And the current AI model placement decision Output inference task scheduling decision Training task scheduling decision ;
[0095] Step 4: Solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm. ;
[0096] Step 5: Solve for the optimal edge server transmission rate allocation decision based on KKT conditions ;
[0097] Step 6: Solve the optimal edge server computing resource allocation decision based on convex optimization methods ;
[0098] Step 7: Calculate the total system overhead based on the results of steps 3 to 6. Then calculate the lower-level rewards. ;
[0099] Step 8: Transfer the lower-level experience tuples Store in the lower-level experience replay pool ;
[0100] Step 9: Transfer the upper-level experience tuples Store in the upper-level experience replay pool ;
[0101] Step 10: Replay experience from the lower-level experience pool Randomly sample batches of experience and update the parameters of the lower-level Critic network. and ;
[0102] Step 11: Replay experience from the upper-level experience pool Based on random sampling batch experience, update the parameters of the upper-layer Critic network. and ;
[0103] Step 12: Determine the time slot Is it Multiples of, and so on, from, respectively and Based on the experience of sampling batches, update the parameters of the lower and upper layer Actor networks. and Update all target network parameters according to the soft update algorithm. ;
[0104] Step 13: Let The value is incremented by 1 to determine if the current episode has ended: if so, the system state is reset and a new episode begins;
[0105] Step 14: Determine if the training process has ended: If yes, end the training process; otherwise, proceed to step 2.
[0106] (III) Beneficial Effects
[0107] Compared with the prior art, the present invention, employing the above technical solution, has the following beneficial effects:
[0108] 1. When jointly optimizing AI model placement, inference and training task scheduling and resource allocation, an inference-training collaborative operation mode is adopted, which comprehensively considers the latency and energy consumption of inference task processing and training task execution (including distributed training data collection), making it more practical in edge intelligent networks where multiple types of AI models coexist.
[0109] 2. When optimizing AI model placement variables, task scheduling variables, and resource allocation variables, we considered that different variables have different optimization frequencies and constructed an adaptive two-layer deep reinforcement learning architecture. The upper layer makes decisions on model placement at a variable frequency, and the lower layer schedules tasks at a fixed frequency, which is more practical.
[0110] 3. Based on TD3 deep reinforcement learning technology, combined with numerical methods such as Dinkelbach, KKT conditions and convex optimization, AI model placement, task scheduling and resource allocation are jointly optimized. This can effectively adapt to the dynamic network conditions of edge environments as well as the dynamic inference and training triggering needs of users, reduce the total system overhead and improve the quality of AI services. Attached Figure Description
[0111] Figure 1 This is a schematic diagram of the system model of an embodiment;
[0112] Figure 2 This is a schematic diagram of the training process of an adaptive two-layer DRL model that combines AI model placement, inference and training task scheduling and resource allocation.
[0113] Figure 3 This is a schematic diagram illustrating the application process of an adaptive two-layer DRL model that combines AI model placement, inference and training task scheduling, and resource allocation.
[0114] Figure 4 This is a graph showing the relationship between the number of users and the average total system overhead in an edge intelligent computing network.
[0115] Figure 5 This is a graph showing the relationship between the size of AI model training data and the average total system overhead in an edge intelligent computing network.
[0116] Figure 6 This is a graph showing the relationship between the computing resources of the edge server and the average total system overhead.
[0117] Figure 7 This is a graph showing the relationship between the transmission rate of the edge server and the average total system overhead. Detailed Implementation
[0118] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0119] This invention proposes a joint optimization method for AI model placement, inference and training task scheduling, and resource allocation in edge intelligent networks. The embodiments include the following steps:
[0120] Step 1: Refer to Figure 1Construct an edge intelligent network system model, an AI model placement model, and an inference training task scheduling model;
[0121] Step 2: Construct a model of inference task processing latency and energy consumption, and train the model of inference task processing latency and energy consumption, thereby establishing a system overhead optimization problem;
[0122] Step 3: Decompose the system overhead optimization problem established in S2, solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm, solve the optimal edge server transmission rate allocation decision based on the Karush-Kuhn-Tucker condition, and solve the optimal edge server computing resource allocation decision based on the convex optimization method.
[0123] Step 4: Construct the AI model placement and inference training task scheduling subproblem into a dual-timescale Markov decision process, define the system's state, actions, and rewards, and design an adaptive two-layer deep reinforcement learning framework.
[0124] Step 5: Refer to Figure 2 and Figure 3 Based on deep reinforcement learning technology, specifically the TD3 algorithm, a joint AI model placement, inference and training task scheduling and resource allocation strategy is trained and applied.
[0125] Furthermore, step one includes:
[0126] (1) Constructing an edge intelligent network system model. The edge computing network consists of... The system consists of several edge servers. Edge servers are deployed at base stations, which are connected to each other via a wired access network. Each edge server has multiple user devices within its coverage area. User devices generate various types of AI inference tasks and offload them to the edge servers for processing; simultaneously, user devices upload data required for AI model training to the edge servers where the corresponding AI models are deployed. The set of edge servers is represented as... Edge server Total computing resources are The total storage resources are represented by This indicates that the cloud server's index is 0, and its total computing resources are... The total storage resources are represented by Indicates. Define the time interval as... .
[0127] (2) The set of user equipment is represented as The set of user devices within the coverage area of the edge server consists of... The set of AI model types is represented as follows. Each type of reasoning task can only be handled by the corresponding type of AI model. The reasoning task is represented as ,in Indicates the first Data size for reasoning tasks Indicates processing the first Computational load required for class-based reasoning tasks. User equipment. The first The number of reasoning tasks is represented as .
[0128] (3) AI model placement model. Definition To deploy the first Storage resources required for AI-like models. Introducing binary variables. ,in Used to indicate whether it is on the edge server The first was deployed on AI-like models. Specifically, if the first... AI-like models are deployed on edge servers Up, settings The placement of AI models must meet the storage capacity constraints of edge servers. Using sets This indicates where all AI models place variables.
[0129] (4) Inference task scheduling model. Introducing binary variables. ,in Used to represent user equipment The first Scheduling decisions for reasoning tasks. Specifically, if the user device... The first Inference tasks are scheduled to edge servers. ,set up Inference task scheduling must satisfy: This means that inference tasks can only be scheduled to edge servers where the corresponding AI models have been deployed; and This means that each inference task must be scheduled to exactly one edge server. Using a set... This represents the scheduling variables for all inference tasks.
[0130] (5) Train the task scheduling model. Introduce binary variables. ,in Used to represent edge servers The first Scheduling decisions for training tasks. Specifically, if the edge server... The first The training task was scheduled to the edge server. ,set up Training task scheduling must meet the following requirements: ,and Using sets This represents the scheduling variables for all training tasks.
[0131] Furthermore, step two includes:
[0132] (1) Construct a model of inference task processing latency and energy consumption.
[0133] Since no AI model is deployed on the user device, the inference tasks generated by the user device can only be offloaded to edge servers with the corresponding AI models for processing. Inference task processing includes four stages: the user device uploads the inference task to its affiliated edge server; the affiliated edge server further transmits the inference task to the target edge server with the corresponding AI model deployed; the target edge server performs the inference computation; and the inference result is returned to the user device along the reverse path. Because the size of the inference result data is much smaller than the size of the inference task data, the latency of the inference result return process is negligible.
[0134] The wireless access system for each edge server is defined based on Orthogonal Frequency Division Multiple Access (OFDMA), where multiple user devices associated with the same edge server transmit through orthogonal channels, thus avoiding inter-channel interference. Bandwidth is determined by... This indicates that the channel gain is determined by... This indicates that the noise power spectral density is derived from... This indicates that the average channel interference power is determined by... Indicate. Assume user equipment Maximum transmit power is Introducing variables Indicates user equipment Offload its inference tasks to the edge server's transmit power. According to Shannon's theorem, the user equipment... The transmission rate to its affiliated edge server is:
[0135]
[0136] User equipment The first The upload latency from the inference task to its associated edge server is:
[0137]
[0138] User equipment Uninstall the The energy consumption of reasoning tasks is:
[0139]
[0140] If the corresponding AI model is not deployed on the auxiliary edge server, the inference task needs to be further offloaded to another edge server where the corresponding AI model is deployed. Define the edge server. and edge servers The total transmission rate between them is Introducing variables Represents edge server and Allocation between transmission user equipment The Transmission rate for inference tasks. User equipment. The first The transmission latency of reasoning tasks between edge servers is:
[0141]
[0142] The corresponding transmission power consumption is:
[0143]
[0144] in For edge servers The transmission power.
[0145] The computation latency of the inference task on the target edge server is:
[0146]
[0147] The corresponding calculated energy consumption is:
[0148]
[0149] in, To calculate the energy factor.
[0150] (2) Construct a training task processing latency and energy consumption model.
[0151] When the inference accuracy of an AI model deployed on an edge server consistently falls below a preset threshold, a corresponding model training task is triggered. Edge server The first The number of training tasks per class is represented as ,in . No. The training task consists of tuples Depiction, among which Indicates the mini-batch size. Indicates the number of training rounds. This indicates the computational load for each batch of data within a single round. This represents the set of user devices that provide training data.
[0152] The training task processing includes five stages: model parameters are transferred from the initiating edge server to the target edge server; training data is uploaded from the user device to its affiliated edge server; training data is transferred from the affiliated edge server to the target edge server; the target edge server performs training calculations; and the updated model parameters are synchronized to all edge servers that have deployed the AI model.
[0153] The model parameter transmission delay is:
[0154]
[0155] in yes The data size of hyperparameters in AI-like models The transmission rate allocated to the training data.
[0156] The corresponding transmission power consumption is:
[0157]
[0158] The training data upload latency is:
[0159]
[0160] The corresponding upload power consumption is:
[0161]
[0162] in For user equipment The uploaded number The size of the training data.
[0163] The training data transmission latency is:
[0164]
[0165] The corresponding transmission power consumption is:
[0166]
[0167] The training computation latency is:
[0168]
[0169] in For the first Total data size for the training tasks For edge servers Computational resources allocated to the training task.
[0170] The corresponding calculated energy consumption is:
[0171]
[0172] The model parameter synchronization delay is:
[0173]
[0174] The corresponding synchronization energy consumption is:
[0175]
[0176] (3) Construct an optimization problem.
[0177] Total system latency This includes the latency of each stage of the inference task and the sum of the latency of each stage of the training task. Total system energy consumption. This includes the sum of user equipment power consumption and edge server power consumption. They are expressed as follows:
[0178]
[0179]
[0180] The total system overhead is defined as the weighted sum of total latency and total energy consumption: ,in and The weighting coefficients for latency and energy consumption are respectively, satisfying... .
[0181] Combining latency and energy consumption models, this paper addresses the optimization problems of joint AI model placement, inference and training task scheduling, and resource allocation in edge intelligent networks. :
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196] in, This indicates where the AI model places the variable set. This represents the scheduling variable for the inference task. Represents the training task scheduling variable. Represents the set of user equipment transmit power allocation variables. This represents the set of variables for allocating transmission rates between edge servers. This represents the set of variables for allocating computing resources to the edge server.
[0197] Furthermore, step three includes:
[0198] (1) Optimization problem The problem is decomposed into three numerical optimization subproblems and one deep reinforcement learning subproblem, which are the user equipment transmit power allocation subproblems. Edge server transmission rate allocation sub-problem The sub-problem of edge server computing resource allocation And the sub-problem of AI model placement and inference training task scheduling. .
[0199] (2) Solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm. The user equipment transmit power variable... Decouple from the original problem and construct a subproblem of transmit power allocation. This subproblem has a linear numerator and a convex denominator, and the constraints are linear. Therefore, the Dinkelbach fractional programming algorithm can be used to efficiently solve the optimal user equipment transmit power allocation decision. .
[0200] (3) Solve the optimal edge server transmission rate allocation decision based on the Karush-Kuhn-Tucker conditions. Place the decision in the given AI model. Reasoning task scheduling decision Training task scheduling decision Under the condition of allocating variables to the transmission rate and Decouple from the original problem and construct a transmission rate allocation subproblem. This subproblem is a convex optimization problem, which can be solved using the KKT conditions to determine the optimal edge server transmission rate allocation decision. .
[0201] (4) Solve the optimal edge server computing resource allocation decision based on the convex optimization method. Given the relevant decision variables, the computing resource allocation variables are... and Decouple from the original problem and construct a subproblem of computing resource allocation. The objective function of this subproblem can be expressed as: and The convex function of the sum of terms can be used to solve the optimal edge server computing resource allocation decision using a convex optimization framework (such as CVXPY). .
[0202] Furthermore, step four includes:
[0203] (1) Construct a dual-timescale Markov decision process for AI model placement and inference training task scheduling. Define the long update period as... Each long update cycle is composed of It consists of short update cycles. An adaptive two-layer deep reinforcement learning architecture is constructed, with the upper policy network responsible for AI model placement decisions and the lower policy network responsible for inference and training task scheduling decisions. The upper policy network adaptively determines whether to update the AI model placement policy and its update frequency, while the lower policy network updates task scheduling decisions in each time slot.
[0204] (2) The Markov decision process placed in the upper-level AI model is defined as follows:
[0205] Upper-level actions It consists of two parts: an update flag And AI model placement decision .when If the AI model's placement decision is updated, it is updated; otherwise, the placement decision from the previous time slot remains unchanged.
[0206] upper state This includes: the number of inference tasks generated by each user device. The number of training tasks generated by each edge server AI model placement strategy update interval And the AI model placement decision from the previous time slot. .
[0207] Upper-level rewards Defined as the negative value of the total system overhead: .
[0208] Introducing penalty items To reduce the update frequency of AI model placement strategies: ,in This represents the penalty coefficient. The larger the update interval, the smaller the penalty; the greater the difference between the old and new placement strategies, the greater the penalty.
[0209] (3) The Markov decision process for lower-level inference and training task scheduling is defined as follows:
[0210] Lower-level actions Includes: all inference task scheduling decisions and all training task scheduling decisions The lower-level task scheduling decision is updated in each time slot.
[0211] Lower state Including: Current AI model placement strategy The number of inference tasks generated by each user device The number of training tasks generated by each edge server Channel conditions Total transmission rate between edge servers and total computing resources of edge servers .
[0212] Lower-level rewards Defined as the negative value of the total system overhead: .
[0213] Furthermore, step five includes:
[0214] (1) Design an adaptive two-layer deep reinforcement learning algorithm based on TD3 (Twin Delayed Deep Deterministic Policy Gradient) and construct six neural networks for the upper policy network: main Actor network Target Actor Network Two main Critic networks and Two target Critic networks ¹and Six neural networks that construct the lower-level policy network: the main Actor network. Target Actor Network Two main Critic networks and Two target Critic networks and .
[0215] (2) Train an adaptive two-layer deep reinforcement learning algorithm. In each time slot, the upper-layer policy network adjusts the algorithm according to the current environmental state. Output Action The lower-level policy network determines whether to update the AI model placement strategy and make new placement decisions based on the current environmental state. (Including the current AI model placement decision) Output action The scheduling decisions for inference and training tasks are determined. The AI model placement and task scheduling decisions are input into a numerical optimization subroutine, which solves for optimal power allocation, transmission rate allocation, and computational resource allocation based on the Dinkelbach algorithm, KKT conditions, and convex optimization methods, respectively. After the environment executes all decisions, the total system overhead is calculated, rewards and penalties are generated, and experience tuples are constructed and stored in the upper-level experience replay pool. and lower-level experience replay pool Once the experience replay pool is full, a batch of experiences is randomly sampled from it. The Actor network parameters are updated using the policy gradient method, the Critic network parameters are updated by minimizing the temporal difference error loss function, and the target network parameters are updated using a soft update algorithm. The specific steps of the training process are as follows:
[0216] Step 1: Randomly initialize the neural network parameters of the upper and lower policy networks, and initialize time slots. Set the interval to 1, and initialize the update interval. =1;
[0217] Step 2: Determine the output of the upper-layer policy network If the value is greater than 0.5, update the AI model's placement decision based on the output of the upper-level policy network. and reset If it is 1; otherwise, keep the AI model placement decision the same as the previous time slot, and at the same time... ;
[0218] Step 3: The lower-level policy network determines the state based on... And the current AI model placement decision Output inference task scheduling decision Training task scheduling decision ;
[0219] Step 4: Solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm. ;
[0220] Step 5: Solve for the optimal edge server transmission rate allocation decision based on KKT conditions ;
[0221] Step 6: Solve the optimal edge server computing resource allocation decision based on convex optimization methods ;
[0222] Step 7: Calculate the total system overhead based on the results of steps 3 to 6. Then calculate the lower-level rewards. ;
[0223] Step 8: Transfer the lower-level experience tuples Store in the lower-level experience replay pool ;
[0224] Step 9: Transfer the upper-level experience tuples Store in the upper-level experience replay pool ;
[0225] Step 10: Replay experience from the lower-level experience pool Randomly sample batches of experience and update the parameters of the lower-level Critic network. and ;
[0226] Step 11: Replay experience from the upper-level experience pool Based on random sampling batch experience, update the parameters of the upper-layer Critic network. and ;
[0227] Step 12: Determine the time slot Is it Multiples of, and so on, from, respectively and Based on the experience of sampling batches, update the parameters of the lower and upper layer Actor networks. and Update all target network parameters according to the soft update algorithm. ;
[0228] Step 13: Let The value is incremented by 1 to determine if the current episode has ended: if so, the system state is reset and a new episode begins;
[0229] Step 14: Determine if the training process has ended: If yes, end the training process; otherwise, proceed to step 2.
[0230] Figure 4 , Figure 5 , Figure 6 and Figure 7 The paper presents a comparison of the disclosed solution with other solutions under different conditions regarding the number of users in different edge computing networks, the size of AI model training data in different edge computing networks, and the computing resources and transmission rates of different edge servers. Experimental results show that the disclosed solution can achieve joint optimization of AI model placement, inference and training task scheduling, and resource allocation in edge intelligent dynamic network environments and under dynamic user needs.
[0231] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Therefore, the present invention is not limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A joint optimization method for AI model placement, inference and training task scheduling, and resource allocation in an edge intelligent network, characterized in that, Includes the following steps: S1: Construct an edge intelligent network system model, an AI model placement model, and an inference training task scheduling model; S2: Construct a model of inference task processing latency and energy consumption, and train the model of inference task processing latency and energy consumption, thereby establishing a system overhead optimization problem; S3: Decompose the system overhead optimization problem established in S2, solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm, solve the optimal edge server transmission rate allocation decision based on the Karush-Kuhn-Tucker condition, and solve the optimal edge server computing resource allocation decision based on the convex optimization method. S4: Construct the AI model placement and inference training task scheduling subproblem into a dual-timescale Markov decision process, define the system's state, action and reward, and design an adaptive two-layer deep reinforcement learning framework. S5: Based on deep reinforcement learning technology, train and apply joint AI model placement, inference and training task scheduling and resource allocation strategies.
2. The joint optimization method for AI model placement, inference and training task scheduling and resource allocation in an edge intelligent network according to claim 1, characterized in that, Step S1 includes: (1) Constructing an edge intelligent network system model. The edge computing network consists of... The system consists of several edge servers. Edge servers are deployed at base stations, which are connected to each other via a wired access network. Each edge server has multiple user devices within its coverage area. User devices generate various types of AI inference tasks and offload them to the edge servers for processing; simultaneously, user devices upload data required for AI model training to the edge servers where the corresponding AI models are deployed. The set of edge servers is represented as... Edge server Total computing resources are The total storage resources are represented by This indicates that the cloud server's index is 0, and its total computing resources are... The total storage resources are represented by Indicates. Define the time interval as... . (2) The set of user equipment is represented as The set of user devices within the coverage area of the edge server consists of... The set of AI model types is represented as follows. Each type of reasoning task can only be handled by the corresponding type of AI model. The reasoning task is represented as ,in Indicates the first Data size for reasoning tasks Indicates processing the first Computational load required for class-based reasoning tasks. User equipment. The first The number of reasoning tasks is represented as . (3) AI model placement model. Definition To deploy the first Storage resources required for AI-like models. Introducing binary variables. ,in Used to indicate whether it is on the edge server The first was deployed on AI-like models. Specifically, if the first... AI-like models are deployed on edge servers Up, settings The placement of AI models must meet the storage capacity constraints of edge servers. Using sets This indicates where all AI models place variables. (4) Inference task scheduling model. Introducing binary variables. ,in Used to represent user equipment The first Scheduling decisions for reasoning tasks. Specifically, if the user device... The first Inference tasks are scheduled to edge servers. ,set up Inference task scheduling must satisfy: This means that inference tasks can only be scheduled to edge servers where the corresponding AI models have been deployed; and This means that each inference task must be scheduled to exactly one edge server. Using a set... This represents the scheduling variables for all inference tasks. (5) Train the task scheduling model. Introduce binary variables. ,in Used to represent edge servers The first Scheduling decisions for training tasks. Specifically, if the edge server... The first The training task was scheduled to the edge server. ,set up Training task scheduling must meet the following requirements: ,and Using sets This represents the scheduling variables for all training tasks.
3. The joint optimization method for AI model placement, inference and training task scheduling and resource allocation in an edge intelligent network according to claim 1, characterized in that, Step S2 includes: (1) Construct a model of inference task processing latency and energy consumption. Since no AI model is deployed on the user device, the inference tasks generated by the user device can only be offloaded to edge servers with the corresponding AI models for processing. Inference task processing includes four stages: the user device uploads the inference task to its affiliated edge server; the affiliated edge server further transmits the inference task to the target edge server with the corresponding AI model deployed; the target edge server performs the inference computation; and the inference result is returned to the user device along the reverse path. Because the size of the inference result data is much smaller than the size of the inference task data, the latency of the inference result return process is negligible. The wireless access system for each edge server is defined based on Orthogonal Frequency Division Multiple Access (OFDMA), where multiple user devices associated with the same edge server transmit through orthogonal channels, thus avoiding inter-channel interference. Bandwidth is determined by... This indicates that the channel gain is determined by... This indicates that the noise power spectral density is derived from... This indicates that the average channel interference power is determined by... Indicate. Assume user equipment Maximum transmit power is Introducing variables Indicates user equipment Offload its inference tasks to the edge server's transmit power. According to Shannon's theorem, the user equipment... The transmission rate to its affiliated edge server is: User equipment The first The upload latency from the inference task to its associated edge server is: User equipment Uninstall the The energy consumption of reasoning tasks is: If the corresponding AI model is not deployed on the auxiliary edge server, the inference task needs to be further offloaded to another edge server where the corresponding AI model is deployed. Define the edge server. and edge servers The total transmission rate between them is Introducing variables Represents edge server and Allocation between transmission user equipment The Transmission rate for inference tasks. User equipment. The first The transmission latency of reasoning tasks between edge servers is: The corresponding transmission power consumption is: in For edge servers The transmission power. The computation latency of the inference task on the target edge server is: The corresponding calculated energy consumption is: in, To calculate the energy factor. (2) Construct a training task processing latency and energy consumption model. When the inference accuracy of an AI model deployed on an edge server consistently falls below a preset threshold, a corresponding model training task is triggered. Edge server The first The number of training tasks per class is represented as ,in . No. The training task consists of tuples Depiction, among which Indicates the mini-batch size. Indicates the number of training rounds. This indicates the computational load for each batch of data within a single round. This represents the set of user devices that provide training data. The training task processing includes five stages: model parameters are transferred from the initiating edge server to the target edge server; training data is uploaded from the user device to its affiliated edge server; training data is transferred from the affiliated edge server to the target edge server; the target edge server performs training calculations; and the updated model parameters are synchronized to all edge servers that have deployed the AI model. The model parameter transmission delay is: in yes The data size of hyperparameters in AI-like models The transmission rate allocated to the training data. The corresponding transmission power consumption is: The training data upload latency is: The corresponding upload power consumption is: in For user equipment The uploaded number The size of the training data. The training data transmission latency is: The corresponding transmission power consumption is: The training computation latency is: in For the first Total data size for the training tasks For edge servers Computational resources allocated to the training task. The corresponding calculated energy consumption is: The model parameter synchronization delay is: The corresponding synchronization energy consumption is: (3) Construct an optimization problem. Total system latency This includes the latency of each stage of the inference task and the sum of the latency of each stage of the training task. Total system energy consumption. This includes the sum of user equipment power consumption and edge server power consumption. They are expressed as follows: The total system overhead is defined as the weighted sum of total latency and total energy consumption: ,in and The weighting coefficients for latency and energy consumption are respectively, satisfying... . Combining latency and energy consumption models, this paper addresses the optimization problems of joint AI model placement, inference and training task scheduling, and resource allocation in edge intelligent networks. : in, This indicates where the AI model places the variable set. This represents the scheduling variable for the inference task. Represents the training task scheduling variable. Represents the set of user equipment transmit power allocation variables. This represents the set of variables for allocating transmission rates between edge servers. This represents the set of variables for allocating computing resources to the edge server.
4. The joint optimization method for AI model placement, inference and training task scheduling and resource allocation in an edge intelligent network according to claim 1, characterized in that, Step S3 includes: (1) Optimization problem The problem is decomposed into three numerical optimization subproblems and one deep reinforcement learning subproblem, which are the user equipment transmit power allocation subproblems. Edge server transmission rate allocation sub-problem The sub-problem of edge server computing resource allocation And the sub-problem of AI model placement and inference training task scheduling. . (2) Solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm. The user equipment transmit power variable... Decouple from the original problem and construct a subproblem of transmit power allocation. This subproblem has a linear numerator and a convex denominator, and the constraints are linear. Therefore, the Dinkelbach fractional programming algorithm can be used to efficiently solve the optimal user equipment transmit power allocation decision. . (3) Solve the optimal edge server transmission rate allocation decision based on the Karush-Kuhn-Tucker conditions. Place the decision in the given AI model. Reasoning task scheduling decision Training task scheduling decision Under the condition of allocating variables to the transmission rate and Decouple from the original problem and construct a transmission rate allocation subproblem. This subproblem is a convex optimization problem, which can be solved using the KKT conditions to determine the optimal edge server transmission rate allocation decision. . (4) Solve the optimal edge server computing resource allocation decision based on the convex optimization method. Given the relevant decision variables, the computing resource allocation variables are... and Decouple from the original problem and construct a subproblem of computing resource allocation. The objective function of this subproblem can be expressed as: and The convex function of the sum of terms can be used to solve the optimal edge server computing resource allocation decision using a convex optimization framework (such as CVXPY). .
5. The joint optimization method for AI model placement, inference and training task scheduling, and resource allocation in an edge intelligent network according to claim 1, characterized in that, Step S4 includes: (1) Construct a dual-timescale Markov decision process for AI model placement and inference training task scheduling. Define the long update period as... Each long update cycle is composed of It consists of short update cycles. An adaptive two-layer deep reinforcement learning architecture is constructed, with the upper policy network responsible for AI model placement decisions and the lower policy network responsible for inference and training task scheduling decisions. The upper policy network adaptively determines whether to update the AI model placement policy and its update frequency, while the lower policy network updates task scheduling decisions in each time slot. (2) The Markov decision process placed in the upper-level AI model is defined as follows: Upper-level actions It consists of two parts: an update flag And AI model placement decision .when If the AI model's placement decision is updated, it is updated; otherwise, the placement decision from the previous time slot remains unchanged. upper state This includes: the number of inference tasks generated by each user device. The number of training tasks generated by each edge server AI model placement strategy update interval And the AI model placement decision from the previous time slot. . Upper-level rewards Defined as the negative value of the total system overhead: . Introducing penalty items To reduce the update frequency of AI model placement strategies: ,in This represents the penalty coefficient. The larger the update interval, the smaller the penalty; the greater the difference between the old and new placement strategies, the greater the penalty. (3) The Markov decision process for lower-level inference and training task scheduling is defined as follows: Lower-level actions Includes: all inference task scheduling decisions and all training task scheduling decisions The lower-level task scheduling decision is updated in each time slot. Lower state Including: Current AI model placement strategy The number of inference tasks generated by each user device The number of training tasks generated by each edge server Channel conditions Total transmission rate between edge servers and total computing resources of edge servers . Lower-level rewards Defined as the negative value of the total system overhead: .
6. The joint optimization method for AI model placement, inference and training task scheduling and resource allocation in an edge intelligent network according to claim 1, characterized in that, Step S5 includes: (1) Design an adaptive two-layer deep reinforcement learning algorithm based on TD3 (Twin Delayed Deep Deterministic Policy Gradient) and construct six neural networks for the upper policy network: main Actor network Target Actor Network Two main Critic networks and Two target Critic networks ¹and Six neural networks that construct the lower-level policy network: the main Actor network. Target Actor Network Two main Critic networks and Two target Critic networks and . (2) Train an adaptive two-layer deep reinforcement learning algorithm. In each time slot, the upper-layer policy network adjusts the algorithm according to the current environmental state. Output Action The lower-level policy network determines whether to update the AI model placement strategy and make new placement decisions based on the current environmental state. (Including the current AI model placement decision) Output action The scheduling decisions for inference and training tasks are determined. The AI model placement and task scheduling decisions are input into a numerical optimization subroutine, which solves for optimal power allocation, transmission rate allocation, and computational resource allocation based on the Dinkelbach algorithm, KKT conditions, and convex optimization methods, respectively. After the environment executes all decisions, the total system overhead is calculated, rewards and penalties are generated, and experience tuples are constructed and stored in the upper-level experience replay pool. and lower-level experience replay pool Once the experience replay pool is full, a batch of experiences is randomly sampled from it. The Actor network parameters are updated using the policy gradient method, the Critic network parameters are updated by minimizing the temporal difference error loss function, and the target network parameters are updated using a soft update algorithm. The specific steps of the training process are as follows: Step 1: Randomly initialize the neural network parameters of the upper and lower policy networks, and initialize time slots. Set the interval to 1, and initialize the update interval. =1; Step 2: Determine the output of the upper-layer policy network If the value is greater than 0.5, update the AI model's placement decision based on the output of the upper-level policy network. and reset If it is 1; otherwise, keep the AI model placement decision the same as the previous time slot, and at the same time... ; Step 3: The lower-level policy network determines the state based on... And the current AI model placement decision Output inference task scheduling decision Training task scheduling decision ; Step 4: Solve the optimal user equipment transmit power allocation decision based on the Dinkelbach algorithm. ; Step 5: Solve for the optimal edge server transmission rate allocation decision based on KKT conditions ; Step 6: Solve the optimal edge server computing resource allocation decision based on convex optimization methods ; Step 7: Calculate the total system overhead based on the results of steps 3 to 6. Then calculate the lower-level rewards. ; Step 8: Transfer the lower-level experience tuples Store in the lower-level experience replay pool ; Step 9: Transfer the upper-level experience tuples Store in the upper-level experience replay pool ; Step 10: Replay experience from the lower-level experience pool Randomly sample batches of experience and update the parameters of the lower-level Critic network. and ; Step 11: Replay experience from the upper-level experience pool Based on random sampling batch experience, update the parameters of the upper-layer Critic network. and ; Step 12: Determine the time slot Is it Multiples of, and so on, from, respectively and Based on the experience of sampling batches, update the parameters of the lower and upper layer Actor networks. and Update all target network parameters according to the soft update algorithm. ; Step 13: Let The value is incremented by 1 to determine if the current episode has ended: if so, the system state is reset and a new episode begins; Step 14: Determine if the training process has ended: If yes, end the training process; otherwise, proceed to step 2.