Mobile edge network cooperative reasoning method based on DRL and BP

By adopting a collaborative inference method based on DRL and BP in the mobile edge network, coordinating the resource use of edge servers and terminal devices, the computing complexity and resource utilization problems of LLM inference in the mobile edge network are solved, and efficient and secure inference services are achieved.

CN120197693APending Publication Date: 2025-06-24UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145742.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In mobile edge networks, the computational complexity of LLM inference leads to long response delays, high bandwidth costs, and the risk of private data leakage. The prior art fails to fully utilize the computing resources of terminal devices and edge servers for efficient inference.

Method used

Using a mobile edge network collaborative inference method based on DRL and BP, a synergistic inference framework for large language models and small models is established, the accuracy of long-term inference tasks is optimized and the completion delay constraints of tasks are ensured, and the resource usage of edge servers and terminal devices is coordinated through DRL and BP algorithms.

Benefits of technology

It realizes efficient coordination of edge servers and terminal devices under resource constraints, completes time-varying inference tasks, improves the inference performance of LLM, protects user data privacy, and improves resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197693A_ABST
    Figure CN120197693A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer artificial intelligence, in particular to a mobile edge network collaborative reasoning method based on DRL and BP, and provides an edge-terminal collaborative reasoning scheme in a large language model, which fully utilizes the computing power of terminal equipment and an edge server, protects the data privacy of a user to a certain extent, and improves the user experience. Under the condition of limited resources, the problems of reasoning task unloading and model caching are solved by utilizing a'learning-to-learning 'normal form in a combined manner. A lightweight and high-popularity model is selected and cached to an edge server by using deep reinforcement learning, and a related reasoning task unloading problem is solved by adopting a distributed belief propagation technology. A numerical result proves the effectiveness and superiority of the proposed scheme. By means of edge terminal cooperation, the communication and computing resource utilization rate in LLM reasoning can be improved. In other words, diversified reasoning requests can be effectively completed through effective small model caching and reasoning task offloading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer artificial intelligence, and specifically relates to a collaborative inference method for mobile edge networks based on DRL and BP. Background Art

[0002] Large language models (LLMs) have advanced content creation and inference capabilities, which can provide immersive intelligent services for users in mobile edge networks. However, due to the high computational complexity of LLM inference, it is usually provided as a cloud service, which results in long response latency, high bandwidth costs, and the risk of private data leakage. At the same time, mobile edge computing (MEC), as an efficient paradigm for supporting various distributed applications, has gradually emerged, and its application scenarios cover many aspects from content delivery to artificial intelligence (AI) applications. Therefore, it is a natural choice to use MEC as a supplementary mode for deploying LLM services to provide services for mobile users. Although this approach provides flexibility, how to utilize the computing and memory resources of edge servers and terminal devices to provide efficient inference services remains a challenging problem.

[0003] Edge computing processes information locally instead of transmitting sensitive data to a centralized cloud. In addition, edge computing allows for a flexible and distributed architecture, which can allocate appropriate computing resources to specific tasks. Therefore, the combination of edge computing and large language models (LLMs) endows LLMs with personalized and domain-specific generation capabilities. However, the high computational requirements of LLM inference often exceed the processing capabilities of resource-limited user equipment (UEs) in wireless access networks, making it difficult to provide LLM inference services for UEs. Some researchers have proposed deploying LLMs on edge servers, but this fails to fully utilize the computing resources of terminal devices and may lead to inefficiency. To jointly utilize the computing capabilities of terminal devices and edge servers for LLM inference, some researchers have considered using split learning in LLMs. Although the inference frameworks in the existing literature support collaborative LLM inference on distributed edge devices and servers through model splitting, split learning requires a large amount of data exchange between the server and the terminal device during the inference process. This communication overhead has become a major burden for providing diverse and unpredictable AI inference services in mobile access networks. Summary of the Invention

[0004] In view of the above problems, the present invention provides a collaborative inference method for mobile edge networks based on DRL and BP.

[0005] The technical solution adopted is a collaborative inference method for mobile edge networks based on DRL and BP, including the following steps:

[0006] S1. Establish an inference framework for the collaborative operation of large language models and small models in a mobile edge network computing system;

[0007] S2. Establish a communication resource allocation model and an inference task offloading model;

[0008] S3. Under the constraints of limited communication resources and computing resources, optimize the accuracy of long-term inference tasks and ensure the task completion delay constraint, and establish a joint optimization problem for inference task offloading and model caching;

[0009] S4. Split the established joint optimization problem according to the time scale, solve it iteratively, and transform the model caching problem on the large time scale into an MDP model and solve it using DRL;

[0010] S5. Transform the inference task offloading problem on the small time scale into an unconstrained optimization problem and solve it using BP.

[0011] Optionally, in S1, it includes a base station, an edge server co-located with the base station, and a set of N terminal devices with processing capabilities;

[0012] The set of terminal devices is denoted as N = {1, 2,..., N}, and a small model is deployed on each terminal device in the set. The small model is independently constructed based on the local data of each terminal device and the processing capabilities of each terminal device. The set of small models is denoted as M = {m1, m2,..., m N}, and the number of parameters of the model deployed on terminal device i is |m i |;

[0013] A large language model is deployed on the edge server, and inference task offloading in the large language model is carried out in time slots, where each time slot t has the same duration t s .

[0014] Optionally, in S2, the time interval for caching in the large language model:

[0015] T s = T * t s ;

[0016] where T represents the number of time slots in a communication round;

[0017] Let represent the r-th communication round of model caching.

[0018] Optionally, in S3, the model caching problem is described as defining the environmental state within the communication round r as s r , s r ∈ S;

[0019] where

[0020] where N r,i represents the number of times the small model deployed on the access terminal device i is accessed within the communication round r;

[0021] N r,f represents the number of requests for the task type f within the communication round r;

[0022] represents the small model caching decision in the communication round r - 1.

[0023] Optionally, in S4, first define G as the available GPU memory of the edge server for caching the small model, then the model caching decision variable needs to satisfy the following constraints:

[0024]

[0025] Then define the action selected by the edge server in the communication round r as a r , a r ∈A, where

[0026] Then let the state transition probability be representing the probability that the environmental state transfers from s r to s r through the action a r+1 ;

[0027] Next, set the reward and define the reward function as:

[0028]

[0029] Then define the long - term reward as the sum of the current reward and the future discounted rewards, expressed as where γ ∈ [0, 1] represents the discount factor, which is used to reflect the discount effect of future rewards on the current decision;

[0030] Then define the expected value of the reward given the current state and action as the state - value function:

[0031] Q(s r , a r ) = E[R r |s r , a r (Formula 2).

[0032] Optionally, in S4, explore the action space by maximizing the value of the Q function, and select the optimal action in the current state. The objective function is replaced with the following form:

[0033]

[0034] Among them, θ k and respectively represent the weight vectors of the estimated Q-network and the target Q-network in communication round k;

[0035] The loss function is expressed as:

[0036] L(θ k ) = E[(y k - Q(s k , a k ; θ k )) 2 (Formula 4);

[0037] During the training process, θ k updates the parameters by performing gradient descent on the loss function, and its update process can be expressed as:

[0038]

[0039] Optionally, in S4, it includes the following steps:

[0040] Initialize the experience replay pool D;

[0041] Initialize the Q-network θ0 and the target Q-network

[0042] FOR Episode = 1 to Maxepisode;

[0043] Initialize the environmental state;

[0044] FOR communication round k = 1, 2,..., R;

[0045] Observe the current environmental state s k ;

[0046] Select a random action ak with probability òk, otherwise select a k = argmax a Q(·);

[0047] Execute the action a k , obtain the reward r k and the environmental state s k+1 ;

[0048] Store (s k , a k , r k , s k+1 ) into the experience replay pool D;

[0049] Randomly select a mini-batch sample of size N from the experience replay pool D (s i , a i , ri ,s i );

[0050] Calculate according to formula (3)

[0051] For the loss function Execute the gradient descent step;

[0052] Update the parameters

[0053] Reset the target Q network every C time slots

[0054] ENDFOR.

[0055] Optionally, in S5, first define the inference task offloading decision of terminal device i at time slot t as where j ∈ L(i), L(i) = {0, i} ∪ N(i);

[0056] When j = 0, is a decision variable indicating whether terminal device i chooses to offload the task to the edge server at time slot t;

[0057] When j = i, is a decision variable indicating whether terminal device i chooses to locally process the task at time slot t;

[0058] When j ∈ N(i), is a decision variable indicating whether terminal device i chooses to offload the task to neighboring terminal device j at time slot t;

[0059] The inference accuracy of the task requested by terminal device i at time slot t can be expressed as

[0060] where represents the inference accuracy when the task of terminal device i is offloaded to the edge server, neighboring terminal device, or locally processed at time slot t.

[0061] Optionally, in S5, transform the optimization problem into the following form:

[0062]

[0063] At the same time, define four functions of the decision variable X:

[0064]

[0065] And make the optimization problem equivalent to the following formula:

[0066]

[0067] Among them,

[0068]

[0069] Among them, η’ i (X) is used to measure the inference accuracy of the requested task;

[0070] Indicator function and impose strict constraints on task offloading and latency requirements;

[0071] Let G1(X) = g 1 (X),

[0072] Obtain:

[0073]

[0074] The beneficial effects of the present invention are as follows;

[0075] 1. A proposed edge-terminal collaborative inference scheme in large language models, which makes full use of the computing capabilities of terminal devices and edge servers, and at the same time protects the data privacy of users to a certain extent;

[0076] 2. A proposed mechanism to upload lightweight models with high popularity to the edge server as plugins to provide specific knowledge for the LLM, which effectively improves the inference performance of the LLM;

[0077] 3. To effectively coordinate the edge server and terminal devices under resource-constrained conditions, complete time-varying inference tasks, and solve the inference task offloading and model caching problems in ETCI. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a schematic diagram of the ETCI framework;

[0079] Figure 2 It is a schematic diagram of the inference process of the LLM-assisted task;

[0080] Figure 3 It is a hierarchical algorithm flowchart for inference task offloading and model caching;

[0081] Figure 4 It is a network topology diagram of one edge server and three users;

[0082] Figure 5 It is a factor graph model of one edge server and 3 users;

[0083] Figure 6 It is a convergence performance graph of the ETCI algorithm;

[0084] Figure 7 Schematic diagram of cumulative rewards for ETCI and CINL;

[0085] Figure 8 Schematic diagram of the inference accuracy of different task types and algorithms;

[0086] Figure 9 Schematic diagram of the influence of transmission power on the average inference accuracy (AIA);

[0087] Figure 10 Schematic diagram of the influence of system bandwidth on the average inference accuracy (AIA);

[0088] Figure 11 Schematic diagram of the relationship between the average inference accuracy and ; Detailed implementation manners

[0089] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0090] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0091] In this embodiment, as Figure 1 and Figure 3 shown, a mobile edge network collaborative inference method based on DRL and BP is provided, including the following steps:

[0092] S1. Establish an inference framework that coordinates a large language model and a small model for the mobile edge network computing system;

[0093] S2. Establish a communication resource allocation model and an inference task offloading model;

[0094] S3. Under the condition of limited communication resources and computing resources, optimize the accuracy of long-term inference tasks and ensure the task completion delay constraint, and establish a joint optimization problem of inference task offloading and model caching;

[0095] S4. Split the established joint optimization problem according to the time scale, solve it iteratively, and transform the model caching problem with a large time scale into an MDP model, which is solved by DRL;

[0096] S5. Transform the inference task offloading problem with a small time scale into an unconstrained optimization problem, which is solved by BP.

[0097] The purpose of such a design is to propose an edge-terminal collaborative inference scheme in a large language model. This scheme makes full use of the computing capabilities of terminal devices and edge servers, and at the same time protects the data privacy of users to a certain extent.

[0098] At the same time, it should be noted that the DRL mentioned in this application refers to deep reinforcement learning, and its English is Deep Reinforcement Learning. The BP mentioned in this application refers to the belief propagation algorithm, and its English is Belief Propagation.

[0099] At the same time, in the specific implementation, in S1, it includes a base station, an edge server co-located with the base station, and a set of N terminal devices with processing capabilities;

[0100] The set of terminal devices is denoted as N = {1, 2,..., N}, and a small model is deployed on each terminal device in the set. The small model is independently constructed based on the local data and processing capabilities of each terminal device. The set of small models is denoted as M = {m1, m2,..., m N}, and the number of parameters of the model deployed on terminal device i is |m i |;

[0101] A large language model is deployed on the edge server, and the inference task offloading in the large language model is carried out in time slots, where each time slot t has the same duration t s .

[0102] In S2, the time interval for caching in the large language model:

[0103] T s = T * t s ;

[0104] where T represents the number of time slots in a communication round;

[0105] Let represent the communication round of the r-th model cache.

[0106] It should be noted that within the communication round r, the server determines the caching strategy of the small model. Define a binary variable to indicate whether the small model deployed on the terminal device i is uploaded and cached to the server within the communication round r: if indicates caching, otherwise Normally, a small model for processing text has millions of parameters, which will obviously occupy a certain amount of Graphics Processing Unit (GPU) memory space. Let G denote the available GPU memory of the edge server for caching small models. Then, the model caching decision variable needs to satisfy the following constraints:

[0107]

[0108] To better utilize the computing resources of the edge server to complete the AI inference tasks offloaded to the edge server, the small models cached in the edge server are combined with the Large Language Model (LLM) so that they can work together to improve the performance of the AI inference tasks. The small models act as plugins, providing knowledge and predictions for specific tasks, while the large language model focuses on general language understanding and uses its powerful generalization ability to expand the knowledge not available in the small models.

[0109] To enable the Large Language Model (LLM) to better understand the specific task knowledge provided by the small models, a context prompt (in-context prompt) is constructed for the LLM. This prompt consists of a set of examples randomly drawn from the public dataset of the edge server, as well as the prediction results of their corresponding small plugin models. The prediction results include the predicted labels and their associated confidence scores. An example of the context prompt is shown in Table 1:

[0110] TABLEI

[0111] AN EXAMPLE OF IN-CONTEXT PROMPT ANDINFERENCE PROCEDURE

[0112]

[0113] In this way, by combining the labels predicted by the small plugin models, the LLM can better understand the relationship between the input examples, the true labels, and the expertise of the plugin models. This will help the LLM make the final prediction. In addition, the confidence scores provide a measure of the uncertainty of the plugin models in the prediction. By incorporating these scores into the context, the LLM can place more trust in highly confident predictions while remaining cautious about uncertain predictions. In addition, the confidence scores can also guide the LLM to focus on more challenging context examples, enabling it to learn from these difficult cases and potentially improve the overall performance.

[0114] Such as Figure 3As shown in the figure, in order to iteratively solve the problems of inference task offloading and model caching, a hierarchical algorithm based on deep reinforcement learning (DRL) and belief propagation (BP) is proposed in this embodiment. In each time slot, the user will make an inference task offloading decision. Since there is no central controller in the considered scenario, the information at the user side mainly depends on local collection and local interaction within the domain. Therefore, a distributed algorithm is more suitable for solving the above problems. In addition, in order to reduce the computational complexity of the proposed hierarchical algorithm and ensure the convergence speed of the algorithm, a low-complexity distributed cooperative task offloading algorithm based on belief propagation (BP) is utilized. In each communication round, the edge server will cache those small models that are frequently accessed but not yet cached. During the task offloading process, since small model caching is a dynamic sequential decision-making problem, this problem is modeled as a Markov decision process (MDP) and solved using a DRL algorithm.

[0115] Since the caching decision of small models is a dynamic sequential decision-making problem, this problem is modeled as a Markov decision process (MDP). The MDP consists of a quadruple <S, A, P, R>, where S represents the environmental state, A represents the action space, P represents the transition probability between states, and R represents the reward value function. The edge server needs to formulate a caching policy for all small models deployed on users in each communication round, and then, based on the model caching state in the previous communication round, determine which cached small models need to be deleted from the base station side and which small models need to be uploaded and cached. After the edge server completes the model caching decision, the state of the model caching will transition to another state, and a reward will be obtained according to the inference accuracy of the tasks in this communication round.

[0116] The detailed description of the MDP process is as follows:

[0117] State: Define the environmental state in communication round r as s r , s r ∈S, where where, N r,i represents the number of times the small model deployed on user i is accessed in communication round r, N r,f represents the number of requests for task type f in communication round r, represents the small model caching decision in communication round r - 1.

[0118] Action: Define the action selected by the edge server in communication round r as a r , a r ∈A, where

[0119] Transition Probability: Let the transition probability be Denote through action a r , the probability that the environmental state transfers from s r to s r+1 .

[0120] Reward: To maximize the long-term inference accuracy while satisfying the GPU memory constraint of the edge server, the reward function is defined as:

[0121]

[0122] Since the sub-problem is transformed into a general Markov decision process (MDP) problem with a finite state space and action space, a value-based reinforcement learning algorithm is adopted to ensure the stable convergence of the cumulative expected reward. However, the diversity of user-requested tasks leads to a very large state space. In this case, maintaining a Q-table becomes impractical because it cannot store all state-action pairs. Therefore, a deep reinforcement learning (DRL) algorithm is considered to solve this MDP problem. Given that the MDP problem has a discrete action space, the double deep Q-network (DDQN) method is adopted to solve it.

[0123] In a communication round, the edge server will select an action according to the current environmental state. When the edge server takes an action, the environmental state will transfer to another state, and the edge server will obtain the corresponding reward. In this way, the edge server continuously interacts with the environment, and the goal is to learn a policy that enables it to make decisions in any observed state and maximize the long-term cumulative reward during the learning process.

[0124] Define the long-term reward as the sum of the current reward and the future discounted reward, expressed as where γ ∈ [0, 1] represents the discount factor, reflecting the discount effect of future rewards on the current decision. Define the expected reward value given the current state and action as the state-value function:

[0125] Q(s r , a r ) = E[R r |s r , a r (Formula 2)

[0126] DDQN explores the action space by maximizing the value of the Q function or adopting an ε-greedy strategy, and selects the optimal action in the current state. In DDQN, the value of the Q function can be estimated by a neural network. The update method of DDQN is the same as that of DQN, but its objective function is replaced by the following form:

[0127]

[0128] Among them, θ k and respectively represent the weight vectors of the estimated Q-network and the target Q-network in communication round k. The goal of DDQN is to minimize the gap between the estimated Q-value and the target Q-value. Therefore, the loss function of DDQN can be expressed as

[0129] L(θ k ) = E[(y k - Q(s k , a k ; θ k )) 2 (Formula 4)

[0130] During the training process, θ k updates the parameters by performing gradient descent on the loss function, and its update process can be expressed as:

[0131]

[0132] Therefore, in S4, the following steps are included:

[0133] Initialize the experience replay pool D;

[0134] Initialize the Q-network θ0 and the target Q-network

[0135] FOR Episode = 1 to Maxepisode;

[0136] Initialize the environmental state;

[0137] FOR communication round k = 1, 2,..., R;

[0138] Observe the current environmental state s k ;

[0139] Select a random action ak with probability òk, otherwise select a k = argmax a Q(·);

[0140] Execute the action a k , obtain the reward r k and the environmental state s k+1 ;

[0141] Store (s k , a k , r k , s k+1 ) into the experience replay pool D;

[0142] Randomly select a mini-batch sample of size N from the experience replay pool D (si , a i , r i , s i );

[0143] Calculate according to formula (3)

[0144] For the loss function Perform a gradient descent step;

[0145] Update the parameters

[0146] Reset the target Q-network every C time slots

[0147] ENDFOR.

[0148] Meanwhile, in this embodiment, in order to reduce the computational complexity of the hierarchical algorithm, it is crucial to provide a low-complexity algorithm for sub-problems within a small time scale. The belief propagation (BP) algorithm is commonly used in iterative processes such as LDPC code decoding in wireless communication. It is good at iteratively propagating probability information between local nodes to approximate the global optimal solution. This characteristic highly coincides with the characteristics of the inference task offloading problem because the decision of this problem is based on local interaction information. Therefore, a low-complexity distributed cooperative task offloading algorithm based on belief propagation is proposed.

[0149] First, define the inference task offloading decision of terminal device i at time slot t as where j ∈ L(i), and L(i) = {0, i} ∪ N(i);

[0150] When j = 0, is a decision variable indicating whether terminal device i chooses to offload the task to the edge server at time slot t;

[0151] When j = i, is a decision variable indicating whether terminal device i chooses to locally process the task at time slot t;

[0152] When j ∈ N(i), is a decision variable indicating whether terminal device i chooses to offload the task to neighboring terminal device j at time slot t;

[0153] The inference accuracy of the task requested by terminal device i at time slot t can be expressed as

[0154] where represents the inference accuracy when the task of terminal device i is offloaded to the edge server, neighboring terminal device, or locally processed at time slot t.

[0155] In S5, the optimization problem is transformed into the following form:

[0156]

[0157] Meanwhile, four functions of the decision variable X are defined:

[0158]

[0159]

[0160] And the optimization problem is equivalent to the following formula:

[0161]

[0162] Where,

[0163]

[0164] Where, η’ i (X) is used to measure the inference accuracy of the requested task;

[0165] The indicator function and impose strict constraints on task offloading and latency requirements;

[0166] Let G1(X) = g 1 (X),

[0167] We get:

[0168]

[0169] As Figure 5 and Figure 6 shown, a factor graph model can be proposed for Formula 12, introducing a variable node μ for each variable n , and introducing a function node F i for each function η m . The mapping rule from to μ n is expressed as:

[0170]

[0171] Where, ξ(i,j) represents the index of neighbor user i in the set L(i). The mapping rule from η’ i (X), G1(X) and G2(X) to the function node F m is expressed as:

[0172]

[0173] The goal is to design a message passing process that enables the gradual approximation of the optimal solution to sub-problem (12). Let denote the message passed from variable node μ n to function node F m in the l-th iteration. denotes the message passed from function node F m to variable node μ n . Since is a binary variable, in fact, only the scalar ratio of the messages passed between each pair of nodes is needed. These message ratios can also be represented in the logarithmic domain as follows:

[0174]

[0175] In this way, the computational complexity and communication overhead are significantly reduced because only half of the messages are calculated and passed.

[0176] Finally, the update formula for message is as follows:

[0177]

[0178] When F m = η’ i , the update formula for message is as follows:

[0179]

[0180] where q n denotes the inference accuracy related to variable μ n .

[0181] When F m = G1, the update formula for message is as follows:

[0182]

[0183] where

[0184] When F m = G2, the update formula for message is:

[0185]

[0186] where

[0187] In practice, messages and reflects the belief in the value of the variable μ n and should be updated according to Formulas 17 to 20. In the l-th iteration, the belief that μ n = x is expressed as

[0188]

[0189] This expression represents the product of all messages passed to μ n . Therefore, in the log domain, the belief ratio can be obtained as follows:

[0190]

[0191] where, is given by Theorem 1. Therefore, the estimated value of μ n can be expressed as

[0192]

[0193] In each iteration, each variable node μ n updates its belief in the associated variable node according to Formula 22 and estimates according to Formula 23 until convergence.

[0194] Therefore, the complete steps for the entire S5 are as follows:

[0195] Map η i,j and g i to F m ; Map x i,j to μ n ;

[0196] Set l = 0 and

[0197] Set l max to a sufficiently large constant;

[0198] WHILE: not converged and l ≤ l max};

[0199] FOR i = 1:N;

[0200] Calculate the message

[0201] for F m = η’ i according to Formula 18, and calculate the message

[0202] ENDFOR;

[0203] For Fm = G1, calculate the message according to Equation 19

[0204] For F m = G2, calculate the message according to Equation 20

[0205] Calculate the belief according to Equation 22 Calculate according to Equation (\ref{y})

[0206] Estimate each variable according to Equation 23

[0207] Check for convergence and set l = l + 1;

[0208] END WHILE;

[0209] Obtain the optimal solution of Equation 12

[0210] In this embodiment, a simulation experiment is conducted to verify the convergence of the solution provided by the present application and evaluate its performance in terms of average inference accuracy (AIA).

[0211] Among them, the following three schemes are used as comparative references:

[0212] 1. Cooperative Edge-Terminal Inference Scheme Based on Belief Propagation and Deep Reinforcement Learning (without using LLM assistance) (CINL): Each user offloads its inference task according to the belief propagation algorithm in a time slot. In each communication round, the DDQN agent selects a small model to cache and caches it on the server.

[0213] 2. LLM-Assisted Cooperative Edge-Terminal Inference Scheme Based on Belief Propagation and Random Model Caching (CIRC): Each user offloads its inference task according to the belief propagation algorithm in a time slot. In each communication round, the edge server randomly selects a small model and caches it on the server.

[0214] 3. LLM-Assisted Inference Scheme Based on Belief Propagation without Edge-Terminal Cooperation (ILBP): Each user offloads its inference task according to the belief propagation algorithm in a time slot.

[0215] The simulation settings are as follows:

[0216] Consider a MEC system that includes an edge server co-deployed with a base station and 3 user equipments (UEs) in 3 wireless communication scenarios, with a coverage area of 100m × 100m. The users are randomly distributed in this area, and the network topology is as Figure 4As shown, the noise power spectral density is assumed to be -174 dBm / Hz. The cellular channel model is assumed to be 128.1 + 37.6 log 10 (d [km]), while the device-to-device (D2D) link channel model is assumed to be 148 + 40 log 10 (d [km]).

[0217] In the inference framework, it is assumed that the GPT3.5-turbo model is deployed on the edge server, and the simulation accesses GPT3.5-turbo only through its application programming interface (API). Additionally, the Bert model is deployed on the user side. Since the Bert model is small, the user side can perform fine-tuning. Considering the different computing capabilities of users, the Bert-large-cased model is deployed on User 1 and User 2 respectively, and the Bert-base-cased model is deployed on User 3. The Bert model on the user side is fine-tuned through the GLUE dataset. Specifically, the Bert model on User 1 is fine-tuned through the SST-2 dataset in GLUE, while the Bert models on User 2 and User 3 are fine-tuned through the COLA dataset in GLUE. To demonstrate the performance of the inference scheme, inference tasks are introduced by splitting the COLA, SST-2, and QNLI datasets in GLUE into multiple data packets.

[0218] In the algorithm, the DDQN network contains two hidden layers with 32 and 64 neurons respectively. A total of 250 training episodes are conducted, and each training episode contains 8 training cycles. In each cycle, the agent interacts with the environment once and trains a mini-batch of data. Unless otherwise specified, the parameters used in the communication scenario and the DRL algorithm are listed in Tables 3 and 4 respectively.

[0219] TABLE III

[0220] PARAMETERS OF THE COMMUNICATION SCENARIO

[0221]

[0222] TABLEIV

[0223] PARAMETERS OF THE DRL ALGORITHM

[0224]

[0225] As Figure 6 shown, the average loss of ETCI starts to converge after approximately 800 training episodes. After 1500 episodes, the average loss of ETCI approaches 0, which means that the ETCI scheme converges well.

[0226] Next, the cumulative rewards of ETCI and CINL were compared. As Figure 7 shown, the cumulative reward of ETCI was always higher than that of the CINL algorithm. When the number of rounds was greater than 175, the cumulative reward of ETCI approached 19.5, which was 1.8 higher than that of the CINL algorithm. Therefore, it can be seen that significant performance improvement was achieved after introducing the large language model (LLM) into the proposed inference scheme.

[0227] The average inference accuracy (AIA) of three task types under four algorithms (ETCI, CIRC, ILBP, and CINL) was compared. The transmission power was set to 23 dBm and the system bandwidth was set to 15 MHz. As Figure 8 shown, ETCI was always superior to the other three algorithms in the three inference tasks. In the CINL algorithm, the inference accuracies of the COLA and QNLI tasks were significantly lower than those of the other three algorithms. Therefore, it can be seen that the large language model (LLM) brought a significant improvement in the inference accuracies of these two tasks.

[0228] In the next experiment, the relationship between the transmission rate and the inference accuracy among the four algorithms was examined by changing the transmission power and bandwidth. First, the relationship between the transmission power and the AIA was investigated. Figure 9 The relationship between the user transmission power and the AIA is shown. It can be seen that the AIA achieved by ETCI was always higher than that of CINL, ILBP, and CIRC. When the user's transmission power increased, the AIA also showed an upward trend. Especially when the transmission power was 26 dBm, the AIA achieved by ETCI could reach 0.819, which was 0.04 higher than that of CIRC, 0.069 higher than that of ILBP, and 0.15 higher than that of CINL.

[0229] Next, the relationship between the system bandwidth and the AIA was compared under the condition that the transmission power was fixed at 23 dBm. As Figure 10 shown, the AIA achieved by the four algorithms increased monotonically with the increase of the system bandwidth. When the system bandwidth was greater than 5 MHz, the AIA of various tasks achieved by ETCI was higher than that of CINL, ILBP, and CIRC, demonstrating its ability to better utilize higher transmission rates to improve the inference accuracy. Especially, the advantage of ETCI was more significant at higher bandwidths.

[0230] Finally, the AIA of the four algorithms under different bandwidth allocation ratios was compared. The bandwidth allocation ratio was used to measure the bandwidth allocation ratio between two communication modes. The system bandwidth was set to 10 MHz and the transmission power was set to 23 dBm. As Figure 11 shown, when When all system bandwidth is allocated to D2D links, the AIA achieved by ILBP is higher than that of ETCI, CINL, and CIRC. This is because the inference task cannot be offloaded to the server through LLM inference. When is less than 0.8, the AIA achieved by ETCI increases with the increase of. When is greater than 0.4, the AIA achieved by ETCI is always higher than other algorithms.

[0231] According to the simulation verification results provided in this embodiment, to optimize the LLM inference performance in ETCI, the inference task offloading and model caching problems are jointly solved by using the "learning to learn" paradigm under limited resources. Specifically, deep reinforcement learning is used to select lightweight and highly popular models to cache in the edge server, and distributed belief propagation technology is adopted to solve the related inference task offloading problem. The numerical results prove the effectiveness and superiority of the proposed ETCI scheme. At the same time, with the help of edge terminal cooperation, the utilization rate of communication and computing resources in LLM inference can be improved. In other words, through effective small model caching and inference task offloading, diverse inference requests can be effectively completed.

[0232] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A collaborative reasoning method for mobile edge networks based on DRL and BP, characterized in that: The following steps are involved: S1. Establish a reasoning framework for the collaboration of large language models and small models for the mobile edge network computing system; S2. Establish a communication resource allocation model and a reasoning task offloading model; S3. Under the condition of limited communication resources and computing resources, optimize the accuracy of long-term reasoning tasks and ensure the completion delay constraints of tasks, and establish the joint optimization problem of reasoning task offloading and model caching; S4. The established joint optimization problem is split according to the time scale, solved iteratively, and the large time scale model caching problem is converted into an MDP model and solved by DRL; S5. Convert the small-time-scale reasoning task offloading problem into an unconstrained optimization problem and solve it using BP.

2. According to claim 1, a DRL and BP based mobile edge network collaborative reasoning method is characterized in that: S1 includes a base station, an edge server co-located with the base station, and a set of N terminal devices with processing capabilities; The terminal device set is denoted as N = {1, 2, ..., N}, and a small model is deployed on each terminal device in the set. The small model is independently constructed based on the local data of each terminal device and the processing capability of each terminal device. The small model set is denoted as M = {m1, m2, ..., m N }, and the number of parameters of the model deployed on terminal device i is |m i |; A large language model is deployed on the edge server, and the reasoning task offloading in the large language model is performed in time slots, where each time slot t has the same duration t s .

3. According to claim 2, a DRL and BP based mobile edge network collaborative reasoning method is characterized in that: In S2, the time interval for caching in the large language model is: T s =T*t s ; Where T represents the number of time slots in a communication round; make Represents the communication round of the rth model cache.

4. According to claim 3, a DRL and BP-based mobile edge network collaborative reasoning method is characterized in that: In S3, the model cache problem is described as: the environment state within the communication round r is defined as s r ,s r ∈S; in Where N r,i represents the number of visits to the small model deployed on terminal device i within communication round r; N r,f represents the number of requests for task type f within communication round r; Represents the small model caching decision in communication round r-1.

5. According to claim 4, a DRL and BP based mobile edge network collaborative reasoning method is characterized in that: In S4, we first define G to represent the amount of GPU memory available on the edge server for caching small models. Then the model caching decision variable needs to satisfy the following constraints: The action selected by the edge server in communication round r is defined as a r , a r ∈A, where Then let the state transition probability be Indicates that through action a r , the environmental state changes from s r Transfer to r+1 probability; Then set the reward and define the reward function as: The long-term reward is then defined as the sum of the current reward and the future discounted reward, expressed as Among them, γ∈[0,1] represents the discount factor, which is used to reflect the discount effect of future rewards on current decisions; The expected value of reward given the current state and action is then defined as the state-value function: Q(s r ,a r )E[R r |s r ,a r ] (House 2).

6. According to claim 5, a DRL and BP based mobile edge network collaborative reasoning method is characterized in that: In S4, the action space is explored by maximizing the value of the Q function and the optimal action is selected in the current state. The objective function is replaced by the following form: Among them, θ k and They represent the weight vectors of the estimated Q network and the target Q network in communication round k respectively; The loss function is expressed as: L(θ k ) = E[(y k - Q(s k , a k ; θ k )) 2 (Formula 4); During the training process, θ k The parameters are updated by performing gradient descent on the loss function. The update process can be expressed as: θ k+1 = θ k + α[y k - Q(s k , a k , θ k )] · ▽Q(s k , a k , θ k (Equation 5).

7. The DRL and BP-based mobile edge network collaborative reasoning method according to claim 6, characterized in that: S4 includes the following steps: Initialize experience replay pool D; Initialize Q network θ0 and target Q network FOR Episode=1ToMaxEpisode; Initialize the environment state; FOR communication round k = 1, 2, ..., R; Observe the current environment status k ; With probability k Choose a random action ak, otherwise choose a k = argmax a Q(·); Execute action a k , get reward r k and environmental state s k+1 ; Will (s k ,a k ,r k ,s k+1 ) is stored in the experience replay pool D; Randomly select a mini-batch of size N from the experience replay pool D (s i ,a i ,r i ,s i ); According to formula (3) For the loss function Perform a gradient descent step; Update the parameter θ k+1 =θ k -β▽L(θ k ); Reset the target Q network every C time slots ENDFOR.

8. According to claim 4, a DRL and BP based mobile edge network collaborative reasoning method is characterized in that: In S5, the reasoning task offloading decision of terminal device i in time slot t is first defined as Where j∈L(i), L(i)={0,i}∪N(i); When j = 0, is a decision variable, indicating whether terminal device i chooses to offload the task to the edge server in time slot t; When j = i, is a decision variable, indicating whether terminal device i chooses to process tasks locally in time slot t; When j∈N(i), is a decision variable, indicating whether terminal device i chooses to offload the task to the neighboring terminal device j in time slot t; The inference accuracy of the task requested by terminal device i in time slot t can be expressed as in represents the inference accuracy of the task of terminal device i when it is offloaded to the edge server, neighboring terminal devices, or local processing at time slot t.

9. The DRL and BP-based mobile edge network collaborative reasoning method according to claim 8, characterized in that: In S5, the optimization problem is transformed into the following form: At the same time, four functions of the decision variable X are defined: And the optimization problem is equivalent to the following formula: in, Among them, η′ i (X) is used to measure the reasoning accuracy of the requested task; Indicator function and Strict constraints are imposed on task offloading and latency requirements; make get: