Online caching and service system for edge device large model based on multi-arm bandit

By adopting the online caching method of multi-arm gambling machine model in edge devices, selecting suitable LLM caches in real time, solving the problem of optimization of large-scale language model cache strategy in edge devices, and achieving efficient resource utilization and performance improvement.

CN120029782AActive Publication Date: 2025-05-23BEIJING NORMAL UNIV AT ZHUHAI

Patent Information

Application Number
CN202510177649.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-23
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

It is difficult for prior art to effectively optimize caching strategies for large-scale language models (LLMs) in edge devices, especially in environments where resources are limited and service quality is unpredictable.

Method used

Using the online caching method of edge device large model based on multi-arm gambling machine (MAB), the output training modules are used to learn and select suitable LLM caches in real time to optimize resource utilization and inference performance.

Benefits of technology

Dynamic optimization of cache strategy is realized, system response speed and task processing efficiency are improved, static cache strategy is overcome, and the overall performance and resource utilization efficiency of the system are significantly improved through nonlinear performance optimization and efficient feedback mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029782A_ABST
    Figure CN120029782A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of edge calculation and large model reasoning optimization, and provides an edge device large model online caching and service system based on a multi-arm bandit, and the system comprises a feature information receiving module which is used for receiving a feature information set of a cloud alternative LLM; the LLM quality prediction module is used for predicting a utility vector for each LLM by using a neural network based on the feature information set and calculating a UCB value and a utility offset; the edge sorting selection module is used for sorting the sum of the UCB value and the utility offset in a descending order, selecting a preset number of LLMs to generate an LLM subset, caching the LLM subset and deploying the LLM subset to an edge end; the asynchronous processing output training module is used for collecting service quality data and neural network gradient data and asynchronously performing training and parameter optimization updating on a neural network model in the LLM quality prediction module; and continuously repeating the steps to realize online caching and service of the edge device large model based on the multi-arm bandit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing and large model reasoning optimization, and in particular to an edge device large model online caching and service system based on a multi-armed gambling machine. Background Art

[0002] With the rapid development of edge computing technology, the demand for smart devices is increasing, especially in large-scale language model (LLM) reasoning applications, such as smart voice assistants, translation services, text summary generation, etc. In order to meet the needs of these services for low latency and efficient computing, edge devices are usually responsible for processing tasks close to users. Because the traditional full model caching method cannot meet resource constraints, in order to adapt to the storage and computing limitations of edge devices, researchers have proposed a variety of model compression techniques, including quantization, knowledge distillation, and pruning. These techniques aim to reduce the memory footprint and computing requirements of the model while maintaining performance. However, the computing resources and storage space of edge devices are limited, and the service quality corresponding to different models is different, so it is necessary to select a suitable model cache to optimize the service quality.

[0003] (1) In the prior art, many joint optimization methods are used to optimize multiple factors such as cache, computation offloading, and bandwidth allocation. These methods usually rely on static or semi-static models and aim to improve the overall performance of the system, such as reducing latency, reducing energy consumption, or improving cache hit rate. However, joint optimization methods usually involve multiple variables and constraints, resulting in high complexity and difficulty in real-time solution. Moreover, since these methods are usually based on predefined models and parameters, they are difficult to adapt to dynamically changing user requests and environments.

[0004] (2) In recent years, the introduction of deep learning and reinforcement learning technologies, especially in the field of task scheduling and resource optimization, has significantly improved the intelligence level of the system. For example, the scheduling strategy based on the deep Q network (DQN) can adaptively learn the processing methods of different tasks and gradually optimize the task offloading decision. However, such methods often assume that the service quality requirements provided by the model are fixed or predictable. For the processing requirements of large-scale language model (LLM) requests, the service quality is often unpredictable, mainly due to the following factors: First, the performance of LLM in processing tasks is affected by the complexity of the task. Different models have different effects in tasks such as generating long texts, answering questions in specific fields, or open-ended conversations. Second, the structure and number of parameters of the model are different. Larger models usually provide better performance, but require more computing resources and time. In addition, the uncontrollability of the request context causes the performance of the same model to fluctuate in different environments, especially when multiple tasks are concurrent. The resource status and network status of the edge device will also affect the reasoning efficiency. In addition, some LLMs may perform well in specific domain knowledge, but perform poorly in other fields. These factors together make it difficult to predict the service quality of different LLMs.

[0005] (3) Some other methods treat different models as homogeneous to simplify the complexity of the optimization problem. These methods usually rely on a unified model caching strategy, aiming to improve the overall performance of the system through fixed parameters and rules, such as reducing latency and increasing cache hit rate. However, this approach fails to fully consider the unique characteristics of the large language model (LLM) architecture, resulting in the inability to effectively utilize the advantages of each model during the optimization process. For example, different LLMs have significant differences in task processing capabilities, response speeds, and resource requirements, and a unified approach ignores these differences, making it difficult to achieve the expected performance results in practical applications. Summary of the invention

[0006] The present invention proposes an online caching method for large models of edge devices based on Multi-Armed Bandit (MAB), which optimizes the resource utilization and reasoning performance of edge devices by selecting a preset number of large language models (LLMs) suitable for caching through real-time learning:

[0007] A feature information receiving module, configured to receive a feature information set of a cloud candidate LLM of an edge server device, wherein the feature information set of the cloud candidate LLM includes model parameter quantity, computing requirements, memory requirements, usage task type, inference delay baseline, and multi-task supportability;

[0008] An LLM quality prediction module, used to predict the utility vector of each LLM using a neural network based on the feature information set, and calculate the UCB value and utility offset corresponding to each LLM utility vector;

[0009] The edge sorting selection module is used to sort the sum of the UCB value and the utility offset in descending order and select the first K of the preset number of t The LLM generates an LLM subset, defines the LLM subset as a super arm, and caches the super arm and deploys it to the edge end to provide services;

[0010] Asynchronous processing output training module, used to collect service quality data and neural network gradient data, and asynchronously train and optimize the parameters of the neural network model in the LLM quality prediction module;

[0011] The above modules are run in a cyclic sequence to realize online caching and service of the large edge device model of the multi-armed gambling machine.

[0012] Optionally, in the LLM quality prediction module, the operation process of the neural network prediction specifically includes:

[0013] Use a neural network function with a fully connected structure to approximate the unknown function in the utility round model:

[0014]

[0015] Among them, σ(·) is the ReLU activation function, m is the width of the neural network hidden layer, is the feature vector of the cloud-based candidate LLM, θ is the parameter matrix of the neural network model, is the weight matrix of the i-th layer of the neural network, is the weight matrix of the last layer, is the weight matrix of the last layer.

[0016] Optionally, the utility round model is:

[0017]

[0018] Among them, u t,i is the utility round of the ith arm, h(·) is an unknown function, ∈ t is a v-sub-Gaussian noise, is the feature vector of the tth round of the ith arm.

[0019] Optionally, the LLM quality prediction module includes a utility vector prediction submodule and a utility offset calculation submodule;

[0020] The utility vector prediction submodule is used to calculate the UCB value of the utility vector based on the parameters and covariance matrix of the neural network model:

[0021]

[0022] in, is the utility vector u t,i The predicted value of t-1 is the scaling factor, m is the width of the hidden layer of the neural network, For the neural network function The gradient of

[0023] The utility offset calculation submodule is used to calculate the utility offset based on the parameters and covariance matrix of the neural network model:

[0024]

[0025] in, is the weight coefficient, K is the capacity of the edge server, L is the processing delay, and λ is the regularization parameter.

[0026] Optionally, in the asynchronous processing output training module, the service quality data is the service quality u obtained by the super arm based on the accuracy A of completing the task, the processing delay L and the computing cost C:

[0027]

[0028] Among them, L β To handle the weight of delay in service quality evaluation, C γ It is the weight value of the calculation cost in the service quality evaluation.

[0029] Optionally, in the asynchronous processing output training module, service quality data and neural network gradient data are collected, and the neural network model in the LLM quality prediction module is asynchronously trained and parameter optimized, specifically including:

[0030] Check the training signal of the neural network model trained in this round, copy the neural network model in the LLM quality prediction module, and train and optimize the copied neural network based on the collected service quality data;

[0031] Based on the trained and optimized neural network model, the gradient descent method is used to update the parameters and covariance matrix of the neural network model;

[0032] The optimized neural network model parameters are copied to the neural network model in the LLM quality prediction module.

[0033] Optionally, the covariance matrix of the neural network model is updated using a gradient descent method, specifically including:

[0034]

[0035] in, is a function The gradient of is the accumulated feedback set during the last period when the training signal is false. When the training signal is false, the parameters of the previous round are kept unchanged, that is, and

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1) Dynamic optimization of cache strategy: Based on the context-combined multi-armed bandit model, the present invention can dynamically select and cache a preset number of large language models (LLMs) that are most suitable for the current task in each round of decision-making. This real-time learning mechanism enables edge devices to flexibly adjust cache strategies according to actual request requirements and system resource conditions, significantly improving the system's response speed and task processing efficiency, and overcoming the limitation that static cache strategies in the prior art are difficult to adapt to changes.

[0038] 2) Nonlinear performance optimization: By defining the context feature vector of LLM and establishing a nonlinear mapping between performance and context, the present invention can more accurately predict the performance of the model in different environments, thereby achieving comprehensive performance optimization. This nonlinear relationship enables the system to better balance latency, accuracy, and computational cost when processing tasks of different complexity, thereby improving the overall throughput of the system.

[0039] 3) Efficient feedback mechanism: By providing real-time feedback on the execution effect of tasks (including latency, accuracy, resource consumption, etc.), the system can continuously learn and improve its decision-making process. The contextual neural network UCB algorithm adopted by the present invention can gradually optimize the cache strategy based on these service quality data, ensuring that the system can continuously improve performance in a dynamic environment and achieve near-optimal performance in long-term operation. It has been theoretically proven that this system performs stably in long-term operation, and its error growth rate will gradually slow down and eventually approach zero.

[0040] 4) Improved resource utilization efficiency: The present invention significantly reduces the storage and computing burden of edge devices by effectively selecting and caching LLMs. Compared with existing full model caching or simple task offloading solutions, the present invention can better utilize the limited resources of edge devices, especially in resource-constrained environments, and significantly improve system performance and resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0042] Figure 1 It is a structural diagram of an online caching and service system for a large model of an edge device based on a multi-armed bandit according to an embodiment of the present invention;

[0043] Figure 2 This is a schematic diagram of a fully connected neural network according to an embodiment of the present invention;

[0044] Figure 3 The present invention is a flowchart of the working method of the edge device large model online caching and service system based on the multi-armed bandit according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] First, the technical contents used in the present invention are described and explained:

[0047] In the present invention, one LLM represents one arm and corresponds to one feature vector, which is encoded by the feature information of the corresponding LLM. That is, through encoding, one LLM corresponds to a group of feature information; all the alternative LLMs have one group of feature information respectively, and multiple LLMs obtain multiple groups of feature information, and a feature information set is composed according to the multiple groups of feature information.

[0048] A set of feature information corresponding to an LLM is defined as the "context" information of this LLM, and a multi-armed bandit containing such a setting is called a "contextual multi-armed bandit".

[0049] It can be understood as follows: there are 6 items of feature information used to characterize an LLM, namely, model parameter quantity, computing requirements, memory requirements, task type used, inference delay baseline, and multi-task support. Different LLMs have different values ​​of these 6 items. The feature information of each LLM is encoded into a vector according to the established encoding scheme. Our algorithm will select multiple arms in each round, and these selected arms form a set, which is called a "super arm". This kind of multi-armed gambling machine is called a "combined multi-armed gambling machine".

[0050] Parameter count: The total number of parameters in a model, usually used to measure the size of the model. The larger the number of parameters, the more complex the model can handle, but it also consumes more resources.

[0051] Example: an LLM with 100 million parameters is usually faster than an LLM with 1 billion parameters, but is limited in complexity.

[0052] Computational requirements: The computing resources required by the model for inference, usually expressed in FLOPs (floating point operations) or other units of measurement for computational power.

[0053] Example: A model that requires 200 GFLOPs of computation will run more efficiently than a model that requires 500 GFLOPs of computation.

[0054] Memory requirements: The memory space occupied by the model during inference, including the memory required to load the model and process data.

[0055] Example: A model with a memory requirement of 4GB is suitable for devices with limited resources, while a model with a memory requirement of 16GB requires more powerful hardware support.

[0056] Applicable task types: The types of tasks that the model can handle, such as text generation, classification, translation, etc. Different models may be optimized for different tasks.

[0057] Example: Some LLMs are particularly good at text classification, while others perform better at text generation.

[0058] Inference latency baseline: The average latency of a model processing a single inference task on standard hardware. The lower the latency, the faster the response.

[0059] For example, if one model has an inference latency of 100 milliseconds and another model has an inference latency of 500 milliseconds, the former is more suitable for application scenarios that require fast response.

[0060] Multi-task support: Whether the model has the ability to process multiple tasks in parallel to improve the system's throughput and processing efficiency.

[0061] Example: An LLM that supports multitasking can handle multiple user requests simultaneously, while an LLM that does not support multitasking can only handle requests one by one.

[0062] Example description: Assume that in an actual application scenario, this embodiment has the following two alternative LLMs:

[0063] Model A:

[0064] Number of parameters: 100 million;

[0065] Computing requirements: 200GFLOPs;

[0066] Memory requirement: 4GB

[0067] Applicable task types: text classification, sentiment analysis;

[0068] Inference latency baseline: 100 ms;

[0069] Multi-tasking support: support;

[0070] Model B:

[0071] Number of parameters: 1 billion;

[0072] Computing requirements: 500GFLOPs;

[0073] Memory requirement: 16GB

[0074] Applicable task types: text generation, question answering system;

[0075] Inference latency baseline: 500 ms;

[0076] Multitasking support: Not supported.

[0077] In order to encode these feature information into numerical feature vectors and input them into the neural network, this embodiment can convert each feature into an appropriate numerical value or vector format according to the following scheme and combine them into a complete feature vector.

[0078] Parameter size: The parameter size is a numerical measure of the size of the model, and is usually logarithmic to reduce the impact of numerical differences.

[0079] Coding scheme: Take the logarithm of the parameter quantity. For example, assuming the number of model parameters is 1 billion, the value of this feature can be log(10^9)=9.

[0080] Computational requirements: Computational requirements are usually expressed in FLOPs, and can also be logarithmically reduced.

[0081] Coding scheme: Take the logarithm of the computational requirement (such as FLOPs). For example, if the computational requirement is 500GFLOPs, the value of this feature can be log(500*10^9).

[0082] Memory requirement: Memory requirement is in GB and is directly processed numerically.

[0083] Encoding scheme: Use the numerical value of memory requirement directly. For example, a 16GB model memory requirement can be directly encoded as 16.

[0084] Applicable task types: Applicable task types are usually discrete categories, such as text classification, generation, translation, etc. They can be represented by one-hot encoding.

[0085] Coding scheme: Set a set of applicable task types, such as [text classification, sentiment analysis, text generation, translation, question answering]. Each task type corresponds to a position, and the applicable task type is set to 1, and the unsuitable task type is set to 0.

[0086] Example: If a model is suitable for text generation and question answering systems, its encoding is [0,0,1,0,1].

[0087] Inference latency baseline: Inference latency represents the response speed and is usually expressed in milliseconds (ms). To reduce the impact of numerical range, the inference latency can be normalized or inversely calculated (i.e. 1 / latency) to make it proportional to the response speed.

[0088] Coding scheme: You can use 1 / inference delay or take the logarithm of the delay. Assuming the inference delay is 500 milliseconds, the coding can be 1 / 500=0.002.

[0089] Multi-tasking support: Multi-tasking support is a Boolean attribute that can be processed using binary encoding.

[0090] Coding scheme: If the model supports multi-task, it is encoded as 1, otherwise it is 0.

[0091] Synthetic feature vector example: Assume that the model has the following features:

[0092] Number of parameters: 1 billion;

[0093] Computing requirements: 500GFLOPs;

[0094] Memory requirement: 16GB

[0095] Applicable task types: text generation, question answering system;

[0096] Inference latency baseline: 500 ms;

[0097] Multitasking support: Not supported.

[0098] According to the above encoding scheme, the feature vector of the model can be expressed as:

[0099] Parameter quantity: log(10^9)=9;

[0100] Calculation requirements: log(500*10^9)≈11.7;

[0101] Memory requirement: 16

[0102] Applicable task type (one-hot encoding): [0,0,1,0,1];

[0103] Inference latency baseline: 1 / 500 = 0.002;

[0104] Multitasking support: 0.

[0105] The final combined feature vector is:

[0106] [9,11.7,16,0,0,1,0,1,0.002,0].

[0107] Example

[0108] Large-scale online caching and service system for edge devices based on multi-armed bandit, such as Figure 1 As shown, the system includes:

[0109] The feature information receiving module is used to receive a feature information set of a cloud candidate LLM of an edge server device, wherein the feature information set of the cloud candidate LLM includes model parameter quantity, computing requirements, memory requirements, usage task type, inference delay baseline and multi-task supportability.

[0110] Initialize neural network model parameters These parameters will be used in subsequent utility prediction and model training. At the same time, set the covariance matrix where λ is the regularization parameter, is the identity matrix.

[0111] In the feature information receiving module, the cloud alternative LLM is as follows: the time window for edge device servers to process a fixed number of requests is defined as a round. In round t, K is selected according to the hardware conditions of the edge device servers. t LLMs are cached. At this time, the number of candidate LLMs in the cloud is recorded as N t .

[0112] Specifically, consider a system that treats each LLM as an arm in the contextual combination semi-bandit framework. All candidate LLMs are stored in the cloud. Each LLM is embedded into a vector through a feature encoder. As a neural network function The input parameters are To predict the effectiveness of LLM. It can be composed of a series of features, such as a vector composed of feature variables including model parameter quantity, computing requirements, memory requirements, applicable task types, inference delay baseline, multi-task support, etc. In this embodiment, the time window for edge servers to process a fixed number of requests is defined as a round t, and t is used for indexing. In round t, K can be selected according to the hardware conditions. t LLMs are cached, and the number of candidate LLMs in the cloud is recorded as N t , assuming N t ≤N for any t∈[T].

[0113] The LLM quality prediction module is used to predict the utility vector of each LLM based on the feature information set using a neural network, and calculate the UCB value and utility offset corresponding to each LLM utility vector.

[0114] In each round t (from 1 to T), the feature information set (parameter quantity, computing requirements, memory requirements, applicable task types, inference latency baseline, multi-task support) of the candidate LLM in the cloud is received to form a feature vector LLM i corresponding eigenvector This information represents the characteristics of the large language model and affects the quality of the subsequent services it provides.

[0115] Use neural networks to predict the service quality of each LLM by calculating the UCB value of the utility vector and the utility offset ω t This step involves using the neural network model parameters from the previous round and the covariance matrix Make predictions to provide a basis for optimal selection of cache LLM.

[0116] Calculate the N cloud-based t UCB value of LLM and the utility offset ω t The pseudo code is shown in Table 1:

[0117] Table 1

[0118]

[0119] In the LLM quality prediction module, the operation process of the neural network specifically includes:

[0120] Due to the lack of historical data, this embodiment uses the service quality predicted by the neural network as a substitute for the actual service quality. Once the service is provided and the feedback of the actual service quality is obtained, it is used to train the neural network model and update its parameters. In the method proposed in this embodiment, the alternative LLM feature information vector is used as its corresponding context vector, and then input into the neural network to predict the utility. Figure 2 As shown, a neural network function with a fully connected structure is used to approximate the unknown function in the utility vector model:

[0121]

[0122] Where σ(·) is the ReLU activation function, is the weight matrix, and m is the width of the hidden layer. The parameter matrix can be expressed as Where p = (1 + d) m + (L-1) m 2, and vec(·) vectorizes the matrix into a flat vector. The gradient of this function is recorded as Given a confidence parameter δ∈(0,1), the scaling factor is defined as:

[0123]

[0124] are some positive constants. The parameters of the network are initialized from a Gaussian distribution. Specifically, for 1≤l≤L-1, we set Each of these Independently generated from N(0,2 / m).

[0125] The utility vector model is:

[0126]

[0127] Where h(·) is an unknown function, ∈ t is a v-sub-Gaussian noise, is the i-th eigenvector in the t-th round.

[0128] The cloud-based LLM is first based on the score U t,i +ω t Sort in descending order and select the first K t LLMs, forming a subset The LLM set to be cached at the edge.

[0129] The LLM quality prediction module includes a utility vector prediction submodule and a utility offset calculation submodule;

[0130] The utility vector prediction submodule is used to calculate the utility vector based on the parameters and covariance matrix of the neural network model:

[0131]

[0132] in, is the utility vector u t,i The predicted value of t-1 is the scaling factor, m is the width of the hidden layer of the neural network,

[0133] The utility offset calculation submodule is used to calculate the utility offset based on the parameters and covariance matrix of the neural network model:

[0134]

[0135] The edge sorting selection module is used to sort the sum of the utility vector and the UCB value in descending order, select the first LLM of a preset number to generate an LLM subset, define the LLM subset as a super arm, and cache the super arm and deploy it to the edge end to provide services.

[0136] Based on the calculated score U t,i +ω t Sort the candidate LLMs in descending order to ensure that the ones cached to the edge are the most beneficial. Select the top K after sorting t LLMs, forming an LLM subset This subset will be used for caching at the edge servers.

[0137] The asynchronous processing output training module is used to cache and deploy services for the super arm, and asynchronously train and optimize the parameters of the neural network model in the LLM quality prediction module. In the asynchronous processing output training module, the service quality data is the service quality u obtained by the super arm based on the accuracy A of completing the task, the processing delay L and the computing cost C:

[0138]

[0139] Among them, L β To handle the weight of delay in service quality evaluation, C γ It is the weight value of the calculation cost in the service quality evaluation.

[0140] In the asynchronous processing output training module, the service quality data is collected to train and optimize the neural network model, and the trained and optimized parameters are used to update the neural network model in the LLM quality prediction module. Specifically, the following steps are performed:

[0141] Check the training signal of the neural network model trained in this round, and train and optimize the neural network model based on the feedback data during the training process;

[0142] When a round of neural network model training cycle just begins, the training signal changes from "true" to "false", indicating that the training is in progress; when a round of training ends, the training signal changes from "false" to "true", indicating that a new round of neural network model training cycle can be started. In order to maintain real-time processing efficiency, this embodiment uses an asynchronous method to train the neural network model in the LLM quality prediction module. The neural network model parameters in the LLM quality prediction module will not be updated before the current training process of the neural network model is completed, but will be updated after the training is completed. If round t is in the neural network model training period, the feedback generated during round t may be used immediately. On the contrary, it will be accumulated and included in the data set of the next round of model training at the end of the neural network model training. This results in delayed feedback, in which the reward of one round may affect the model training of subsequent rounds. This strategy ensures that edge inference remains timely and uninterrupted, while efficiently handling model updates. Specifically, for any i∈S t , the real utility u t,i After passing d t,iAfter round t, the agent receives the feedback set {u s,i |(s,i)∈T t r},in make Represents the set of all feedback indexes received up to round t, expressed as

[0143] The optimization goal is to maximize the overall quality of the edge servers, while the utility value is initially unknown, by minimizing the cumulative expected regret to achieve this goal. It quantifies the potential reward loss due to suboptimal decisions and is formally defined as

[0144]

[0145] in is the utility vector, It is the offline optimal super arm in round t.

[0146] Based on the trained and optimized neural network model, the parameters and covariance matrix of the neural network model are updated by using the gradient descent method; the utility vector and utility offset are calculated based on the updated parameters and covariance matrix.

[0147] The covariance matrix of the neural network model is updated using the gradient descent method, which specifically includes:

[0148]

[0149] in, is a function The gradient of is the set of feedback accumulated during the duration of the most recent training signal being false.

[0150] On the edge, select a subset of LLMs Perform cache deployment and record the service quality of each LLM during the processing as feedback for selecting this model. This data is used for subsequent neural network model training to help improve the accuracy of model prediction.

[0151] The neural network model is optimized using the service quality data collected in each round. The service quality data is the accuracy A, processing delay L and computational cost C of the selected LLM to complete the task, and the service quality Q is calculated as:

[0152]

[0153] By using the service quality data as the prediction target of the neural network model, the system can gradually train the model parameters and improve its prediction ability, thereby improving the accuracy of utility calculations for subsequent LLMs and ultimately optimizing the performance and service quality of edge devices.

[0154] Since the goal of this embodiment is to improve the overall LLM service quality of the edge server, the LLM with higher actual service quality is more suitable for caching on the edge server. t,i is defined as the quality of service Q provided by the LLM, that is, u t,i =Q. Define the quality of service as:

[0155]

[0156] Among them, A is accuracy, which indicates the relevance and accuracy of the model generation results; L is processing delay, which indicates the time from receiving a request to generating a response; C is computational cost, which indicates the computing resources consumed when processing a request (such as CPU time, memory usage, etc.); β and γ are weight parameters, which indicate the importance of latency and computational cost in service quality evaluation. There is a complex nonlinear relationship between these characteristics and service quality. This design enables the evaluation of service quality to reflect the multidimensional characteristics of model performance and provide more accurate and dynamic performance indicators. Service quality cannot be predicted in advance due to the diversity of task complexity, variability of model performance, and dynamic context influence. Specifically, the types and complexities of tasks involved in different user requests vary, and different LLMs may perform differently under the same task. In addition, the context and environmental changes of requests (such as network status and device load) are often unpredictable, making it difficult to accurately evaluate service quality. Feedback delays further limit the model's predictive ability in the initial stage, and the relationship between service quality and influencing factors is usually nonlinear. These factors lead to the real utility u t,i Information that is unpredictable and can only be observed after the service is provided (accuracy, processing delay, and computational cost) is used as feedback to the arm. The utility is modeled as:

[0157]

[0158] where h(·) is an unknown function, ∈ t is a ν-sub-Gaussian noise, conditioned on the history and For the selected Super Arm S t ,award Defined as S t The sum of the utilities of the middle arms represents the overall quality of service provided by this LLM cache combination on the edge server.

[0159]

[0160] At the end of each round, the training signal is checked. If the training signal is "true", the training phase of the neural network model is immediately carried out, the training signal is changed to "false" and continues until the end of training, and then changed back to "true". At the same time, all the feedback data collected so far is used to train the neural network, and the gradient descent method is used to obtain the updated parameters. And calculate the covariance matrix

[0161]

[0162] in, is a function The gradient of Represents the accumulated feedback set during the last period when the training signal value was false. If the training signal is false, the parameters of the previous round are kept unchanged, that is, and

[0163] Repeat the above steps for each round. During this process, the system will dynamically adjust according to the feature information and service feedback of each LLM. (Asynchronous processing) Ensure that the LLM cache scheduling process and the neural network model training process are asynchronous to avoid the model training process affecting the real-time cache scheduling performance and ensure the efficiency and timely response of the edge server. The pseudo code of the neural network training method is shown in Table 2:

[0164] Table 2

[0165]

[0166] like Figure 3 As shown, the present invention aims at the resource limitations, unpredictable service quality and inflexible caching strategy faced by edge devices in the large model reasoning process in the prior art, and proposes an online caching and service system for large models of edge devices based on a multi-armed bandit (MAB). The method dynamically selects and caches small large language models suitable for the current task, adapts to changing user requests in real time to optimize the overall service quality. At the same time, combined with contextual features, the multi-armed bandit model intelligently manages resources to maximize the performance of edge devices. In addition, the present invention uses feedback learning mechanisms and nonlinear modeling methods to accurately evaluate the performance of different mini LLMs in specific contexts, ensuring that the selected model can achieve the best results in practical applications. In response to dynamically changing network conditions and user needs, the dynamic caching method of the present invention can continuously update and optimize caching decisions based on real-time service quality data, enhance the adaptability and flexibility of the system, and improve service quality.

[0167] Specifically, the algorithm pseudo code of the online caching and service system of the edge device large model of the multi-armed bandit (MAB) is shown in Table 3:

[0168] Table 3

[0169]

[0170] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. An online caching and service system for large edge device models based on multi-armed bandit machines, characterized in that: The system includes: A feature information receiving module, configured to receive a feature information set of a cloud candidate LLM of an edge server device, wherein the feature information set of the cloud candidate LLM includes model parameter quantity, computing requirements, memory requirements, usage task type, inference delay baseline, and multi-task supportability; An LLM quality prediction module, used to predict the utility vector of each LLM using a neural network based on the feature information set, and calculate the UCB value and utility offset corresponding to each LLM utility vector; The edge sorting selection module is used to sort the sum of the UCB value and the utility offset in descending order and select the first K of the preset number of t The LLM generates an LLM subset, defines the LLM subset as a super arm, and caches the super arm and deploys it to the edge end to provide services; Asynchronous processing output training module, used to collect service quality data and neural network gradient data, and asynchronously train and optimize the parameters of the neural network model in the LLM quality prediction module; The above modules are run in a cyclic sequence to realize online caching and service of the large edge device model of the multi-armed gambling machine.

2. The edge device large model online caching and service system based on the multi-armed bandit according to claim 1 is characterized in that: In the LLM quality prediction module, the operation process of the neural network prediction specifically includes: Use a neural network function with a fully connected structure to approximate the unknown function in the utility round model: Among them, σ(·) is the ReLU activation function, m is the width of the neural network hidden layer, is the feature vector of the cloud-based candidate LLM, θ is the parameter matrix of the neural network model, is the weight matrix of the i-th layer of the neural network, is the weight matrix of the last layer, is the weight matrix of the last layer.

3. The edge device large model online caching and service system based on the multi-armed bandit according to claim 2 is characterized in that: The utility round model is: Among them, u t,i is the utility round of the ith arm, h(·) is an unknown function, ∈ t is a v-sub-Gaussian noise, is the feature vector of the tth round of the ith arm.

4. The edge device large model online caching and service system based on the multi-armed bandit according to claim 3 is characterized in that: The LLM quality prediction module includes a utility vector prediction submodule and a utility offset calculation submodule; The utility vector prediction submodule is used to calculate the UCB value of the utility vector based on the parameters and covariance matrix of the neural network model: in, is the utility vector u t,i The predicted value of t-1 is the scaling factor, m is the width of the hidden layer of the neural network, For the neural network function The gradient of The utility offset calculation submodule is used to calculate the utility offset based on the parameters and covariance matrix of the neural network model: in, is the weight coefficient, K is the capacity of the edge server, L is the processing delay, and λ is the regularization parameter.

5. The edge device large model online caching and service system based on the multi-armed bandit according to claim 1 is characterized in that: In the asynchronous processing output training module, the service quality data is the service quality u obtained by the super arm based on the accuracy A of completing the task, the processing delay L and the computing cost C: Among them, L β To handle the weight of delay in service quality evaluation, C γ It is the weight value of the calculation cost in the service quality evaluation.

6. The edge device large model online caching and service system based on the multi-armed bandit according to claim 5 is characterized in that: In the asynchronous processing output training module, the service quality data and the neural network gradient data are collected, and the neural network model in the LLM quality prediction module is asynchronously trained and parameter optimized, specifically including: Check the training signal of the neural network model trained in this round, copy the neural network model in the LLM quality prediction module, and train and optimize the copied neural network based on the collected service quality data; Based on the trained and optimized neural network model, the gradient descent method is used to update the parameters and covariance matrix of the neural network model; The optimized neural network model parameters are copied to the neural network model in the LLM quality prediction module.

7. The edge device large model online caching and service system based on the multi-armed bandit according to claim 6 is characterized in that: The covariance matrix of the neural network model is updated using the gradient descent method, which specifically includes: in, is a function The gradient of is the accumulated feedback set during the last period when the training signal is false. When the training signal is false, the parameters of the previous round are kept unchanged, that is, and

Citation Information

Patent Citations

  • Cognitive service caching method and system oriented to resource consumption application

    CN113868153A

  • Dynamic task replication method, device and system in edge computing environment

    CN114090218A

  • Edge network computing system with deep reinforcement learning based task scheduling

    US20230153124A1

Cited By

  • Recommendation method and system based on reinforcement learning driving large model adaptive prompt

    CN120849720A