Service deployment and model selection system and method for reasoning precision perception in edge environment
By combining sparse self-attention mechanism, pruning search and deep reinforcement learning in the edge computing environment, a service deployment and model selection system was designed to solve the problem of balancing inference accuracy and latency in edge computing, and to achieve efficient service quality improvement and resource utilization.
Patent Information
- Application Number
- CN202511102382.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-11
AI Technical Summary
In edge computing environments, it is difficult to balance inference accuracy and latency, achieve efficient service deployment and model selection, meet diverse user needs and resource constraints, and select the best model configuration to maximize service benefits.
We adopt a service deployment and model selection system that is aware of inference accuracy in edge environments. Combining probabilistic sparse self-attention mechanism, pruning search and deep reinforcement learning, we predict traffic distribution through self-attention mechanism, design pruning search strategy, introduce deep imitation learning to guide DRL agent, and optimize model selection and resource allocation.
It improved service quality, reduced latency, increased inference accuracy satisfaction rate, and enhanced system resource utilization efficiency, demonstrating faster convergence speed and higher stability, and adapting to dynamic service environments.
Smart Images

Figure CN120935185A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge computing technology, and in particular to a service deployment and model selection system and method for inference accuracy awareness in edge environments. Background Technology
[0002] With the rapid development of artificial intelligence technology, deep neural networks (DNNs) have been widely used to support and enhance many emerging intelligent applications (such as autonomous driving, video analytics, and chatbots). However, classic cloud computing models are struggling to efficiently meet users' increasing demands for inference speed and privacy protection. To alleviate this problem, edge computing pushes computing and storage resources to the network edge to improve system response speed and reduce pressure on the core network. In edge computing, edge service providers (ESPs) provide users with DNN model deployments to fulfill service requests, but still face a series of technical challenges such as task scheduling, resource allocation, and cost control.
[0003] Most existing work focuses on reducing latency and energy consumption for computationally intensive tasks by optimizing scheduling strategies, but neglects the impact of DNN inference accuracy on service quality. In real-world DNN applications, inference accuracy and service experience are strongly correlated. For example, in autonomous driving scenarios, even slight deviations in object detection inference accuracy can lead to serious traffic accidents. Furthermore, to meet diverse user needs and resource constraints, many DNN models offer network architectures of different sizes. For instance, the YOLOv8 series includes five versions: Nano, Small, Medium, Large, and Extra-Large, each with differences in model complexity, computational cost, and inference accuracy. By appropriately selecting model configurations, the service experience can be further improved while meeting different user needs. However, due to the limited computing and storage resources of edge systems, it is difficult to deploy all different models for all services online to edge servers. For specific edge scenarios, the user needs they serve usually follow certain patterns. By analyzing the demand distribution and appropriately deploying specified services and models, the service experience can be improved while meeting resource constraints. However, achieving efficient service deployment and model selection in edge environments still faces the following key challenges:
[0004] (1) It is difficult to balance inference accuracy and latency to improve service experience quality. Dynamic computing resources and diverse task attributes make accurate characterization of experience quality extremely difficult. For dynamic inference requests from users, network architectures of different scales will produce different inference accuracy and latency. At the same time, factors such as the allocated computing frequency and upload bandwidth will also affect inference latency. In addition, users usually have different Service Level Objectives (SLOs) for different types of services. For example, for object detection services, users tend to have real-time perception of the detected objects; for image generation services, users tend to generate higher quality images. Therefore, how to balance inference accuracy and latency according to service type to improve experience quality needs to be carefully considered. Although some works have considered the impact of inference accuracy and latency on experience quality, most of them are designed for a specific inference service and cannot be effectively extended to other inference service types, making it difficult to meet the dynamic needs of practical applications for diverse inference services.
[0005] (2) Difficulty in efficiently deploying services and models under memory constraints to meet user service preferences. The differentiated characteristics of different services and models (e.g., memory consumption, inference accuracy, and computational latency) pose a significant challenge to selecting appropriate model deployment schemes for services. For users in edge environments, their request traffic patterns for various types of services may differ significantly, and these patterns change dynamically over time. Therefore, service deployment schemes need to effectively address this dynamism while also ensuring fairness among different services to avoid long-distance cloud transmission latency caused by some requests not receiving timely responses. Although some works have attempted to predict the request traffic patterns of different services, they are limited by inherent long-distance dependency issues and cannot effectively capture long-term sequential dependencies, thus making it difficult to accurately predict the long-term changing trends of user request patterns over time.
[0006] (3) Difficulty in selecting the optimal model configuration to maximize service benefits. Based on the service and model deployment scheme, how to further select the optimal model configuration to maximize service benefits is a key issue. Compared to high-precision deep DNN models, shallow DNN models usually have faster inference speeds, but may also experience a significant drop in inference accuracy, especially when facing complex task inputs. This invention has found through experiments that the performance differences of models of different sizes on various tasks are diverse. For example, for some tasks, the accuracy of deep DNN models is actually lower than that of shallow DNN models. At the same time, this difference dynamically changes with different task inputs (i.e., the performance differences shown for tasks with different characteristics may change). As an emerging intelligent decision-making algorithm, Deep Reinforcement Learning (DRL) has been applied to many optimization problems in edge environments (such as computation offloading and resource allocation), and can serve as a potentially feasible solution for selecting model configurations. However, when facing new environments, classic DRL requires training the agent from scratch and gradually optimizing the strategy, leading to excessive training time and resource overhead. Furthermore, existing solutions typically ignore the problem scale of model selection and resource allocation decisions, making it difficult to achieve fast and stable convergence. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a service deployment and model selection system and method with inference accuracy awareness in edge environments, which comprehensively considers multiple factors such as inference latency, inference accuracy and memory limitations, aiming to improve the utilization efficiency of edge resources while taking into account the quality of user services.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: a service deployment and model selection system for inference accuracy awareness in an edge environment, comprising a user terminal and an edge server; the user terminal includes a terminal device and an inference request module; the edge server includes a service type selection module and a model selection module; the system operates based on multiple time slots; in each time slot t∈T={1,2,...,T}, the user's inference task i is denoted as a quintuple <d i ,s i ,η i ,acc i ,ρ i >, where d i Indicates the task size, s i ∈{s1,s2,...,s J} represents the service type, η i ={η i,1 ,η i,2 ,...η i,K} represents the computational density under different configurations, acci ={acc i,1 ,acc i,2 ,...acc i,k} represents the inference accuracy under different configurations, ρ i This indicates the task priority; the computational density and inference accuracy of a task will vary depending on the resource allocation.
[0009] In a preferred embodiment, the task processing includes an upload and an execution phase. During the upload phase of task i, the upload rate is defined according to the Shannon channel capacity formula.
[0010]
[0011] Among them, b i Let p represent the bandwidth allocated to task i, and g represent the signal power. i σ represents the channel gain. 2 This represents the noise power during signal transmission.
[0012] The upload latency of task i is defined as
[0013]
[0014] Once a task is uploaded to the edge server, the proposed system will select a model configuration for it to complete inference. If the model corresponding to the user's requested service is not deployed, the request will be forwarded to the remote cloud for processing. The latency of forwarding user task i from the edge server to the remote cloud is defined as...
[0015]
[0016] in, This represents the latency for forwarding a request to the remote cloud under configuration k corresponding to service type j; x j,k ∈{0,1} represents the model deployment decision; if the configuration k corresponding to the service type j required by task i is deployed on the edge server, x j,k =1; otherwise, x j,k =0;
[0017] When performing DNN inference on an edge server, the execution latency of task i is defined as follows:
[0018]
[0019] Where f represents the computing power of the edge server; y i,k ∈{0,1} represents the model configuration decision assigned to task i; if task i chooses configuration k, y i,k =1; otherwise, y i,k =0;
[0020] The total latency required to process task i is defined as
[0021] T i total =T i up +T i tran +T i exe (5)
[0022] For service requests initiated by users, the computational density and inference accuracy vary depending on the model configuration; based on the model configuration selected by the user, the inference accuracy of task i is expressed as follows:
[0023]
[0024] In a preferred embodiment, the ESP receives a benefit from the user when the total processing time and inference accuracy of the task both meet the user's requirements; this process is defined as follows:
[0025]
[0026] in, and These represent the minimum tolerable precision and the maximum tolerable latency for service j, respectively.
[0027] The higher the priority of a user's task, the greater the benefit ESP receives; within time slot t, the service benefit is defined as...
[0028]
[0029] Where m represents the number of users served by ESP within time slot t.
[0030] In a preferred embodiment, based on the proposed system model, the optimization objective is to maximize the long-term system benefit under system memory constraints; this optimization problem P1 is formally expressed as:
[0031]
[0032] Where, m j,k This represents the memory required for configuration k corresponding to the deployment service type j, and M is the total available memory in the edge environment.
[0033] This invention also provides a method for inference accuracy-aware service deployment and model selection in edge environments. Based on the aforementioned service deployment and model selection system for inference accuracy-awareness in edge environments, every T... deployIn each time slot, the latest service deployment plan is obtained, the service deployment plan is executed, and the service deployment status is updated; the service traffic distribution of future time slots is obtained through a traffic prediction model based on a self-attention mechanism, and the optimal model deployment plan for each service is obtained through pruning search based on this.
[0034] In each time slot t, a suitable model configuration and resource allocation scheme is selected for the inference task request, and the task completion status and service traffic are collected after the task is completed. The size of the task and the computational density attributes of the task under different configurations are used as state information to input the agent for model configuration decision-making. The behavior cloning loss is constructed by obtaining expert actions through offline search, and is combined with the online loss in the process of agent exploration and environment exploration to update the policy network, thereby accelerating the agent's fitting of the optimal model configuration strategy.
[0035] In a preferred embodiment, obtaining the latest service deployment plan, executing the service deployment plan, and updating the service deployment status are achieved through the following method: First, initialize the candidate service deployment plan p. best and its potential value v best And construct encoder input X en and decoder input X de , where X his X represents the historical inference request traffic. input X represents the traffic sequence collected in the current window. 0 This represents temporal sampling of the target sequence; probabilistic sparse self-attention is introduced when constructing the encoder, and self-attention distillation is used between layers to reduce computational overhead; feature extraction from layer j to layer j+1 is defined as...
[0036]
[0037] Where Convld(·) represents a one-dimensional convolution over a time series, ELU represents the activation function, MaxPool(·) represents the max pooling operation, [·] AB This indicates sparse self-attention;
[0038] Adding a max-pooling layer after Conv1d reduces the input downsampling to half its original size, thereby reducing the computational complexity of the training process from... Reduce to Where L represents The length of the encoder is then determined. Next, the encoder output and decoder input are fed into the decoder, which contains one layer of multi-head sparse probabilistic self-attention mechanism and one layer of multi-head attention mechanism. Subsequently, the encoder output is fed into a multilayer perceptron (MLP) to predict the desired traffic sequence. Then, based on the traffic prediction results, all feasible service deployment schemes are iteratively searched. During the search, schemes with memory consumption exceeding the threshold M are first excluded because they exceed the constraint C2 in P1. Further, schemes with memory consumption below the threshold σM are excluded, ensuring that the memory consumption of all candidate deployment schemes lies within [σM, M], where σ is a pruning coefficient used to control the balance between search efficiency and deployment scheme overhead. If a deployment scheme has a lot of remaining memory, its value is further increased by deploying more models. For candidate service deployment schemes, the potential value of deploying each model is quantified, defined as...
[0039]
[0040] in, and These represent the model's average accuracy and average latency on the test set, respectively, with δ representing the weighting coefficient; next; a penalty term is designed based on the latency of task transmission to the cloud; the potential value of a service deployment scheme is defined as...
[0041]
[0042] Where I(·) represents the indicator function; when service j has no deployed model, I(·) = 1; otherwise, I(·) = 0;
[0043] Finally, based on the memory consumption and potential value of each model, an iterative search is performed, and the best service deployment scheme with the highest potential value is output.
[0044] In a preferred embodiment, a Domain-Specific Identifier (DIL) is introduced to train the DRL agent by learning from expert demonstrations, and a Domain-Specific Knowledge Base (BC) is introduced to help the DRL agent quickly learn expert knowledge in a specified domain. During the interaction between the DRL agent and the environment, the corresponding state space, action space, and reward function are defined as follows:
[0045] (1) State space: includes the service type of the task, data size, and optional model attributes; at the same time, pixel quantity and image entropy are added to the state space; in vision tasks, pixel quantity PC i Image entropy is the product of the image's width and height in pixels. It measures the amount of information and uncertainty in an image; higher image entropy indicates richer information. According to Shannon entropy, image entropy is defined as...
[0046]
[0047] Where, ph This represents the probability of the h-th grayscale pixel appearing;
[0048] Therefore, the state space in time slot t is defined as
[0049] s t ={d i ,s i ,ρ i PC i H i} (14)
[0050] (2) Action Space; The goal of the DRL agent is to select an appropriate model configuration and bandwidth allocation for each inference request based on the current state space to achieve a good balance between inference accuracy and latency; therefore, the action space in time slot t is defined as
[0051] a t ={y i,k ,b i} (15)
[0052] (3) Reward function; First, the reward function needs to maximize the cumulative system revenue; Second, if the model selected by the DRL agent is not deployed, its task is rescheduled to the available model with the lowest accuracy, and a penalty term r is fed back. penalty The remaining available bandwidth is taken into account in the reward function, denoted as r. res Therefore, the reward function in time slot t is defined as follows:
[0053]
[0054] in,
[0055]
[0056] In a preferred embodiment, selecting a suitable model configuration and resource allocation scheme for the inference task request is achieved by the following method: For each received inference task, construct the action state input s according to formula (14). t Through offline simulated environment interaction, expert models dynamically search for and select strategies, and store these strategies in the form of <state-action-reward> triples in the expert experience pool for subsequent guidance of the agent's strategy training process; subsequently, the state s t Input is fed into the Actor network to obtain model selection and resource allocation actions; upon receiving action a... t Afterwards, the environment will execute the action and provide a reward and the next state, which will be stored in the local experience pool for updating model parameters; a discount factor γ is introduced to calculate the reward for the current task, defined as follows:
[0057]
[0058] Generalized advantage estimation (GAE) is introduced as the objective of network updates, defined as follows:
[0059]
[0060] δ t =R t +γV φ (s t+1 )-V φ (s t (20)
[0061] Where λ represents the discount rate of the advantage function, δ t This represents the timing difference (TD), where V is the state-value function;
[0062] The range of each update of the Actor network is controlled by pruning the objective function, which is defined as follows:
[0063]
[0064] Where r(θ) represents the sampling rate between the old and new policies; by adjusting r(θ), the sampling bias caused by the policy update is corrected, which is defined as...
[0065]
[0066] Next, by minimizing L critic (φ) is used to optimize the Critic network, which is defined as follows:
[0067]
[0068] For the action space, its behavior cloning loss function is defined as follows:
[0069]
[0070] in, Indicates that the expert is in state s t The following action, Indicates in s t Lower output The probability of;
[0071] By combining the behavioral cloning loss with the losses of the Actor and Critic networks in DRL, the total loss function is defined as follows:
[0072] L total =L actor (θ)+L critic (φ)+λL BC (θ) (25)
[0073] Here, λ represents the weighting factor used to control the relative importance of BC and DRL; by controlling the change of λ, the exploration cost of the policy is significantly reduced in the early stage of training to accelerate model convergence; in the later stage of training, the model performance is further optimized to generalize to other state spaces.
[0074] Compared with existing technologies, this invention has the following advantages: By effectively combining probabilistic sparse self-attention mechanisms, pruning search, deep reinforcement learning (DRL), and deep transfer learning (DIL) techniques, this invention proposes a novel SDMS method to address the revenue optimization problem of precision-aware service deployment and model selection in edge computing environments, aiming to improve service quality and enhance system resource utilization efficiency. Based on real user communication traffic datasets and DNN inference datasets, this invention verifies the effectiveness of the proposed SDMS method in improving service revenue, reducing service latency, and increasing inference accuracy satisfaction through extensive experiments. Experimental results show that compared with five other benchmark methods, the SDMS method exhibits superior latency and accuracy performance under different service requirements and model configurations. Furthermore, compared with advanced PredRNN and PPO methods, the SDMS method also demonstrates faster convergence speed and higher stability. Simulation experiments further verify the good generalization ability of the SDMS method under different inference traffic scales, indicating its broad applicability in dynamic service environments. Attached Figure Description
[0075] Figure 1 This is a schematic diagram of the model structure of a service deployment and model selection system for edge environment precision awareness according to a preferred embodiment of the present invention;
[0076] Figure 2 This is a flowchart illustrating the SDMS method proposed in a preferred embodiment of the present invention;
[0077] Figure 3 This is a schematic diagram comparing the inference latency and accuracy of different YOLOv8 models on COCO dataset samples, representing a preferred embodiment of the present invention.
[0078] Figure 4 This is a schematic diagram comparing the convergence of different methods in a preferred embodiment of the present invention;
[0079] Figure 5 This diagram illustrates the comparison of accuracy and latency satisfaction rates of different methods according to a preferred embodiment of the present invention on different services.
[0080] Figure 6 This is a schematic diagram of an ablation experiment on each core component of the SDMS method according to a preferred embodiment of the present invention;
[0081] Figure 7This diagram illustrates a comparison of service revenue under different traffic volumes for different methods of a preferred embodiment of the present invention. Detailed Implementation
[0082] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0083] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0084] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0085] A novel method for inference accuracy-aware service deployment and model selection in edge environments, referencing Figure 1-7 This application:
[0086] (1) A unified service deployment and model selection system model for edge environment inference accuracy awareness is proposed to cope with diverse service request types and model scale configurations. The proposed system model comprehensively considers multiple factors such as inference latency, inference accuracy and memory limitations, aiming to improve the utilization efficiency of edge resources while taking into account the quality of user services.
[0087] (2) A proactive service deployment method based on a self-attention mechanism was designed. Based on historical user service request preferences, a probabilistic sparse self-attention mechanism was used to predict the distribution of user request traffic for different services in future time slots. Furthermore, a potential value function was constructed for different attributes of each model, and an efficient pruning search strategy was designed to obtain the optimal service deployment scheme that matches the current user request traffic distribution.
[0088] (3) A model selection and resource allocation method based on Deep Imitation Learning (DIL) guided DRL is designed. By introducing generalized advantage estimation and pruning objective function, a suitable model is efficiently selected to improve the service quality of DNN inference. In particular, by introducing DIL, the proposed method can make full use of expert experience to guide the DRL agent to accelerate model convergence and achieve a good balance between exploration and utilization.
[0089] (4) Based on real user communication traffic and DNN inference datasets, extensive experiments verified the effectiveness of the proposed SDMS method. The results show that the SDMS method can make service deployment and model selection decisions quickly and effectively. Compared with other benchmark methods, the SDMS method demonstrates superior service benefits, latency, and accuracy satisfaction in different scenarios.
[0090] I. System Model
[0091] 1. Specifically, the proposed service deployment and model selection system model for edge environment inference accuracy awareness is as follows: Figure 1 As shown, to reduce user inference latency, ESP leases computing resources from edge servers to provide DNN inference services (e.g., knowledge answering, object detection, autonomous driving, etc.). To cater to the needs of different users and provide customized services, ESP offers various model configurations for its different service types (e.g., YOLOv8N, YOLOv8S, etc. in the YOLOv8 series). These models have different network scales and will produce different inference accuracies and latency. Therefore, ESP needs to provide users with appropriate model configurations to further improve the service experience. However, limited by the available resources at the edge, ESP cannot deploy all models online. Therefore, it is necessary to rationally select and deploy only some models to improve the efficiency of edge resource utilization while ensuring service quality. At the same time, ESP needs to allocate appropriate bandwidth resources for inference requests initiated by different user devices to ensure stable data transmission.
[0092] Specifically, the proposed system operates on a multi-slot basis. In each slot t∈T={1,2,...,T}, the user's inference task i is denoted as a quintuple <d i ,s i ,η i ,acc i ,ρ i >, where d i Indicates the task size, s i ∈{s1,s2,...,s J} represents the service type, η i ={η i,1 ,η i,2 ,...η i,K} represents the computational density under different configurations, acc i ={acc i,1 ,acc i,2 ,...acc i,k} represents the inference accuracy under different configurations, ρ iThis indicates task priority. The computational density and inference accuracy of a task vary depending on resource configuration. Therefore, performance under different resource configurations will directly affect task latency, accuracy satisfaction, and resource consumption, thereby impacting service revenue.
[0093] 2. Communication and Computation Model
[0094] The task processing involves two phases: uploading and execution. Specifically, in the uploading phase of task i, according to the Shannon channel capacity formula, its uploading rate is defined as...
[0095]
[0096] Among them, b i Let p represent the bandwidth allocated to task i, and g represent the signal power. i σ represents the channel gain. 2 This represents the noise power during signal transmission.
[0097] Therefore, the upload latency of task i is defined as
[0098]
[0099] Once a task is uploaded to the edge server, the proposed system will select a suitable model configuration for it to complete the inference. If the model corresponding to the user's requested service is not deployed, the request will be forwarded to the remote cloud for processing. The latency of forwarding user task i from the edge server to the remote cloud is defined as...
[0100]
[0101] in, This represents the latency for forwarding requests to the remote cloud under configuration k corresponding to service type j. j,k ∈{0,1} represents the model deployment decision. If the configuration k corresponding to the service type j required by task i is deployed on the edge server, x j,k =1; otherwise, x j,k =0.
[0102] When performing DNN inference on an edge server, the execution latency of task i is defined as follows:
[0103]
[0104] Where f represents the computing power of the edge server (i.e., CPU frequency). y i,k ∈{0,1} represents the model configuration decision assigned to task i. If task i chooses configuration k, y i,k =1; otherwise, y i,k =0.
[0105] Therefore, the total latency required to process task i is defined as
[0106] T i total =T i up +T i tran +T i exe (5)
[0107] For service requests initiated by users, different model configurations correspond to different computational densities and inference accuracies. Based on the model configuration selected by the user, the inference accuracy of task i can be expressed as follows:
[0108]
[0109] 3. Profit Model
[0110] The goal of the ESP service is to deliver satisfactory inference results within a latency that is tolerable for the user. Therefore, ESP can generate revenue from the user when both the total processing time and inference accuracy of the task meet their needs.
[0111] This process is defined as
[0112]
[0113] in, and These represent the minimum tolerance precision and the maximum tolerance latency for service j, respectively.
[0114] The higher the priority of a user's task, the greater the benefit ESP receives. Within time slot t, the service benefit is defined as...
[0115]
[0116] Where m represents the number of users served by ESP within time slot t.
[0117] 4. Problem Definition
[0118] Based on the proposed system model, the optimization objective of this invention is to maximize the long-term system benefit under system memory constraints. This optimization problem P1 can be formally expressed as:
[0119]
[0120] Where, m j,k This represents the memory required for configuration k corresponding to the deployment service type j, and M is the total available memory in the edge environment.
[0121] II. A Novel Service Deployment and Model Selection Method for Inference Accuracy Awareness in Edge Environments: SDMS
[0122] 1. In response to problem P1, this invention proposes a novel inference precision-aware service deployment and model selection method SDMS in edge environments, which aims to maximize service benefits through accurate analysis and matching of diverse service requests and model configurations. Figure 2 The proposed SDMS method is summarized, and its main steps are shown in Algorithm 1. Every T... deploy In each time slot, Algorithm 2 is called to obtain the latest service deployment plan, execute the service deployment plan, and update the service deployment status (lines 2-5). Specifically, this invention designs a traffic prediction model based on a self-attention mechanism to obtain the service traffic distribution of future time slots, and uses pruning search to obtain the optimal model deployment plan for each service.
[0123] In each time slot t, Algorithm 3 is invoked to select a suitable model configuration and resource allocation scheme for the inference task request, and the task completion status and service traffic are collected after the task is completed (lines 6-12). Specifically, this invention inputs attributes such as the task size and computational density under different configurations as state information into the agent for model configuration decisions. The behavioral cloning loss is constructed by obtaining expert actions through offline search, and combined with the online loss from the agent's environmental exploration process to update the policy network, thereby accelerating the agent's fitting of the optimal model configuration strategy.
[0124]
[0125] 2. Proactive service deployment based on traffic prediction
[0126] Service deployment aims to rationally deploy services suitable for the accuracy and latency requirements of the current task under edge memory constraints. However, user request traffic varies significantly across different services and changes dynamically over time. Therefore, it is necessary to allocate appropriate memory to different services and dynamically adjust it according to changes in system state. Furthermore, since the range of inference latency and accuracy generated by different services can vary considerably, it is necessary to select the appropriate model based on the attributes of the different models included in each service.
[0127] By leveraging historical traffic variation characteristics, future user request patterns can be predicted, thereby supporting service deployment. To improve service deployment stability and avoid frequent model switching, it is necessary to effectively capture long-term sequence dependencies to achieve accurate traffic prediction.
[0128] Compared to classic Recurrent Neural Networks (RNNs) and Convolutional Neural Networks (CNNs), the emerging Transformer demonstrates outstanding capabilities in capturing long-term dependencies and dynamic correlations between features. It can explicitly model the correlations of features at different time steps, thus more accurately understanding the global spatiotemporal patterns of user service requests. Furthermore, thanks to its self-attention mechanism, the Transformer can handle input data at different time scales without adjusting the model structure, making it suitable for inference request traffic prediction problems. Meanwhile, service deployment requires consideration of fairness to reduce high latency caused by tasks being forwarded to the cloud due to some services not having any configuration deployed. Specifically, since each model has different resource requirements, memory costs, and performance characteristics, service deployment must fully consider the capabilities and load of each node to ensure efficient task allocation, improving service quality while meeting memory constraints. Through reasonable load balancing and fairness mechanisms, unnecessary task forwarding can be reduced, effectively lowering transmission latency and optimizing overall system performance. In addition, each model has different potential value and memory costs, requiring efficient service deployment to improve system utility while meeting memory constraints.
[0129] To address the aforementioned issues, this invention proposes a proactive service deployment method based on traffic prediction, the key steps of which are shown in Algorithm 2. First, the candidate service deployment scheme p is initialized. best and its potential value v best And construct encoder input X en and decoder input X de (Lines 1-2), where X his X represents the historical inference request traffic. input X represents the traffic sequence collected in the current window. 0 This represents the temporal sampling of the target sequence (i.e., the flow sequence to be predicted). Classical self-attention mechanisms require calculating the attention weights of each position with all other positions during encoder construction, leading to high computational complexity. To alleviate this problem, inspired by this concept, this invention introduces probabilistic sparse self-attention during encoder construction and uses self-attention distillation between layers to reduce computational overhead. Specifically, feature extraction from layer j to layer j+1 is defined as...
[0130]
[0131] Where Convld(·) represents a one-dimensional convolution over a time series, ELU represents the activation function, MaxPool(·) represents the max pooling operation, [·] ABThis indicates sparse self-attention.
[0132] Adding a max-pooling layer after Conv1d reduces the input downsampling to half its original size, thereby reducing the computational complexity of the training process from... Reduce to Where L represents The length of the data is shown in line 3. Next, the encoder output and decoder input are fed into the decoder (line 4), which contains one layer of multi-head sparse probabilistic self-attention and one layer of multi-head attention. Then, the encoder output is fed into a multilayer perceptron (MLP) to predict the desired traffic sequence (line 5). Next, based on the traffic prediction results, all feasible service deployment schemes are iteratively searched (lines 6–14). During the search, schemes with memory consumption exceeding the threshold M are first excluded because they exceed the constraint C2 in P1; further, schemes with memory consumption below the threshold σM are excluded, ensuring that the memory consumption of all candidate deployment schemes lies within [σM, M], where σ is the pruning coefficient used to control the balance between search efficiency and deployment scheme overhead (lines 7–9). If a deployment scheme has a lot of remaining memory, its value can be further increased by deploying more models. Therefore, the optimal service deployment scheme tends to select a set of models whose total memory is close to but does not exceed the memory limit.
[0133] For candidate service deployment schemes, this invention quantifies the potential value of deploying each model, which is defined as follows:
[0134]
[0135] in, and These represent the model's average accuracy and average latency on the test set, respectively, and δ represents the weighting coefficient.
[0136] Next, to avoid unfair service deployment due to uneven traffic distribution, this invention incorporates a penalty based on the latency of task transmission to the cloud. Therefore, the potential value of a service deployment scheme is defined as...
[0137]
[0138] Here, I(·) represents the indicator function. When service j has no deployed model, I(·) = 1; otherwise, I(·) = 0. Finally, an iterative search is performed based on the memory consumption and potential value of each model, and the optimal service deployment scheme with the highest potential value is output (lines 10-13).
[0139]
[0140] 3. Model selection and resource allocation based on guided DRL
[0141] For service deployment schemes, a dynamic selection method for diverse model configurations needs to be designed, combined with real-time resource allocation strategies, to collaboratively optimize the accuracy and latency requirements of different inference requests. As an emerging intelligent decision-making algorithm, DRL can be considered a potentially feasible solution to the model selection and resource allocation problem. However, classic DRL typically requires starting from scratch and gradually optimizing strategies when facing new environments. This not only consumes significant training time and system resources but may also lead to convergence difficulties for the DRL agent, especially when dealing with large-scale state and action spaces. With the diversification of service types and model configurations in edge environments, the complexity of the strategies that the DRL agent needs to fit increases exponentially, greatly increasing the difficulty of fitting reasonable model selection and resource allocation strategies. To address this challenge, this invention introduces DIL, training the DRL agent by learning expert-demonstrated behaviors to make its strategies approximate expert strategies, avoiding potentially extensive trial-and-error exploration. As a widely used supervised learning method in DIL, Behavior Cloning (BC) can directly utilize expert knowledge to help the DRL agent learn the mapping relationship between states and actions. Compared to interacting with the environment from scratch, the introduction of BC can help DRL agents quickly learn expert knowledge in a specified domain, thereby accelerating their convergence speed and improving the practicality of DRL in resource-constrained edge environments.
[0142] In the interaction between the DRL agent and the environment, the corresponding state space, action space, and reward function are defined as follows:
[0143] (1) State Space. This includes attributes such as the service type of the task, data size, and available models. Additionally, to more accurately reflect the inference density and complexity of the task, pixel count and image entropy are added to the state space. In visual tasks, pixel count PC... i Image entropy is the product of the image's width and height in pixels. It measures the amount of information and uncertainty in an image; higher image entropy indicates that the image contains richer information. According to Shannon entropy, image entropy is defined as...
[0144]
[0145] Where, p h This represents the probability of the h-th grayscale pixel appearing.
[0146] Therefore, the state space in time slot t is defined as
[0147] s t ={d i ,s i ,ρi PC i H i}. (14)
[0148] (2) Action Space. The goal of a DRL agent is to select an appropriate model configuration and bandwidth allocation for each inference request based on the current state space to achieve a good balance between inference accuracy and latency. Therefore, the action space in time slot t is defined as follows:
[0149] a t ={y i,k ,b i}. (15)
[0150] (3) Reward Function. First, the reward function needs to maximize the cumulative system revenue (calculated according to formula (7)). Second, if the model selected by the DRL agent is not deployed, its task is scheduled to the available model with the lowest accuracy, and a penalty term r is fed back. penalty Furthermore, to prevent early-arriving tasks from monopolizing remaining resources, the remaining available bandwidth is considered in the reward function, denoted as r. res Therefore, the reward function in time slot t is defined as follows:
[0151]
[0152] in,
[0153]
[0154] Based on the above definitions, this invention proposes a model selection and resource allocation method based on guided DRL, the key steps of which are shown in Algorithm 3. For each received inference task, the action state input s is constructed according to formula (14). t This invention considers constructing expert actions for some tasks. However, unlike tasks with explicit optimal decision rules (e.g., chess games), the reward patterns of model-selected scenarios have multi-dimensional coupling, making it difficult to directly derive the optimal solution. Therefore, this invention designs an expert knowledge acquisition method based on heuristic search. Through offline simulated environment interaction, expert model selection strategies are dynamically searched, and these strategies are stored in an expert experience pool in the form of <state-action-reward> triples for subsequent guidance of the agent's strategy training process (lines 6-8). Subsequently, the state s... t Input to the Actor network to obtain model selection and resource allocation actions (line 9). Receive action a. tAfterwards, the environment will execute the action and provide a reward and the next state, which will be stored in the local experience pool to update the model parameters (lines 10-11). Since resource allocation decisions have long-term effects across time slots, this invention introduces a discount factor γ to calculate the reward for the current task (line 12), defined as...
[0155]
[0156] To reduce the impact of environmental noise on gradient estimation, this invention introduces Generalized Advantage Estimation (GAE) as the objective of network updates (line 13), which is defined as follows:
[0157]
[0158] δ t =R t +γV φ (s t+1 )-V φ (s t ), (20)
[0159] Where λ represents the discount rate of the advantage function, δ t V represents temporal-difference (TD), where V is a state-value function.
[0160] Furthermore, to improve the stability of the advantage value during the policy update process, this invention controls the range of each update of the Actor network by pruning the objective function, which is defined as follows:
[0161]
[0162] Here, r(θ) represents the sampling rate between the old and new policies. Adjusting r(θ) can correct the sampling bias caused by policy updates, and it is defined as follows:
[0163]
[0164] Next, by minimizing L critic (φ) is used to optimize the Critic network, which is defined as follows:
[0165]
[0166] For the action space, its behavior cloning loss function is defined as follows:
[0167]
[0168] in, Indicates that the expert is in state s tThe following action, Indicates in s t Lower output The probability of.
[0169] By combining the behavioral cloning loss with the losses of the Actor and Critic networks in DRL (line 14), the total loss function is defined as
[0170] L total =L actor (θ)+L critic (φ)+λL BC (θ), (25)
[0171] Here, λ represents the weighting factor used to control the relative importance of BC and DRL. By controlling the change of λ, the exploration cost of the policy can be significantly reduced in the early stage of training, thereby accelerating model convergence. In the later stage of training, the model performance can be further optimized to generalize to other state spaces.
[0172]
[0173]
[0174] III. Experiment
[0175] The effectiveness and superiority of the proposed SDMS method were comprehensively evaluated through experiments.
[0176] 1. Experimental setup
[0177] Table 1. Detailed information on the models used for different service types.
[0178]
[0179] Based on a workstation equipped with an i5-12400F CPU@4.4GHz and an NVIDIA GeForce RTX 4070TiS, this invention uses Python to build the proposed system and utilizes PyTorch to implement the proposed SDMS and other comparative methods. This invention performs DNN inference tasks offline and collects the basic attributes of each sample task (including data size, number of pixels, and image entropy, etc.) and their inference accuracy and latency on different models, converting the inference time into inference density to be input into the state space of the DRL. Specifically, this invention uses ResNet to perform image classification tasks on the ImageNet dataset, YOLOv8 to perform object detection tasks on the COCO dataset, and DeepLabv3 to perform semantic segmentation tasks on the COCO dataset. Furthermore, this invention uses different models to perform DNN inference tasks, and the different model structures and their detailed information corresponding to each service type are shown in Table 1. In the experiments, this invention randomly extracts image IDs from the dataset and reads the corresponding task attributes for evaluation. Based on the above settings, the proposed system can be used to simulate the execution time and inference accuracy of DNN inference in real-world scenarios.
[0180] This invention uses a real cellular traffic dataset from Milan to simulate user service requests, which includes three types of services (i.e., messaging, calling, and internet). This dataset records traffic over a two-month period, with a sampling frequency of 10 minutes. This invention selects three regions (i.e.,
[0181] Traffic records with IDs 4259, 4456, and 5060 are used as the number of user requests for the three services within a time slot. User upload power is distributed in the range [80, 120] mW. The channel gain and Gaussian white noise power are 10⁻⁴ W and 10⁻⁸ W, respectively. For different tasks, their priorities are integers distributed between [1, 3], with maximum tolerable latency of [0.5, 0.7, 1.2] s and minimum tolerable precision of [50, 45, 45]%. The number of time slots is 100, the number of training rounds is 300, the service deployment interval is 10, the cloud transmission latency is 0.4 s, and the total available memory of the edge servers is 500 M. In the optimal strategy search process, the precision-latency weight coefficient is 0.05, and the pruning coefficient is 0.5.
[0182] To verify the superiority of the proposed SDMS method, this invention compares it with the following benchmark methods:
[0183] (1) PredRNN: This method uses RNN-based service traffic prediction and leverages a gating mechanism to capture dependencies in the traffic sequence. The model selection is consistent with the SDMS method.
[0184] (2) PPO-TO: Proximal Policy Optimization (PPO) is used for model selection, prioritizing the deployment of the model with the lowest latency and the highest accuracy for each service.
[0185] (3) DQNM: Deep Q-Network (DQN) is used for model selection, and the model with the lowest latency and the highest accuracy is deployed first for each service.
[0186] (4) Greedy-Delay: A greedy algorithm is used to select the model with the lowest latency for each task.
[0187] (5) Greedy-Acc: A greedy algorithm is used to select the model with the highest accuracy for each task.
[0188] 2. Experimental Results and Analysis
[0189] First, this invention evaluates the inference latency and accuracy of different YOLOv8 models on COCO dataset samples. For example... Figure 3 As shown, compared to the YOLOv8S model, the YOLOv8L model achieves higher inference latency on most samples. This is because the YOLOv8L model employs a more complex network structure, thus requiring more computational resources. However, the inference accuracy of the YOLOv8L model is not always higher than that of the YOLOv8S model; in fact, its inference accuracy is lower on some samples. This is because the YOLOv8L model struggles to effectively handle the bias between the training data distribution and the samples, leading to overfitting. It is noteworthy that the differences in both inference latency and inference accuracy between the YOLOv8L and YOLOv8S models continuously change across different samples, exhibiting strong dynamics. This prompts this invention to focus on how to achieve a better balance between inference latency and accuracy through real-time optimization of service deployment and model selection strategies.
[0190] Next, the present invention compares the convergence of the proposed SDMS with other methods. For example... Figure 4 As shown, the Greedy-Delay and Greedy-Acc methods employ static decision-making, lacking a learning process. These two methods perform worse than other dynamic decision-making methods. This is because their decision-making approach is singular and does not fully consider the differences between models.
[0191] Dynamic performance across different samples makes it difficult to effectively balance inference accuracy and latency. Compared to the DQN method, the PPO method achieves higher rewards and exhibits a smoother convergence curve. This is because the PPO method introduces a pruning operation to ensure that the difference between the old and new strategies is not too large, solving the problem of Q-value overestimation in the DQN method. By introducing traffic prediction and pruning search, the performance of the SDMS and PredRNN methods is further significantly improved. This indicates that accurate traffic prediction helps improve the rationality of service and model deployment in dynamically changing edge environments, thereby increasing task success rate under limited memory budgets. Compared to the PredRNN method, which uses RNNs for temporal prediction, the SDMS method, by introducing an Informer, can more efficiently capture long-term temporal dependencies, achieving better performance than PredRNN. Notably, compared to the PPO and DQN methods, the SDMS method, by combining with DIL, achieves faster and more stable convergence. The results validate the superiority of the SDMS method.
[0192] Subsequently, this invention compares the accuracy and latency satisfaction rate of the proposed SDMS with other methods under different services and models. For example... Figure 5 As shown, the SDMS method achieved the highest accuracy satisfaction rate for both image classification and instance segmentation services, while slightly lower than the Greedy-Acc method for object detection. This is because the Greedy-Acc method always selects the model with the highest accuracy, resulting in a high accuracy satisfaction rate across different services, but also leading to a longer inference latency. Similarly, the Greedy-Delay method exhibited the highest latency satisfaction rate across different services, but revealed the lowest accuracy satisfaction rate. Compared to other methods, the SDMS method achieves superior performance in both accuracy and latency satisfaction rate through a reasonable model deployment and configuration selection strategy. The results validate the effectiveness of the SDMS method in balancing model inference accuracy and latency.
[0193] To evaluate the contribution of each core component in the SDMS method to the overall performance, this invention removed the following core components through ablation experiments and then observed the changes in performance:
[0194] (1) TP (Traffic Prediction): Traffic prediction is not used; historical time slot traffic is used as an estimate of future time slot traffic.
[0195] (2) PS (Pruning Search): Instead of using pruning search, exhaustive search (ES) is used to obtain the best service deployment solution.
[0196] (3) BC (Behavior Cloning): No behavior cloning was used and no expert strategy was introduced to guide the training.
[0197] like Figure 6 As shown, the accuracy and latency satisfaction rate of the SDMS method are significantly affected when the TP component is not used. This is because without accurate traffic prediction, the SDMS method can only rely on historical traffic to infer future traffic, and the large deviation leads to performance degradation when facing load fluctuations. When ES is used to replace the PS component in SDMS, the accuracy and latency satisfaction rate are improved by about 0.5% and 0.9%, respectively. This is because ES obtains the optimal solution by traversing all possible solutions, but this also significantly increases the overhead of finding the optimal service deployment solution. In contrast, the pruning search scheme designed in this invention achieves performance comparable to exhaustive search with extremely low overhead and is more suitable for latency-sensitive inference systems. The most significant decrease in accuracy satisfaction rate occurs when the BC component is not used. By introducing expert policies to guide training, SDMS can learn expert behavior policies more quickly. When the BC component is removed, the SDMS method struggles to achieve efficient training. The results show that these three core components guarantee the performance of the SDMS method from different aspects, and their collaborative work ensures the high performance and strong adaptability of the SDMS method in dynamic and complex environments.
[0198] Furthermore, this invention analyzes the execution latency required by each core component and ES in the SDMS method to verify the efficiency of each core component in the SDMS method. As shown in Table 2, the execution latency of ES reaches 492.51ms. This is because ES needs to traverse all possible service deployment schemes, resulting in a huge amount of computation and excessive latency. In contrast, the proposed PS component can reduce the latency to 12.85ms by effectively filtering most of the lower-value schemes, which is only 2.6% of that of ES. At the same time, the additional latency required by TP and BC is only 9.65ms and 1.53ms, respectively. The above experiments verify the efficiency of the components designed in the SDMS method and do not impose too much additional burden on the system.
[0199] Table 2 shows the execution latency of each core component in the SDMS method and ES.
[0200]
[0201] Finally, this invention compares the service benefits of the proposed SDMS with other methods under different traffic volumes. Figure 7As shown, the service benefits of different methods increase significantly with the increase in traffic volume. When the traffic multiplier is 0.5x, the difference in service benefits between SDMS and other methods is small due to the small number of tasks. As the traffic multiplier increases, the difference in service benefits between different methods gradually widens. Specifically, SDMS and PredRNN methods introduce traffic prediction technology, and their service benefits are higher than other methods. This is because accurate traffic prediction can help to formulate a more reasonable service deployment plan, thereby achieving a more balanced model selection and improving system service benefits. PPO and DQN methods both adopt static service deployment strategies, making it difficult to make reasonable model selections, so their service benefits are lower than those of the SDMS method. Greedy-Delay and Greedy-Acc methods use a single model selection strategy and cannot perceive the impact of different request traffic on the service deployment status, so the difference in service benefits between them and the SDMS method gradually increases with the increase in traffic volume.
[0202] IV. Steps for using this application:
[0203] (1) The SDMS system continuously collects historical user request data from edge servers and uses a probabilistic sparse self-attention mechanism to predict traffic and accurately estimate the demand distribution of various services in the future, providing a basis for resource pre-allocation.
[0204] (2) Based on the prediction results, the edge service provider dynamically adjusts the model deployment strategy of the edge nodes, prioritizes the deployment of high-value model combinations within the memory limit, and regularly updates the deployment plan to adapt to changes in request traffic.
[0205] (3) When a user initiates an inference request, the system analyzes the task characteristics (including data volume, service type, priority, etc.) in real time, and combines them with the current resource status to select the optimal model configuration and bandwidth allocation scheme through deep reinforcement learning algorithms, so as to achieve the best balance between accuracy and latency.
[0206] (4) Edge nodes perform task processing according to system decisions, and perform local inference on service requests of deployed models; requests that are not deployed are intelligently routed to the cloud for processing, and additional transmission latency is recorded for subsequent optimization.
[0207] (5) During operation, the system continuously collects data such as task execution status, resource utilization and user feedback. Through imitation learning, it continuously optimizes the decision-making model, forms a closed-loop management mechanism of prediction-deployment-execution-optimization, and gradually improves the overall service quality.
Claims
1. A service deployment and model selection system for inference accuracy awareness in edge environments, characterized in that... The system includes a user terminal and an edge server. The user terminal includes a terminal device and an inference request module. The edge server includes a service type selection module and a model selection module. The system operates based on multiple time slots. In each time slot t∈T={1,2,...,T}, the user's inference task i is denoted as a quintuple <d i ,s i ,η i ,acc i ,ρ i >, where d i Indicates the task size, s i ∈{s1,s2,...,s J } represents the service type, η i ={η i,1 ,η i,2 ,...η i,K } represents the computational density under different configurations, acc i ={acc i,1 ,acc i,2 ,...acc i,k } represents the inference accuracy under different configurations, ρ i This indicates the task priority; the computational density and inference accuracy of a task will vary depending on the resource allocation.
2. The service deployment and model selection system for inference accuracy awareness in an edge environment according to claim 1, characterized in that, The task processing includes upload and execution phases. During the upload phase of task i, according to the Shannon channel capacity formula, its upload rate is defined as... Among them, b i Let p represent the bandwidth allocated to task i, and g represent the signal power. i σ represents the channel gain. 2 This represents the noise power during signal transmission. The upload latency of task i is defined as Once a task is uploaded to the edge server, the proposed system will select a model configuration for it to complete inference. If the model corresponding to the user's requested service is not deployed, the request will be forwarded to the remote cloud for processing. The latency of forwarding user task i from the edge server to the remote cloud is defined as... in, This represents the latency for forwarding a request to the remote cloud under configuration k corresponding to service type j; x j,k ∈{0,1} represents the model deployment decision; if the configuration k corresponding to the service type j required by task i is deployed on the edge server, x j,k =1; otherwise, x j,k =0; When performing DNN inference on an edge server, the execution latency of task i is defined as follows: Where f represents the computing power of the edge server; y i,k ∈{0,1} represents the model configuration decision assigned to task i; if task i chooses configuration k, y i,k =1; otherwise, y i,k =0; The total latency required to process task i is defined as T i total =T i up +T i tran +T i exe (5) For service requests initiated by users, the computational density and inference accuracy vary depending on the model configuration; based on the model configuration selected by the user, the inference accuracy of task i is expressed as follows:
3. The service deployment and model selection system for inference accuracy awareness in an edge environment according to claim 1, characterized in that, ESP generates revenue from the user when the total processing time and inference accuracy of the task meet the user's requirements; this process is defined as follows: in, and These represent the minimum tolerable precision and the maximum tolerable latency for service j, respectively. The higher the priority of a user's task, the greater the benefit ESP receives; within time slot t, the service benefit is defined as... Where m represents the number of users served by ESP within time slot t.
4. The service deployment and model selection system for inference accuracy awareness in an edge environment according to claim 1, characterized in that, Based on the proposed system model, the optimization objective is to maximize the long-term system benefit under system memory constraints; this optimization problem P1 is formally expressed as: Where, m j,k This represents the memory required for configuration k corresponding to the deployment service type j, and M is the total available memory in the edge environment.
5. A method for service deployment and model selection with inference accuracy awareness in edge environments, characterized in that, Based on the edge environment inference accuracy-aware service deployment and model selection system described in any one of claims 1-4, every T deploy In each time slot, the latest service deployment plan is obtained, the service deployment plan is executed, and the service deployment status is updated; the service traffic distribution of future time slots is obtained through a traffic prediction model based on a self-attention mechanism, and the optimal model deployment plan for each service is obtained through pruning search based on this. In each time slot t, a suitable model configuration and resource allocation scheme is selected for the inference task request, and the task completion status and service traffic are collected after the task is completed. The size of the task and the computational density attributes of the task under different configurations are used as state information to input the agent for model configuration decision-making. The behavior cloning loss is constructed by obtaining expert actions through offline search, and is combined with the online loss in the process of agent exploration and environment exploration to update the policy network, thereby accelerating the agent's fitting of the optimal model configuration strategy.
6. The service deployment and model selection method for inference accuracy awareness in an edge environment according to claim 5, characterized in that, Obtaining the latest service deployment plan, executing the service deployment plan, and updating the service deployment status are achieved through the following method: First, initialize the candidate service deployment plan p. best and its potential value v best And construct encoder input X en and decoder input X de , where X his X represents the historical inference request traffic. input X represents the traffic sequence collected in the current window. 0 This represents temporal sampling of the target sequence; probabilistic sparse self-attention is introduced when constructing the encoder, and self-attention distillation is used between layers to reduce computational overhead; feature extraction from layer j to layer j+1 is defined as... Where Convld(·) represents a one-dimensional convolution over a time series, ELU represents the activation function, MaxPool(·) represents the max pooling operation, [·] AB This indicates sparse self-attention; Adding a max-pooling layer after Conv1d reduces the input downsampling to half its original size, thereby reducing the computational complexity of the training process from... Reduce to Where L represents The length of the encoder is then determined. Next, the encoder output and decoder input are fed into the decoder, which contains one layer of multi-head sparse probabilistic self-attention mechanism and one layer of multi-head attention mechanism. Subsequently, the encoder output is fed into a multilayer perceptron (MLP) to predict the desired traffic sequence. Then, based on the traffic prediction results, all feasible service deployment schemes are iteratively searched. During the search, schemes with memory consumption exceeding the threshold M are first excluded because they exceed the constraint C2 in P1. Further, schemes with memory consumption below the threshold σM are excluded, ensuring that the memory consumption of all candidate deployment schemes lies within [σM, M], where σ is a pruning coefficient used to control the balance between search efficiency and deployment scheme overhead. If a deployment scheme has a lot of remaining memory, its value is further increased by deploying more models. For candidate service deployment schemes, the potential value of deploying each model is quantified, defined as... in, and Let δ represent the model's average accuracy and average latency on the test set, respectively; then, design a penalty term based on the latency of task transmission to the cloud; the potential value of a service deployment scheme is defined as... Where I(·) represents the indicator function; when service j has no deployed model, I(·) = 1; otherwise, I(·) = 0; Finally, based on the memory consumption and potential value of each model, an iterative search is performed, and the best service deployment scheme with the highest potential value is output.
7. The service deployment and model selection method for inference accuracy awareness in edge environments according to claim 5, characterized in that, By introducing DIL (Distributed Instruction) to train the DRL agent through learning from expert demonstrations, and introducing BC (Brainstorming) to help the DRL agent quickly learn expert knowledge in a specified domain, the state space, action space, and reward function are defined as follows during the interaction between the DRL agent and the environment: (1) State space: includes the service type of the task, data size, and optional model attributes; at the same time, pixel quantity and image entropy are added to the state space; in vision tasks, pixel quantity PC i Image entropy is the product of the image's width and height in pixels. It measures the amount of information and uncertainty in an image; higher image entropy indicates richer information. According to Shannon entropy, image entropy is defined as... Where, p h This represents the probability of the h-th grayscale pixel appearing; Therefore, the state space in time slot t is defined as s t ={d i ,s i ,r i ,PC i ,H i } (14) (2) Action Space; The goal of the DRL agent is to select an appropriate model configuration and bandwidth allocation for each inference request based on the current state space to achieve a good balance between inference accuracy and latency; therefore, the action space in time slot t is defined as to t ={and i,k ,b i } (15) (3) Reward function; First, the reward function needs to maximize the cumulative system revenue; Second, if the model selected by the DRL agent is not deployed, its task is rescheduled to the available model with the lowest accuracy, and a penalty term r is fed back. penalty The remaining available bandwidth is taken into account in the reward function, denoted as r. res Therefore, the reward function in time slot t is defined as follows: in, 8. The service deployment and model selection method for inference accuracy awareness in an edge environment according to claim 7, characterized in that, The appropriate model configuration and resource allocation scheme for the inference task request is selected by the following method: For each received inference task, construct the action state input s according to formula (14). t ; By interacting with the offline simulated environment, the expert model is dynamically searched to select strategies, and these strategies are stored in the expert experience pool in the form of <state-action-reward> triples for subsequent guidance of the agent's strategy training process. Then, state s t Input is fed into the Actor network to obtain model selection and resource allocation actions; upon receiving action a... t Afterwards, the environment will execute the action and provide a reward and the next state, which will be stored in the local experience pool for updating model parameters; a discount factor γ is introduced to calculate the reward for the current task, defined as follows: Generalized advantage estimation (GAE) is introduced as the objective of network updates, defined as follows: δ t =R t +γV φ (s t+1 )-V φ (s t ) (20) Where λ represents the discount rate of the advantage function, δ t This represents the timing difference (TD), where V is the state-value function; The range of each update of the Actor network is controlled by pruning the objective function, which is defined as follows: Where r(θ) represents the sampling rate between the old and new policies; by adjusting r(θ), the sampling bias caused by the policy update is corrected, which is defined as... Next, by minimizing L critic (φ) is used to optimize the Critic network, which is defined as follows: For the action space, its behavior cloning loss function is defined as follows: in, Indicates that the expert is in state s t The following action, Indicates in s t Lower output The probability of; By combining the behavioral cloning loss with the losses of the Actor and Critic networks in DRL, the total loss function is defined as follows: L total =L actor (θ)+L critic (φ)+λL BC (i) (25) Here, λ represents the weighting factor used to control the relative importance of BC and DRL; by controlling the change of λ, the exploration cost of the policy is significantly reduced in the early stage of training to accelerate model convergence; in the later stage of training, the model performance is further optimized to generalize to other state spaces.
Citation Information
Cited By
Geographic position driving-oriented edge-cloud collaborative server-free deep learning reasoning method
CN121413774A
Geo-location driven edge-cloud collaborative serverless deep learning inference method
CN121413774B