Edge service deployment and user distribution method based on dobby machine learning

Through the multi-arm slot machine learning method, the joint optimization model of edge service deployment and user allocation is built, which solves the service delay and cost problems caused by limited edge server resources, and realizes efficient service deployment and user allocation, and maximizes service effectiveness.

CN120186033APending Publication Date: 2025-06-20NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320531.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In an edge computing environment, edge server resources are limited, resulting in service latency optimization may lead to reduced resource utilization and increased costs, and it is difficult for the prior art to effectively deploy application service instances and allocate users to meet service quality and cost constraints.

Method used

A joint optimization mathematical model of edge service deployment and user allocation is constructed using a multi-arm slot machine learning method, and the development-utilization trade-off is optimized for edge service deployment and user allocation decisions to maximize service utility.

Benefits of technology

It realizes optimizing service latency and resource utilization in an edge computing environment, reducing costs, while ensuring high levels of service satisfaction and overall service utility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186033A_ABST
    Figure CN120186033A_ABST
Patent Text Reader

Abstract

The invention provides an edge service deployment and user distribution method based on dobby machine learning, which belongs to the technical field of mobile edge computing, and comprises the following steps: constructing an edge service deployment model and a service delay model; based on resource utilization and service satisfaction, a service utility quantitative representation model is constructed, a service utility maximization-oriented edge service deployment and user allocation joint optimization mathematical model is constructed aiming at user specific constraints in an edge computing environment, and aiming at problem characteristics, it is assumed that each edge server is regarded as an arm. Each user is regarded as an independent player, the problem is converted into a dobby machine problem, the problem is solved based on a dobby machine learning method, development-utilization tradeoff in the dobby machine problem is processed by adopting an upper confidence bound strategy, and an edge service deployment scheme and a user distribution scheme with theoretical performance guarantee are obtained. According to the invention, the overall service effectiveness of the service provider can be maximized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mobile edge computing, and in particular to a method for edge service deployment and user allocation based on multi-armed bandit learning. Background Art

[0002] Compared with cloud servers, edge servers have limited computing, storage, and communication resources, making the cost of resource provision high. In addition, each edge server can cover a relatively small geographical area, depending on the co-located wireless infrastructure. Therefore, multiple application service instances are required to serve geographically distributed users in a certain area. At the same time, under the constraints of resources and budgets, edge servers usually only allow a limited number of application service instances to be deployed. Deploying more application service instances and providing lower service latency can indeed help improve service satisfaction, but over-optimizing service latency may lead to reduced resource utilization and higher costs. In this case, service providers may not be able to achieve better service utility. This makes it crucial to jointly consider resource utilization and service satisfaction, which is not only related to the user experience, but also involves the competitiveness and long-term sustainability of application service providers in a highly competitive market. This requires considering the location of deploying application service instances and how to allocate users to edge servers with the required services, while also meeting a series of constraints such as user-specific service latency requirements, budget limitations, and resource limitations of edge servers, and establishing a more practical service provision strategy to improve resource utilization and reduce costs, and optimize the overall service utility. Unfortunately, most existing studies only focus on the edge service deployment problem, aiming to maximize user coverage under budget constraints, but may lead to waste of leased resources of edge servers due to the lack of an effective user allocation strategy. In addition, sharing the same application service instance by multiple users on the same edge server may lead to resource competition and communication interference, thus reducing their service satisfaction. Therefore, the service utility of service providers in terms of resource utilization and service satisfaction may be affected, and due to resource and budget limitations, it is difficult to meet the different needs of all users by deploying the corresponding requested application service instances on edge servers, ultimately resulting in the service provider may not be able to provide continuous and reliable services for its users, thereby reducing the service quality and damaging the revenue of the service provider. Summary of the Invention

[0003] To achieve the above objectives, the present invention adopts the following technical solutions:

[0004] Step S1: Considering the high competition for edge-dispersed limited resources among multi-users with heterogeneous requirements and the communication interference generated by sharing the same application service instance on the same edge server, construct an edge service deployment and service latency model; based on resource utilization and service satisfaction, construct a service utility quantification representation model.

[0005] Step S2: Considering the constraints such as the user-specific service latency requirements, edge service deployment budget, and edge server resources in the edge computing environment, construct a joint optimization mathematical model for edge service deployment and user allocation aiming at maximizing service utility.

[0006] Step S3: According to the problem characteristics, assume that each edge server is regarded as an arm and each user is regarded as an independent player, and transform the problem into a multi-armed bandit problem.

[0007] Step S4: Solve the problem based on the multi-armed bandit learning method, and adopt the upper confidence bound strategy to handle the exploration-exploitation trade-off in the multi-armed bandit problem, so as to obtain an edge service deployment scheme and a user allocation scheme with theoretical performance guarantees.

[0008] Furthermore, in Step S1, consider an edge computing system composed of n edge servers, denoted as Each edge server provides limited storage resources CPU frequency resources for computing and communication bandwidth resources to host and execute application service instances on demand.

[0009] Let be the set of services to be deployed on the edge servers of the system. The storage resources, CPU frequency resources for computing, and bandwidth resources required for deploying service a k are respectively denoted by Z k , F k and ω k . Generally, each service is associated with a deployment budget B k , indicating the maximum cost for deploying its instance.

[0010] Given the total number m of users with request applications, the user set is denoted as Each user u j can be specified by , where z j represents the data size related to the service request of u j , and represents the maximum service latency tolerated by u j . Assume that each user only requests one application service instance denoted by a j . Note that a user who requests multiple services can actually be converted into multiple users. Let be an indicator function, where represents that user u j requests service a k . Otherwise, Considering the constraints of edge server resources and application provider budgets, applications usually cannot be deployed on every edge server. Additionally, due to the interference between the locations of deployed application service instances and concurrent users, not all user service requests can be satisfied. Edge service deployment and user allocation model:

[0011] Let the application service instance deployment decision and user allocation decision be represented by binary variables x and y respectively, where

[0012]

[0013] Specifically, if application a k is deployed on edge server s i then x ik = 1; otherwise x ik = 0. Similarly, if user u j is allocated to edge server s i then y ij = 1; otherwise y ij = 0. Note that users can only be allocated to edge servers on which the application service instances they request are deployed. Therefore, if u j requests service a k then we have y ≤ x ij i.e.: ik That is:

[0014]

[0015] Given an original service, multiple instances are usually deployed to serve users. Therefore, for each service, its total deployment cost is where c ik is the deployment cost of deploying a i on s k The total cost incurred by deploying a k instances should not exceed the given deployment budget B k i.e.:

[0016]

[0017] When deploying multiple edge services to the same edge server, the following resource constraints should be met. The first is the storage resource constraint, the second is the computing resource constraint, and the third is the bandwidth resource constraint.

[0018]

[0019] Service latency model:

[0020] Given on edge server s iApplication a deployed on k , that is, x ik = 1, user u j Receives service a k , that is At edge server s i The computing latency Is calculated as follows:

[0021]

[0022] Where f k Is the computing intensity of service a k , Is the set of users who request service a k And are assigned to edge server s i , that is Is the number of users who request service a k And are assigned to edge server s i .

[0023] In a scenario considering interference constraints, user u j At edge server s i The maximum achievable transmission rate can be calculated as follows:

[0024]

[0025] Where p j Is the transmission power of user u j , a ij Is the channel gain between user u j And edge server s i , Is the noise power spectral density. Represents the cumulative interference of other users who are assigned to edge server s i And request the same service a k . z j Represents the size of the service data that user u j Wants to transmit. Then, user u j Receives the transmission latency of service from edge server s i Calculated as follows: Is calculated as follows:

[0026]

[0027] The service latency perceived by the user consists of the computing latency And the transmission latency These two parts.

[0028] Due to the limited coverage area of each edge server, users within its coverage area can directly access it through the base station covering a specific geographical area. Such an edge server is called the local edge server of the user, and the user can directly receive services through the zero-hop model. Let denote the set of local edge servers of user u j . dist(s i , u j ) represents the Euclidean distance between edge server s i and user u j , and radius(s i ) is the coverage radius of s i .

[0029] To increase the coverage area of the deployed application, users are allowed to receive services from other edge servers through a multi-hop model, that is, the local edge server acts as an intermediate relay. The edge server accessed by user u j through the multi-hop network connection is called the neighbor edge server of u j . In this case, for user u j to receive services from its neighbor edge server will increase the additional transmission delay to transmit it to its local edge server . Considering that the transmission delay is proportional to the number of hops between edge servers s i′ and s i , their service transmission paths are simulated by the shortest connection path among them. Let τ be the transmission delay per hop in the path, which is related to the minimum bandwidth and data transmission information in the path, and the network propagation delay can be ignored in the real-world edge computing environment. Therefore, the service delay L j for user u i to receive services from edge server s ij can be calculated as follows:

[0030]

[0031] where H ij is the number of hops between user u j and edge server s i . If s i is the local edge server of u j , that is then H ij = 0. Otherwise, H ij is the minimum number of hops from the local edge server of u j to s i , because u j may have multiple local edge servers, that is

[0032] Service utility quantification representation model:

[0033] Let ρ ik represent the CPU utilization rate of the computing resources of service a k on edge server s i and can be expressed as:

[0034]

[0035] where the parameter ζ k ∈(0.9, 1) is determined based on the size of the computing data required by a k and the value of ρ ik increases with the number of users i sharing the computing resources of a k on s However, even if continues to increase, the CPU utilization efficiency can converge to a certain value. Note that any user must satisfy the condition

[0036] The service satisfaction and service delay have an S-shaped curve relationship, and it rapidly increases and tends to converge as the service delay decreases. Therefore, given the maximum acceptable delay j of user u the service satisfaction π j perceived by user u i when receiving the service from edge server s ij can be expressed as follows:

[0037]

[0038] where α is the maximum service satisfaction and β is a constant that controls the growth rate of the user service satisfaction level. The term represents the service satisfaction sensitivity related to the service delay. We set α = 2, so the value range of π ij is from 0 to 1, that is, π ij ∈[0, 1].

[0039] For each service a k its service utility is defined as a combination of resource utilization and service satisfaction, and we can get:

[0040]

[0041] where γ > 0 is an empirically chosen coefficient. It is used to normalize and balance resource utilization and service satisfaction according to the preferences of application providers. By adjusting γ, we can find an appropriate balance between optimizing resource utilization and ensuring high service satisfaction according to the priorities of application providers. Generally, a large γ gives priority to resource utilization rather than service satisfaction to ensure the service utility of application providers, and vice versa. However, the value of γ may need to be further fine-tuned according to empirical data and specific application requirements.

[0042] Furthermore, in step S2, considering the specific service latency requirements of users, the deployment budget of services, and the resource limitations of edge servers, as well as the edge service deployment budget constraint, by simultaneously solving the two sub-problems of application deployment and user allocation in the edge computing environment, the problem of maximizing service utility can be formulated as:

[0043]

[0044] Furthermore, in step S3, according to the characteristics of problem P1, it is transformed. Recall the problem P1 we proposed. Our goal is to maximize the overall service utility, which is the cumulative utility obtained by allocating users to edge servers to meet their requirements. Therefore, problem P1 is transformed into a multi-armed bandit problem. By regarding edge servers as arms and assuming that each user is an independent participant, the widely used upper confidence bound strategy is adopted to solve the exploration-exploitation trade-off in the multi-armed bandit problem.

[0045] User allocation is only performed after edge service deployment. It can be inferred that only when any corresponding user is allocated to an edge server, the requested service will be deployed on the server, that is, for each service x ik = 1 if and only if and Therefore, to achieve problem transformation, we can use the user allocation decision y ij to represent the corresponding edge service deployment decision x ik , that is:

[0046]

[0047] At this time, recalling the service utility By simplifying problem P1 into an equivalent problem that only involves variable y, problem P2 can be reconstructed as:

[0048]

[0049]

[0050] Further, in step S4, the UMESP algorithm is adopted to solve P2, and an edge service deployment plan and a user allocation plan are obtained. The key idea of the algorithm is that, under the constraints of the deployment budget and edge server resources, it iteratively selects a candidate edge server with the largest upper confidence bound from all in to determine a new user allocation, and the algorithm iteratively selects the most cost-effective application deployment plan.

[0051] Let be the set of users who request service a k and select edge server s i in trial t. Let k be the instantaneous overall service utility of service a i deployed on edge server s

[0052]

[0053] where is the set of users selected from with satisfactory service latency, that is,

[0054] Given user u k who requests application a j (that is, ) in the t-th trial and selects edge server s i , its instantaneous allocation reward r ij (t) can be expressed as follows:

[0055]

[0056] where represents the cumulative instantaneous overall service utility of service a j when, in the t-th trial, all users except user u i are allocated to edge server s k and request service a k , that is, all users in . A larger r ij (t) indicates that user u j is more likely to be allocated to edge server s i , and vice versa. In each trial, the instantaneous allocation reward of the user is dynamically updated according to the actual feedback, reflecting the long-term selection trend of the user. Through the continuous selection and update process, the optimal allocation between the user and the edge server is optimized, and finally the overall service utility is maximized.

[0057] Given a set of requests for service a kand assigned to the optimal edge server For users, it is easy to obtain that these users are assigned to The overall service utility generated will not be lower than that obtained by assigning them to any other edge server s i In That is always holds. In addition, the overall service utility is jointly determined by the assigned users. Therefore, for each user Assigning it to s i The immediate assignment reward generated is not less than that obtained by assigning it to any edge server s i Generated, that is, r i*j (t) ≥ r ij (t). Therefore, for each service a k , maximizing its overall service utility is equivalent to maximizing its immediate assignment reward by assigning its users to edge servers.

[0058] Latency violation penalty. The perceived latency of any user cannot exceed its maximum tolerable service latency. Therefore, when user u j Selects edge server s i , it should be checked whether the service latency L j Perceived by user u ij (t) meets its maximum tolerable service latency If this constraint is violated, then π ij = 0, because L ij (t) > This means that it is not allowed to assign user u j To edge server s i , that is, y ij (t) = 0. In this case, these users cannot obtain the immediate assignment reward, that is, for each user We set r ij (t) = 0. On the contrary, if That is, π ij > 0, then in trial t, user u j Is assigned to edge server s i Is valid, that is, there is y ij (t) = 1. At the same time, the immediate assignment reward of these users can be obtained.

[0059] Resource and budget violation penalty. When user u j Selects edge server s i In trial t, and Holds, this means that the requested service Will be deployed on s i , that is, x ik= 1. Meanwhile, if and only if at this time x ik = 1. Due to the limitations of edge server resources and edge service deployment budgets, we need to check whether the current edge service deployment decisions for each application satisfy these constraints. To better reflect the immediate allocation reward r j obtained by user u i when selecting edge server s ij (t), we use the cost - benefit marginal utility score to evaluate the reward. This score is defined as the increment of the overall service utility of a j generated by allocating u i to edge server s k . The cost - benefit utility score is defined as to measure the cost - benefit of each application deployment plan.

[0060] The service requests of users within a certain period are random, and the global information of the current state of the edge server cannot be observed. That is, the highest immediate allocation reward that can be obtained by allocating a user to any edge server is unknown. We use historical observations to estimate the expected allocation reward when the user is allocated in the next trial. Denote as the total number of times user u j selects edge server s i in trial t, and denote as the corresponding sample average allocation reward. They are updated as follows:

[0061]

[0062] Generally, we use the sample average allocation reward to represent the estimate of the expected allocation reward when user u j selects edge server s i in the next trial. In addition, the immediate allocation reward r j of each user u ij (t) is used to update its corresponding average allocation reward. By calculating the expected allocation reward, we can evaluate the potential allocation reward that a user may obtain from each edge server, thereby gradually increasing the probability of selecting the optimal edge server.

[0063] The key challenge of the multi - armed bandit problem is to determine whether to continue exploring new edge servers or to exploit the known optimal choices for user allocation. We adopt the UCB strategy to balance exploration and exploitation in each trial to guide the selection of edge servers. For each user u j in the t - th trial, the edge server selected according to the UCB strategy can be expressed as follows:

[0064]

[0065] The first term is the average allocation reward for user u j to select the edge server s in experiment t i , and the expected value is determined using historical information based on past rewards. The second term is the exploration term. It represents user u j selecting s in experiment t i 's confidence radius, which helps to explore new possibilities. By following the upper confidence bound strategy, we can effectively manage the exploration - exploitation trade - off and jointly decide which edge server to select. This strategy enables us to identify the optimal edge server with the highest allocation reward after a given series of experiments.

[0066] At the end of experiment T, for each user u j , by selecting the edge server that provides the maximum immediate allocation reward to obtain its final allocation decision y i*j (T)=1. At the same time, determine the deployment decision of the corresponding service a k , that is, x i*j (T)=1.

[0067] The specific process of the algorithm is as follows:

[0068] Step S4 - 1: Starting from initialization, the algorithm sets the application service instance deployment decision variable x and the user allocation decision variable y as empty sets. At t = 0 of the experiment, for each user and each server , initialize by normalizing the geographical distance between the user and each s i and update the corresponding to 1.

[0069] Step S4 - 2: Perform T iterations of calculation. At the start of each iteration of calculation, update as an empty set to store the users who request service a k in the t - th experiment and select the edge server s i , and update as an empty set to store the overall service utility of service a k deployed on the edge server s i in the t - th experiment.

[0070] Next, select the edge server with the highest upper confidence bound for all users and increment the corresponding total number of selections . Add the users who request service a k and select the edge server s i in experiment t to the set​ in

[0071] Step S4-3: Next, initialize an empty set to store the set of users selected from that satisfy the service delay constraint. Calculate the service delay L ij perceived by each user in the set to all servers, and check whether it satisfies its allowed maximum service delay If it satisfies, add the user to the set; if not, set its corresponding immediate allocation reward r ij (t) to 0. Then, according to the obtained set If it is not an empty set, calculate the service delay L ij (t) and service satisfaction π ij (t) of each user, and its corresponding service utility If it is an empty set, then set its corresponding service utility to 0.

[0072] Then, we sort in descending order by the utility scores of all applications on each edge server and continuously select the service-edge server pair with the maximum score until the resources and budget reach the limit. For all unselected deployment schemes x ik (t) = 1, we update the corresponding and x ik (t) = 0. At the same time, for all relevant users, their corresponding allocation decisions y ij (t), service satisfaction π ij (t) and immediate allocation reward r ij (t) are all set to 0. Then, update the expected allocation reward of each user to each server

[0073] Step S4-4: After T iterations, obtain the final allocation decision y j of user u i*j (t) ← 1, where for u j with the maximum immediate allocation reward, and obtain the edge service deployment decision x i*j (t) ← 1 according to the user allocation decision.

[0074] Step S4-5: Finally, return x and y as the final edge service deployment and user allocation schemes.

[0075] Regarding the prior art, the beneficial effects of the present invention are as follows: In the edge computing environment, deploying more application service instances by edge servers to optimize service latency may lead to waste of leased resources and higher costs. Considering the scenarios of limited heterogeneous resources of edge servers, heterogeneous service latency requirements of users, and budget constraints for edge service deployment, by jointly optimizing the two sub-problems of edge service deployment and user allocation considering resource utilization and service satisfaction, it aims to maximize the overall service utility of service providers, ensuring limited resources while maintaining a high level of service satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 It is a specific flowchart in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.

[0078] Step S1: Considering the high competition for edge-dispersed limited resources by heterogeneous requirements of multiple users and the communication interference generated by sharing application service instances on the same edge server, construct an edge service deployment and service latency model; based on resource utilization and service satisfaction, construct a service utility quantification representation model.

[0079] Consider how to deploy various types of services and allocate users on edge servers with limited resources to meet the service requirements of users. Consider an edge computing system composed of n edge servers, denoted as Each edge server provides limited storage resources CPU frequency resources for computing and communication bandwidth resources to host and execute application service instances on demand.

[0080] Let be the set of services to be deployed on the edge servers of the system. Deploying service a k The required storage resources, CPU frequency resources for computing, and bandwidth resources are respectively denoted by z k , F k and ω k Generally speaking, each service is associated with a deployment budget B k indicating the maximum cost for deploying its instance.

[0081] Given the total number m of users with request applications, the user set is represented as Each user u j can be specified by where zj Denote the data size related to the service request of u j and the maximum service latency tolerated by u Denote the maximum service latency tolerated by u j . Assume that each user only requests one application service instance denoted by a j . Note that a user who requests multiple services can actually be converted into multiple users. Let be an indicator function, where denotes that user u j requests service a k . Otherwise,

[0082] Edge service deployment and user assignment model:

[0083] Denote the application service instance deployment decision and the user assignment decision by binary variables x and y respectively, where

[0084]

[0085] Specifically, if application a k is deployed on edge server s i , then x ik = 1; otherwise x ik = 0. Similarly, if user u j is assigned to edge server s i , then y ij = 1; otherwise y ij = 0. Note that a user can only be assigned to an edge server on which the application service instance they request is deployed. Therefore, if u j requests service a k , that is then we have y ij ≤ x ik , that is

[0086]

[0087] Given an original service, multiple instances are usually deployed to provide services for users. Therefore, for each service, its total deployment cost is where c ik is the deployment cost of deploying a i on s k . The total cost generated by deploying a k instances should not exceed the given deployment budget Bk, that is

[0088]

[0089] When deploying multiple edge services to the same edge server, the following resource constraints should be met. The first is the storage resource constraint, the second is the computing resource constraint, and the third is the bandwidth resource constraint.

[0090]

[0091] Service latency model:

[0092] Given the application a i deployed on the edge server s k , that is, x ik = 1, the user u j receives the service a k , that is the computing latency of the edge server s i is calculated as follows:

[0093]

[0094] where f k is the computing intensity of the service a k , is the set of users who request the service a k and are assigned to the edge server s i , that is is the number of users who request the service a k and are assigned to the edge server s i .

[0095] In a scenario considering interference constraints, the maximum achievable transmission rate of the user u j on the edge server s i can be calculated as follows:

[0096]

[0097] where p j is the transmission power of the user u j , g ij is the channel gain between the user u j and the edge server s i , is the noise power spectral density. represents the cumulative interference of other users who are assigned to the edge server s i and request the same service a k . z j represents the size of the service data to be transmitted by the user u j . Then, the transmission latency j for the user u i to receive the service from the edge server s is calculated as follows:​

[0098]

[0099] The service latency perceived by the user consists of the computing latency and the transmission latency These two parts.

[0100] Since the coverage area of each edge server is limited, users within its coverage area can directly access it through a base station covering a specific geographical area. Such an edge server is called the local edge server of the user, and the user can directly receive services through the zero-hop model. Let denote the set of local edge servers of user u j , dist(s i , u j ) represents the Euclidean distance between edge server s i and user u j , and radius(s i ) is the coverage radius of s i .

[0101] To increase the coverage area of the deployed application, users are allowed to receive services from other edge servers through a multi-hop model, that is, the local edge server acts as an intermediate relay. The edge server accessed by user u j through the multi-hop network connection is called the neighbor edge server of u j . In this case, for user u j receiving services from its neighbor edge server will increase the additional transmission latency to transmit it to its local edge server . Considering that the transmission latency is proportional to the number of hops between edge servers s i′ and s i , their service transmission paths are simulated using the shortest connection path among them. Let τ be the transmission latency per hop in the path, which is related to the minimum bandwidth and data transmission information in the path, and the network propagation latency can be ignored in the real-world edge computing environment. Therefore, the service latency L j for user u i to receive services from edge server s ij can be calculated as follows:

[0102]

[0103] where H ij is the number of hops between user u j and edge server s i . If s i is the local edge server of u j , that is Then H ij = 0. Otherwise, H ij is the minimum number of hops from the local edge server of u j to s i because u j may have multiple local edge servers, that is

[0104] Service utility quantization representation model:

[0105] Let ρ ik represent the CPU utilization rate of the computing resources of service a k on the edge server s i and can be expressed as:

[0106]

[0107] where the parameter ζ k ∈(0.9, 1) is determined based on the size of the computing data required by a k , and the value of ρ ik increases with the increase in the number of users i sharing the computing resources of a k on s . However, even if continues to increase, the CPU utilization efficiency can converge to a certain value. Note that any user must satisfy the condition

[0108] The service satisfaction has an S-shaped curve relationship with the service delay and increases rapidly and converges as the service delay decreases. Therefore, given the maximum acceptable delay j of user u the service satisfaction π j perceived by user u i when receiving the service from the edge server s ij can be expressed as follows:

[0109]

[0110] where α is the maximum service satisfaction and β is a constant that controls the growth rate of the user service satisfaction level. The term represents the service satisfaction sensitivity related to the service delay. We set α = 2, so the value range of π ij is from 0 to 1, that is, π ij ∈[0, 1].

[0111] For each service a k , its service utility Defined as a combination of resource utilization and service satisfaction, we can obtain:

[0112]

[0113] where γ > 0 is a coefficient selected empirically. It is used to normalize and balance resource utilization and service satisfaction according to the preferences of application providers. By adjusting γ, we can find an appropriate balance between optimizing resource utilization and ensuring high service satisfaction according to the priorities of application providers. Generally, a large γ gives priority to resource utilization rather than service satisfaction to ensure the service utility of application providers, and vice versa. However, the value of γ may need to be further fine-tuned according to empirical data and specific application requirements.

[0114] Step S2: For the constraints such as the user-specific service latency requirements, edge service deployment budget, and edge server resources in the edge computing environment, construct a joint optimization mathematical model for edge service deployment and user allocation aiming at maximizing service utility.

[0115] Considering simultaneously meeting the user's specific service latency requirements, the deployment budget of the service, and the resource limitations of the edge server, as well as the edge service deployment budget constraint, by considering solving two sub-problems of application deployment and user allocation in the edge computing environment simultaneously, the problem of maximizing service utility can be constructed as:

[0116]

[0117] Step S3: For the characteristics of this problem, assume that each edge server is regarded as an arm and each user is regarded as an independent player, and transform this problem into a multi-armed bandit problem.

[0118] Recall the problem P1 we proposed. Our goal is to maximize the overall service utility, which is the cumulative utility obtained by allocating users to edge servers to meet their requirements. Therefore, transform the problem P1 into a multi-armed bandit problem by regarding edge servers as arms, assuming that each user is an independent participant, and using the widely used upper confidence bound strategy to solve the exploration-exploitation trade-off in the multi-armed bandit problem.

[0119] User allocation is only performed after edge service deployment. It can be inferred that the requested service will only be deployed on the server when any corresponding user is allocated to the edge server, that is, for each service x ik = 1 if and only if and Therefore, to achieve problem transformation, we can use the user allocation decision y ij to represent the corresponding edge service deployment decision x ik, that is:

[0120]

[0121] At this time, review the service utility Simplify problem P1 into an equivalent problem that only involves variable y, and problem P2 can be reconstructed as:

[0122]

[0123] Step S4: Based on the multi-armed bandit learning method, solve the problem and adopt the upper confidence bound strategy to handle the exploration-exploitation trade-off in the multi-armed bandit problem, and obtain an edge service deployment plan and a user allocation plan with theoretical performance guarantees.

[0124] Solve P2 based on the multi-armed bandit learning method to obtain an edge service deployment plan and a user allocation plan. The key idea of the algorithm is that, under the constraints of the deployment budget and edge server resources, it iteratively selects a candidate edge server with the maximum upper confidence bound from all to determine a new user allocation.

[0125] Let be the set of users who request service a k and select edge server s i in trial t. Let k be the immediate overall service utility of service a i deployed on edge server s

[0126]

[0127] where is the set of users with satisfactory service latency selected from , that is

[0128] Given user u k who requests application a j (that is, ) in the t-th trial, and select edge server s i , its immediate allocation reward r ij (t) can be expressed as follows:

[0129]

[0130] where represents that in the t-th trial, except for user u j , all other users are allocated to edge server s i and request service ak At this time, Service A k The cumulative immediate overall service utility, that is, all users within A larger r ij (t) indicates that user u j Is more likely to be assigned to the edge server s i , and vice versa. In each trial, the immediate allocation reward of the user is dynamically updated according to the actual feedback, reflecting the long-term selection trend of the user. Through the continuous selection and update process, the optimal allocation between the user and the edge server is optimized, ultimately maximizing the overall service utility.

[0131] Given a set of users requesting Service A k And assigned to the optimal edge server The overall service utility generated by assigning these users to Will not be lower than assigning them to any other edge server s i In , that is Always holds. In addition, the overall service utility is jointly determined by the assigned users. Therefore, for each user Assigning it to s i The immediate allocation reward generated is not less than that generated by assigning it to any edge server s i , that is, r i*j (t) ≥ r ij (t). Therefore, for each Service A k , maximizing its overall service utility is equivalent to maximizing its immediate allocation reward by assigning its users to the edge server.

[0132] Latency violation penalty. The perceived latency of any user cannot exceed its maximum tolerable service latency. Therefore, when user u j Selects the edge server s i , it should be checked whether the perceived service latency L j (t) of user u ij Meets its maximum tolerable service latency If this constraint is violated, then π ij = 0, because This means that it is not allowed to assign user u j To the edge server s i , that is, y ij (t) = 0. In this case, these users cannot obtain the immediate allocation reward, that is, for each user We set r ij (t) = 0. On the contrary, if That is, π ij> 0, then in trial t, user u j is assigned to edge server s i is valid, that is, there is y ij (t) = 1. At the same time, the immediate allocation reward for these users can be obtained.

[0133] Resource and budget violation penalty. When user u j selects edge server s i in trial t, and holds, this means that the requested service will be deployed on s i , that is, x ik = 1. At the same time, if and only if then x ik = 1. Due to the limitations of edge server resources and edge service deployment budgets, we need to check whether the current edge service deployment decisions for each application meet these constraints. To better reflect the immediate allocation reward r j obtained by user u i when selecting edge server s ij (t), we use the cost - benefit marginal utility score to evaluate the reward. This score is defined as the increment of the overall service utility of a j generated by assigning u i to edge server s k . The cost - benefit utility score is defined as to measure the cost - benefit of each application deployment scenario.

[0134] The service requests of users within a certain period are random, and the global information of the current state of edge servers cannot be observed. That is, the highest immediate allocation reward that may be obtained by assigning a user to any edge server is unknown. We use historical observations to estimate the expected allocation reward when the user is assigned in the next trial. Let be denoted as the total number of times user u j selects edge server s i in trial t, and let be denoted as the corresponding sample average allocation reward. They are updated as follows:

[0135]

[0136] Generally, we use the sample average allocation reward to represent the estimate of the expected allocation reward when user u j selects edge server s i in the next trial. In addition, the immediate allocation reward r j for each user u ij(t) is used to update its corresponding average allocated reward. By calculating the expected allocated reward, we can evaluate the potential allocated reward that a user may obtain from each edge server, thereby gradually increasing the probability of selecting the optimal edge server.

[0137] The key challenge in the multi-armed bandit problem is to determine whether to continue exploring new edge servers or to exploit the known optimal choices for user allocation. By adopting the UCB strategy to balance exploration and exploitation in each trial to guide the selection of edge servers. For each user u in the t-th trial j , the edge server selected according to the UCB strategy can be expressed as follows:

[0138]

[0139] where the first term is the average allocated reward for user u j in trial t for selecting edge server s i , and the expected value is determined based on historical information according to past rewards. The second term is the exploration term. It represents the confidence radius for user u j to select s i in trial t, which helps to explore new possibilities. By following the upper confidence bound strategy, we can effectively manage the trade-off between exploration and exploitation and jointly decide which edge server to select. This strategy enables us to identify the optimal edge server with the highest allocated reward after a given series of trials.

[0140] At the end of trial T, for each user u j , its final allocation decision y is obtained by selecting the edge server that provides the maximum immediate allocated reward i*j (T) = 1. At the same time, the deployment decision of the corresponding service a k , i.e., x i*j (T) = 1, is determined.

[0141] The algorithm process is as follows:

[0142] Step S4-1: Starting from initialization, the algorithm sets the service instance deployment decision variable c and the user allocation decision variable y as empty sets. At t = 0, for each user and each server , initialize i by normalizing the geographical distance between the user and each s and update the corresponding to 1.

[0143] Step S4-2: Perform T iterations of calculation. At the start of each iteration of calculation, update is an empty set to store the service a for the t-th trial request k and select the edge server s i of the user, and update is an empty set to store the service a in the t-th trial k deployed on the edge server s i of the overall service utility.

[0144] Next, for all users, select the edge server with the highest upper confidence bound, and increment the corresponding total number of selections by one. Add the users who request service a in trial t k and select the edge server s i to the set .

[0145] Step S4-3: Initialize an empty set to store the set of users selected from that satisfy the service delay constraint. Calculate the service delay L perceived by each user in the set to all servers ij , and check whether it satisfies its maximum allowed service delay If it satisfies, add the user to the set; if not, set its corresponding immediate allocation reward r ij (t) to 0. Then, according to the obtained set If it is not an empty set, calculate the service delay L ij (t), service satisfaction π ij (t), and its corresponding service utility If it is an empty set, then set its corresponding service utility to 0.

[0146] Next, sort in descending order by the utility scores of all applications on each edge server , and continuously select the service-edge server pair with the maximum score until the resources and budget reach the limit. For all unselected deployment plans x ik (t) = 1, update its corresponding and x ik (t) = 0. At the same time, for the allocation decisions y ij (t), service satisfaction π ij (t), and immediate allocation reward r ij (t) of all relevant users are all set to 0.

[0147] Then, update each user to each server Expected allocated reward

[0148] Step S4-4: After performing T iterations, obtain the final allocation decision y of user u j i*j (t) ← 1, where For u j with the maximum immediate allocation reward, obtain the edge service deployment decision x based on the user allocation decision i*j (t) ← 1.

[0149] Step S4-5: Finally, return x and y as the final edge service deployment and user allocation solutions.

[0150] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modification or change made by those of ordinary skill in the art according to the disclosed content of the present invention shall be included in the protection scope recorded in the claims.​

Claims

1. An edge service deployment and user allocation method based on multi-armed bandit learning, characterized in that: The method comprises the following steps: Step S1: construct an edge service deployment model and a service delay model; based on resource utilization and service satisfaction, construct a service utility quantitative representation model; Step S2: Based on the user-specific constraints in the edge computing environment, a joint optimization mathematical model of edge service deployment and user allocation for maximizing service utility is constructed; Step S3: According to the characteristics of the problem, assume that each edge server is regarded as an arm and each user is regarded as an independent player, and transform the problem into a multi-armed bandit problem; Step S4: Solve the problem based on the multi-armed bandit learning method, and adopt the upper confidence bound strategy to deal with the development-utilization trade-off in the multi-armed bandit problem, and obtain the edge service deployment plan and user allocation plan with theoretical performance guarantee.

2. According to claim 1, a method for edge service deployment and user allocation based on multi-armed bandit learning is characterized in that: In step S1, the edge service deployment model is defined: Consider an edge computing system consisting of n edge servers, denoted as Each edge server Provides limited storage resources Calculated CPU frequency resources and communication bandwidth resources Host and execute application service instances on demand, is a set of services to be deployed on the edge server of the system. Deploy service a on the edge server k The required storage resources, CPU frequency resources and bandwidth resources are represented by Z k 、F k and ω k Indicates that each service Both with deployment budget B k is associated with, indicating the maximum cost of deploying its instance. Given the total number of users m requesting applications, the user set is represented as Each user u j Can be Specify, where z j Indicates that j The size of the data associated with the service request, Indicates u j The maximum service delay tolerated, Edge service deployment and user allocation model: The application service instance deployment decision and user allocation decision are represented by binary variables x and y, respectively. If you apply a k Deployed on edge servers i Up, that is, x ik =1; otherwise x ik = 0, similarly, if user u j Assigned to edge servers i ,y ij =1; otherwise y ij = 0, the user can only be assigned to the edge server where the application service instance he requested is deployed. Therefore, if u j Request service a k ,Right now Then there is y ij ≤x ik ,Right now Given an original service, multiple instances are usually deployed to provide services to users. Therefore, for each service, the overall deployment cost is where c ik is in i Deploy a k The deployment cost of deploying a k The total cost incurred by the instance should not exceed the given deployment budget B k ,Right now When multiple edge services are deployed on the same edge server, the following resource constraints should be met: the first is storage resource constraints, the second is computing resource constraints, and the third is bandwidth resource constraints.

3. The edge service deployment and user allocation method based on multi-armed bandit learning according to claim 1 is characterized in that: In step S1, the service delay model is defined: Given an edge server s i Application deployed on k , that is, x ik =1, user u j Receiving servicea k ,Right now On edge servers i The calculation delay The calculation is as follows: where f k It is a service k The computational intensity, Is a request service k and are distributed to edge servers i The user set, that is Is a request service k and are distributed to edge servers i The number of users, In the scenario of considering interference limitation, user u j On edge servers i The maximum achievable transfer rate can be calculated as follows: where p j Is user u j Transmission power, g ij Is user u j and edge servers i The channel gain between is the noise power spectral density, Indicates that it is assigned to the edge server s i and request the same service a k The cumulative interference of other users, z j Represents user u j The size of the service data to be transferred, then, user u j From the edge servers i Transmission delays for receiving services The calculation is as follows: The user-perceived service delay is determined by the computational delay and transmission delay These two parts consist of make Represents user u j The local edge server set, dist(s i ,u j ) represents the edge server s i and user u j The Euclidean distance between i ) is i The coverage radius, User j From the edge servers i Service delay L for receiving service ij , calculated as follows: Among them, H ij Is user u j and edge servers i The number of hops between them, if s i is u j The local edge server, i.e. Then H ij =0, otherwise, H ij is u j Local edge server to s i The minimum number of hops, because u j There may be multiple local edge servers, i.e.

4. The edge service deployment and user allocation method based on multi-armed bandit learning according to claim 1 is characterized in that: In step S1, the service utility quantification representation model is defined: Let ρ ik Indicates service a k On edge servers i The CPU utilization of the computing resources on is expressed as: Among them, the parameter ζ k ∈(0.9,1) based on a k The size of the required computational data is determined by ρ ik The value of i Share on k The number of users of computing resources increases with the increase of Given user u j The maximum acceptable delay User j From the edge servers i Service satisfaction perceived by the recipient ij It is expressed as follows: Among them, α is the maximum service satisfaction, β is a constant that controls the growth rate of user service satisfaction level, and represents the sensitivity of service satisfaction related to service delay, assuming α = 2, so π ij The value range is from 0 to 1, which is π ij ∈[0,1], For each service a k , its service utility Defined as a combination of resource utilization and service satisfaction, we get:

5. The edge service deployment and user allocation method based on multi-armed bandit learning according to claim 1 is characterized in that: In step S2, Considering the user's specific service delay requirements, the service deployment budget and the resource constraints of the edge server, as well as the edge service deployment budget constraints, the problem of maximizing service utility is constructed by solving the two sub-problems of application deployment and user allocation in the edge computing environment at the same time:

6. The edge service deployment and user allocation method based on multi-armed bandit learning according to claim 1 is characterized in that: In step S3, For each service If and only if and Therefore, in order to achieve problem transformation, the user allocation decision y ij To represent the corresponding edge service deployment decision x ik ,Right now: At this point, review the service utility Simplifying problem P1 into an equivalent problem involving only variable y, problem P2 is reconstructed as:

7. The edge service deployment and user allocation method based on multi-armed bandit learning according to claim 1 is characterized in that: In step S4, set up To request service a in trial t k and select edge servers i A collection of users, To serve a in test t k Deployed on edge servers i The immediate overall service utility is calculated as follows: in is from The set of users with satisfactory service delay selected from Given the request application a in the tth trial k User u j ,Right now and select edge servers i , whose immediate distribution reward r ij (t) is expressed as follows: in Indicates that in the tth trial, except for user u j In addition, all users are assigned to edge servers s i and request service a k When serving a k The immediate overall service utility accumulation, Delay violation penalty: any user’s perceived delay cannot exceed its maximum tolerable service delay. j Select Edge Servers i When the user u j Perceived service delay L ij (t) Whether the maximum tolerable service delay is met If this constraint is violated, then π ij =0, because This means that user u is not allowed to j Assign to edge servers i , that is, y ij (t) = 0. In this case, these users cannot obtain instant distribution rewards, that is, for each user Setting ij (t) = 0, on the contrary, if That is π ij > 0, then in experiment t, user u j Assigned to edge servers i is valid, that is, there is y ij (t) = 1, and at the same time, these users can get instant rewards; Resource and budget violation penalties, when the user u j Select edge server s in experiment t i ,and When established, this means that the service it requests Will be deployed in s i Up, that is, x ik =1, and if and only if Time ik =1, in order to better reflect the user u j When selecting Edge Servers i The instant distribution reward r ij (t), the reward is evaluated using a cost-benefit marginal utility score, which is defined as the sum of u j Assign to edge servers i The resulting k The cost-effectiveness utility score is defined as To measure the cost-effectiveness of each application deployment solution; Will Represented as user u j Select edge server s in experiment t i The total number of Represents the average reward distribution for the corresponding samples, updated as follows: Use sample average to distribute rewards To represent user u j Select edge servers in the next experiment i is an estimate of the expected distribution reward. In addition, each user u j The instant distribution reward r ij (t) used to update its corresponding average distribution reward; For each user u in the tth trial j ,The edge server selected according to the UCB policy is represented as follows: The first of these is the user uj's choice of edge server s in experiment t i The average distribution of rewards, using historical information to determine the expected value based on past rewards, the second is the exploration item, indicating that user u j Choose s in trial t i The confidence radius of .

8. The edge service deployment and user allocation method based on multi-armed bandit learning according to claim 7 is characterized in that: The specific process of the algorithm in step S4 is as follows: Step S4-1: Starting from initialization, the algorithm sets the application service instance deployment decision variable x and the user allocation decision variable y to an empty set. At trial t = 0, for each user And each server By normalizing the user and each s i Initialize the geographical distance between And update the corresponding is 1; Step S4-2: Perform T iterations. At the beginning of each iteration, update is an empty set to store the tth trial request service a k and select edge servers i and update is an empty set to store service a in the tth trial k Deployed on edge servers i The overall service effectiveness, Next, the edge server with the highest upper confidence bound is selected for all users, and the corresponding total number of selections is Add one and request service a in experiment t k and select edge servers i Add users to the collection middle; Step S4-3: Next, initialize the empty set To store from The set of users selected from satisfying the service delay constraint is calculated The service delay L perceived by each user to all servers in the set ij , and checks whether it meets its maximum allowed service delay If satisfied, add the user to In the set; if it is not satisfied, the corresponding immediate distribution reward r ij (t) is set to 0, and then, according to the obtained set If it is not an empty set, calculate the service delay L for each user ij (t), service satisfaction π ij (t), and its corresponding service utility If it is an empty set, then its corresponding service utility Set to 0, then press All Apply On each edge server Utility score on Sort in descending order and keep selecting the service-edge server pair with the largest score until the resource and budget constraints are reached. For all unselected deployment solutions x ik (t) = 1, update its corresponding and x ik (t) = 0, and at the same time, for all relevant users, the allocation decision y ij (t), service satisfaction π ij (t) and instant distribution reward r ij (t) are all set to 0, and then update each user To each server Expected distribution reward Step S4-4: After T iterations, the user u is obtained j The final allocation decision in For u j With the largest instant allocation reward, edge service deployment decisions are made based on user allocation decisions Step S4-5: Finally, x and y are returned as the final edge service deployment and user allocation solution.