Sharing-based service model pre-deployment method in edge computing network
By decomposing the service model pre-deployment problem in edge computing networks into three sub-problems and using specific algorithms to optimize the model deployment, the efficiency and latency problems of service model deployment in resource-constrained networks are solved, and efficient resource utilization and latency optimization are achieved.
Patent Information
- Application Number
- CN202510129167.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
In resource-constrained edge computing networks, how to efficiently deploy and manage service models, especially considering the correlation between different service models, to reduce latency and optimize resource utilization.
The service model pre-deployment problem in the edge computing network is broken down into three sub-problems. The shared service model pre-deployment strategy is built through descending chain first adaptation algorithm, a collection division algorithm based on neighborhood search, and a hybrid ant colony algorithm to optimize the model deployment, including determining the number of deployments of the sub-service model, grouping and mapping to the edge server.
The model deployment resource loss and data transmission delay are optimized, ensuring the feasibility and efficiency of the deployment strategy, and reducing the model deployment cost and data transmission delay.
Smart Images

Figure CN120075888A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of edge computing, and particularly relates to a method for pre-deploying a service model based on sharing in an edge computing network. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, emerging applications such as intelligent diagnosis assistants, image recognition, and personalized recommendations have been widely used on terminal devices. These applications usually require the deployment of complex service models, such as neural network models, deep neural network models, etc., on terminal devices, which exceed the limited computing and storage capabilities of terminal devices. To solve this problem, Mobile Edge Computing (MEC) has emerged as a new distributed computing paradigm to meet the growing demand for Quality of Service (QoS) of applications. By deploying service models on server devices at the network edge (such as base stations, wireless access points, etc.), terminal devices can offload computing tasks to nearby edge servers through wireless communication, and the corresponding service models already deployed on the edge servers provide computing services, thereby reducing the computing and storage loads of the terminal devices themselves. Compared with cloud computing, edge computing can improve the data privacy and security of terminal devices, while enhancing the availability and flexibility of services.
[0003] Although the resources of edge servers are more abundant than those of terminal devices, their computing, communication, and storage capabilities are still limited. Therefore, how to efficiently deploy and manage service models in an edge computing network, and how to reasonably utilize the computing, communication, and storage resources of the network to improve service quality, are one of the key problems faced by edge computing networks. In recent years, researchers have carried out relevant research on the service model deployment problem in edge computing networks and achieved many beneficial results. However, in traditional service model deployment research, many scholars only focus on the deployment of service models themselves and do not pay attention to the correlation between different service models. In practical applications, different service models are often structurally related, especially service models with complex functions, which are usually composed of multiple sub-service models with different functions. Since the server resources in edge computing networks are usually limited, if the correlation of service models is ignored and sub-service models are not shared, it may partially result in the inability to deploy all service models in the network or overloading of edge servers. Therefore, in resource-constrained edge computing networks, how to design a sub-model sharing mechanism based on the correlation of service models and effectively deploy the shared sub-models to reduce latency has become an urgent and challenging issue. Summary of the Invention
[0004] Objective of the Invention: Aiming at the resource-constrained edge computing network, a sharing-based pre-deployment method for service models is provided to jointly optimize the deployment cost of service models and the data transmission latency between sub-service models. First, aiming at the pre-deployment problem of service models in the edge computing network, a sharing-based pre-deployment architecture for service models is proposed and system modeling is carried out, with the optimization objective of minimizing the deployment cost of service models and the data transmission latency between sub-service models. Second, the sharing-based pre-deployment problem of service models in the edge computing network is decomposed into three sub-problems, and three corresponding methods are proposed to construct a complete sharing-based pre-deployment strategy for service models.
[0005] Technical Solution: A sharing-based pre-deployment method for service models in an edge computing network according to the present invention includes the following steps:
[0006] S1. Carry out system modeling for the sharing-based pre-deployment problem of service models in the edge computing network, with the optimization objective of minimizing the weighted sum of the deployment cost of service models and the data transmission latency between sub-service models;
[0007] S2. Decompose the sharing-based pre-deployment problem of service models in the edge computing network into three sub-problems and solve them separately, including:
[0008] S201. The first sub-problem determines the deployment quantity of each sub-service model and the sub-service model sharing scheme, models it as a bin-packing problem with the optimization objective of minimizing the model deployment cost, proposes a descending-chain first-fit algorithm DCFF, and finally obtains the quantity of sub-service model instances to be deployed and the sub-service model sharing scheme, constructs an undirected graph, with all sub-service model instances to be deployed as vertices, and the data volume required to be transmitted between any two consecutive sub-service models belonging to the same service model as the edge weight value corresponding to the two vertices;
[0009] S202. The second sub-problem determines the grouping of sub-service models, and each group of sub-service models is deployed on the same edge server, models it as a maximum k-uncut problem with capacity constraints with the objective of maximizing the sum of intra-group connection weights (Maximum k-Uncut Problem), proposes a neighborhood search-based set partitioning algorithm NLS, and finally obtains k groups of sub-service model instances, and each group of sub-service model instances is deployed on the same server; constructs an undirected graph with k vertices, each vertex corresponding to a virtual edge server, and the connection weight between any two vertices is the data volume required to be transmitted between the sub-service models deployed on the two virtual edge servers;
[0010] S203. The third sub - problem determines the service model deployment plan for each edge server in the network, that is, mapping k groups of sub - service model instances to k edge servers, which is modeled as a Quadratic Assignment Problem (QAP) with the goal of minimizing data transmission latency, and a meta - heuristic optimization method HAS based on a hybrid ant colony algorithm is proposed.
[0011] Further, step S1 includes the following sub - steps:
[0012] S101. Build a system model; model the edge computing network as an undirected graph Each edge server The amount of resources provided for the pre - deployed service model is C v For any two edge servers The communication bandwidth between them is B u,v The m computing services provided by the network for terminal devices are represented as a service model set Each service model Is represented as a triple Among them, p k Is the utilization rate of this service model, Is an ordered sequence of sub - service models that make up this service model. Assume that there are no duplicate sub - service models in the sequence, and the output of the previous sub - service model Is used as the input of the subsequent sub - service model Its connection is represented as All connections form a set The output data volume of the sub - service model That is, the input data volume of the sub - service model Is represented as Assume that the set of all sub - service models is That is The amount of resources required to deploy each sub - service model Is represented as C(F i )
[0013] S102. Build a service model pre - deployment system architecture based on sharing; assume that the edge server Deploys M vi Instances of sub - service models That is Represents this instance set, that is Assume Represents the service model S k Using instance Otherwise it is 0; for each sub - service model instance pre - deployed on the edge server v It can be shared by multiple service models without conflict if and only if the sum of the usage rates of these service models is not greater than 1, that is When the sum of the usage rates of different service models sharing the same sub-model instance is not greater than 1, and there are differences in the composition and order of their sub-service model sequences, the possibility of contention for computing resources of the same server at the same time can be reduced;
[0014] S103. Construct a deployment cost model and a data transmission delay model; for each edge server The deployment cost is defined as the sum of the required resources for deploying the set of sub-service model instances and needs to satisfy Therefore, the total deployment cost of the edge computing network is
[0015]
[0016] When a service model can share sub-service model instances without conflict, the computing delay of this service model is not affected by the deployment strategy. Therefore, the delay model no longer includes the computing delay of sub-service models, but only includes the data transmission delay between sub-service models; the delay model is expressed as the total delay caused by data transmission. For Its transmission delay is expressed as
[0017]
[0018] where represents mapped to e(u, v), otherwise 0; the total data transmission delay of the edge computing network is
[0019]
[0020] S104. Formalize the problem definition; aiming to propose a sharing-based pre-deployment and scheduling strategy for service models, minimize the weighted sum of the deployment cost of service models and the data transmission delay between sub-service models.
[0021] Furthermore, in step S104, the specific method of minimizing the weighted sum of the deployment cost of service models and the data transmission delay between sub-service models is as follows:
[0022] Input:
[0023] Edge computing network topology;
[0024] Set of service models;
[0025] Set of sub-service models;
[0026] The resource volume of edge servers and the bandwidth capacity between edge servers;
[0027] The resource volume required by the sub-service model;
[0028] Output:
[0029] The deployment decisions for all sub-service models;
[0030] The usage decisions for sub-model instances of all service models; The routing decisions for all service models;
[0031] Optimization objectives:
[0032]
[0033] s.t.
[0034] C1:
[0035] C2:
[0036] C3:
[0037]
[0038] C4:
[0039] Among them, C1 ensures that the resource requirements of sub-models deployed on edge servers do not exceed the upper limit that edge servers can provide. C2 ensures that any sub-service model of any service model S k can only be served by one sub-model instance in the network. C3 ensures the law of conservation of traffic. C4 restricts that the sum of the usage rates of service models sharing the same sub-model instance is not greater than 1.
[0040] Furthermore, step S201 includes the following sub-steps:
[0041] S201-1. Create an instance graph initially empty, and for any create an instance set Ins i , initially empty, and Ins i is always sorted by the remaining capacity of the instance;
[0042] S201-2. For each model Under the conflict-free sharing architecture, the capacity of each sub-service model instance is 1, and each of its sub-service models has a demand capacity of p for the instancek ; Create sub-service model instances and for each sub-service model Determine associated instances such that the sum of the sub-model capacity requirements of the instance services is not greater than 1, and the connection weight between instances is the sum of the data transfer amounts between two instances, finally obtaining a complete instance graph
[0043] Furthermore, step S202 includes the following sub-steps:
[0044] S202-1. Determine the minimum number of instances n that can be deployed on each edge server m and the number of groups k. Each group represents a set of instances. Finally, all instances within the same instance set will be deployed on the same edge server;
[0045] S202-2. Randomly and initially partition the vertex set into k sets, corresponding to the sets of sub-service model instances deployed on k virtual edge servers, denoted as where,
[0046] S202-3. Iteratively optimize the set partition. For any If swapping u I into set v I swapped into can reduce the data transfer amount between instance sets, then perform the swap until no vertex pair can be optimized;
[0047] S202-4. Update the connections between to obtain a k-partition graph
[0048] Furthermore, step S203 includes the following sub-steps:
[0049] S203-1. Randomly generate m initialization sequences π 1 (1), …, π m (1). Each sequence represents a sequence of edge servers and is a feasible solution for deployment; Perform local search on π 1 (1), …, π m (1), retain the optimal solution π * found by the search, initialize the pheromone trail matrix T of, and then perform I max rounds of cyclic iteration. The initial solution selection for each round of iteration depends on a reinforcement flag, which is initialized to the on state, and then successively execute S203-2, S203-3, S203-4;
[0050] S203-2. For any sequence π k (i), where i represents the number of iterations, first modify its sequence through R swap operations to obtain After the swap operation is completed, for perform local search to obtain
[0051] S203-3. Determine the initial value π k (i + 1) for the next iteration between π according to the reinforcement status flag; if for any π k (i + 1) it satisfies π k (i + 1) = π k (i), then update the reinforcement flag to the off state. If the optimal solution π k changes, turn on the reinforcement flag, and then update the pheromone trail matrix T; *
[0052] S203-4. If no improvement to the current optimal solution is detected in the most recent S (S ≤ I max ) iterations, then activate the diversification mechanism. The diversification process includes clearing all information in the pheromone trail by re-initializing the pheromone trail matrix, taking the optimal solution found so far as the current solution of one ant, and randomly generating new current solutions for m - 1 ants.
[0053] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method of the present invention.
[0054] The present invention also discloses a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.
[0055] The present invention also discloses a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the method of the present invention are implemented.
[0056] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The method of the present invention decomposes the pre-deployment problem of the sharing-based service model in the edge computing network into three sub-problems, reducing the problem complexity while ensuring the feasibility of the deployment strategy. First, the first sub-problem takes minimizing the model deployment cost as the optimization objective, and finally obtains the number of required deployed sub-service model instances and the sub-service model sharing scheme. At this stage, a descending chain first-fit algorithm is proposed, which has a provable approximation ratio of 2. The second sub-problem determines the grouping of sub-service models, and each group of sub-service models is deployed on the same edge server. It is modeled as a maximum k-non-cut problem with capacity constraints aiming to maximize the sum of connection weights within the group, and aims to maximize the amount of transmitted data within the group. At this stage, a set partitioning algorithm based on neighborhood search is proposed, which has a provable approximation ratio of to be filled in. The third sub-problem determines the service model deployment scheme for each edge server in the network, that is, mapping k groups of sub-service model instances to k edge servers. It is modeled as a quadratic assignment problem aiming to minimize the data transmission delay, and a meta-heuristic optimization method based on a hybrid ant colony algorithm is proposed. In summary, through the innovative three-stage method, the present invention optimizes the resource loss of model deployment and data transmission delay, and provides strict performance guarantees for the first two stages, ensuring the feasibility of the method of the present invention. Brief Description of the Drawings
[0057] Figure 1 It is a model deployment model diagram of a resource-constrained edge computing network provided by the present invention.
[0058] Figure 2 It is an example diagram of the first stage of deployment provided by the present invention.
[0059] Figure 3 It is an example diagram of the second stage of deployment provided by the present invention.
[0060] Figure 4 It is an example diagram of the third stage of deployment provided by the present invention.
[0061] Figure 5 It is a diagram of the deployment resource consumption generated by each method to complete all requests under the change of the request quantity and the total delay generated by each method to complete all requests under the change of the request quantity. Detailed Embodiment
[0062] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0063] Aspects of the present invention are described with reference to the accompanying drawings, in which a number of illustrative embodiments are shown. Embodiments of the present invention are not limited to those described in the drawings. It should be understood that the present invention can be implemented by any one of the various concepts and embodiments introduced above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the present invention are not limited to any embodiment. Additionally, some aspects disclosed in the present invention can be used alone or in any suitable combination with other aspects disclosed in the present invention.
[0064] The present invention provides a sharing-based model pre-deployment method in an edge network, including the following steps:
[0065] S1. Systematically model the problem of sharing-based model pre-deployment in the edge network, with the optimization objective being to minimize the weighted sum of the model deployment cost and the data transmission delay. Specifically:
[0066] S101. Construct a system model. Model the edge computing network as an undirected graph Each edge server The amount of resources provided for the pre-deployed service model is C v , and the communication bandwidth between any two edge servers is B u,v . The m computing services provided by the network for the terminal device are represented as a service model set Each service model is represented as a triple where p k is the usage rate of this service model, is an ordered sequence of sub-service models that make up this service model. Without loss of generality, it is assumed that there are no duplicate sub-service models in the sequence, and the output of the previous sub-service model serves as the input of the subsequent sub-service model , and its connection is represented as All connections form a set The output data volume of the sub-service model (i.e., the input data volume of the sub-service model ) is represented as (Specifically, is the input data volume of the service model S k ). Let the set of all sub-service models be That is, The amount of resources required to deploy each sub-service model is represented as C(F i ). As Figure 1As shown in the figure, the edge computing network in this embodiment includes 5 edge servers and 2 service models. Each edge server provides 2 CPUs of resource for deployment, and the links between edge servers have different bandwidths. There are also differences in the composition structure, utilization rate, and transmitted data volume of different service models.
[0067] S102. Build a pre-deployment system architecture for service models based on sharing. Assume that the edge server Deploy M vi instances of sub-service models of, represent this instance set, that is Assume represent the service model S k usage instance On the contrary, it is 0. For each instance of the sub-service model pre-deployed on the edge server v it can be shared by multiple service models without conflict, if and only if the sum of the utilization rates of these service models is not greater than 1, that is When the sum of the utilization rates of different service models sharing the same sub-model instance is not greater than 1, and there are differences in the composition and order of their sub-service model sequences, it can effectively reduce the possibility of contention for computing resources of the same server at the same time. Therefore, this architecture is reasonable and can optimize the overall performance and stability of the system while ensuring resource utilization.
[0068] S103. Build a deployment cost model and a data transmission delay model. The deployment cost of each edge server is defined as the sum of the required resources for deploying the sub-service model instance set and needs to meet Therefore, the total deployment cost of the edge computing network is
[0069]
[0070] When the service model can share the sub-service model instance without conflict, the computing delay of this service model is not affected by the deployment strategy. Therefore, the delay model no longer includes the computing delay of the sub-service model, but only includes the data transmission delay between sub-service models. The delay model is expressed as the total delay generated by data transmission. For its transmission delay is expressed as
[0071]
[0072] where represents mapped to e(u, v), otherwise it is 0. The total data transmission delay of the edge computing network is
[0073]
[0074] S104. Formal definition of the problem. The objective of the present invention is to propose a pre - deployment and scheduling strategy for service models based on sharing, aiming to minimize the weighted sum of the deployment cost of service models and the data transmission delay between sub - service models, that is:
[0075] Input:
[0076] 1. Edge computing network topology
[0077] 2. Set of service models
[0078] 3. Set of sub - service models
[0079] 4. Resource amount of edge servers and bandwidth capacity between edge servers
[0080] 5. Required resource amount of sub - service models
[0081] Output:
[0082] 1. Deployment decisions for all sub - service models
[0083] 2. Decision on the use of sub - model instances of all service models
[0084] 3. Routing decisions for all service models
[0085] Optimization objective:
[0086]
[0087] s.t.
[0088] C1:
[0089] C2:
[0090] C3:
[0091]
[0092] C4:
[0093] C1 ensures that the resource requirements of sub - models deployed on edge servers do not exceed the upper limit that can be provided by edge servers, and C2 ensures that for any service model S kAny sub-service model can only be served by one sub-model instance in the network. C3 ensures the law of conservation of traffic, and C4 limits the sum of the usage rates of service models sharing the same sub-model instance to no more than 1.
[0094] S2. Decompose the problem of pre-deploying shared models in the edge network into three sub-problems for separate solutions, including:
[0095] S201. The first sub-problem determines the deployment quantity of each sub-service model and the sub-service model sharing scheme, models it as a bin-packing problem with the optimization goal of minimizing the model deployment cost, proposes a descending-chain first-fit algorithm (DCFF), and finally obtains the quantity of sub-service model instances to be deployed and the sub-service model sharing scheme. Construct an undirected graph, with all sub-service model instances to be deployed as vertices, and the amount of data to be transmitted between any two consecutive sub-service models belonging to the same service model as the edge weight of the corresponding two vertices. As Figure 2 shown, in this embodiment, finally, one instance is created for each sub-model, and the instances of sub-model 3 and sub-model 4 are shared by these two models. Step S201 includes the following sub-steps:
[0096] S201-1. Create an instance graph initially empty, and for any create an instance set Ins i , initially empty, and Ins i is always sorted in descending order according to the remaining capacity of the instance.
[0097] S201-2. Without loss of generality, assume that the service models are sorted in descending order of usage rate, that is, for the set there is p 1 ≥…≥p m .
[0098] S201-3. For each model In a conflict-free sharing architecture, the capacity of each sub-service model instance is 1, and the demand capacity of each of its sub-service models for the instance pair is p k . For the sub-service model If there is an instance in the set Ins i with a remaining capacity greater than p k , take the instance v I with the largest remaining capacity and associate it with the sub-service model , and update the remaining capacity of the instance v I ; if not, create a new instance v Inew with a capacity of 1, associate it with the sub-service model, and update the remaining capacity of the instance v Inew , Insi = Ins i ∪ {v Inew}}. Update the sub-service model Associate instance v cur with its previous sub-service model Associate instance v pre The connection e(v cur , v pre ) between them. If e(v cur , v pre ) ∈ ε I , increase its weight Otherwise, ε I = ε I ∪ {e(v cur , v pre )}, and the weight of the edge e(v cur , v pre ) is
[0099] S202. The second sub-problem determines the grouping of sub-service models. Each group of sub-service model instances is deployed on the same edge server, and it is modeled as a maximum k-uncut problem with capacity constraints aiming to maximize the sum of intra-group connection weights (Maximum k-Uncut Problem). A neighborhood search-based set partitioning algorithm (NLS) is proposed, and finally k groups of sub-service model instances are obtained, with each group of sub-service model instances deployed on the same server. Construct an undirected graph with k vertices, where each vertex corresponds to a virtual edge server, and the connection weight between any two vertices is the amount of data to be transmitted between the sub-service models deployed on the two virtual edge servers. As Figure 3 shown, in this embodiment, two groups are finally obtained. The instances of sub-model 1 and the instances of sub-model 2 are in the same group, and the instances of sub-model 3 and the instances of sub-model 4 are in the same group. The specific steps are as follows:
[0100] S202-1. Under the limitation of the resource amount C v , each edge server can deploy at least sub-service model instances. The number of vertices in the instance graph output by step S201 (i.e., the number of sub-service model instances to be deployed in the edge computing network) is n I . Therefore, when there are edge servers in the edge computing network, the instances in the instance graph will definitely be able to be deployed in the edge computing network . Without loss of generality, let
[0101] S202-2. Take the vertex set The random initial division is into k sets, corresponding to the set of sub-service model instances deployed by k virtual edge servers, denoted as in, set up The initialization points must satisfy n min ≤n max ≤n min +1.
[0102] S202-3. Optimize set partitioning For any sub-service model instance v I and u I , which is The weight on is w(v I ,u I ), therefore, for any sub-service model instance v I Its collection of subservice model instances The weight of Iteratively optimize the set partitioning, for any If satisfied Then u I Exchange to collection v I Exchange to This continues until there are no more vertex pairs that satisfy this requirement.
[0103] S202-4, Update The connection between them is the k-partition graph
[0104] S203, the third sub-problem is to minimize the data transmission delay and map each virtual edge server to a real edge server. This stage problem is a generalization of the classic QAP problem. This stage proposes a meta-heuristic optimization method based on the hybrid ant colony algorithm (HAS). Figure 4 As shown, this embodiment finally maps two virtual edge servers to two real edge servers.
[0105] S203-1. Based on the edge computing network generate The shortest distance matrix D between edge servers is divided into two parts according to the k-partition graph. generate The traffic matrix A between virtual edge servers, the first k×k area in the matrix A contains valid data, A i,j (1≤i≠j≤k) is the instance set and The amount of data transferred between The rest are all zero. Randomly generate m initialization sequences π 1(1),…, π m (1), where m is the number of ants in the hybrid ant colony algorithm, and each sequence represents a sequence of edge servers, where denotes deploying the instance set on edge server h, satisfying is a feasible solution for the deployment. For π 1 (1),…, π m (1), local search is performed. Specifically, two complete neighborhood detections are carried out. If an improved solution is obtained in the first neighborhood detection, the local search is terminated; otherwise, the second neighborhood detection is carried out. For any sequence π k (1), the neighborhood detection first creates the set In each round of loop, an instance set i that is not included in I is selected, such that I = I ∪ {i}. Then, the selection order of all instance sets other than the instance set i is randomly determined, and the instance sets j are selected in order. If swapping and improves the solution, the swap is executed. The loop iterates until |I| = |V|. Initialize the pheromone trail matrix T, and each element τ ij represents the weight of deploying the instance set on edge server j, defined as τ ij = 1 / (Q·f(π * ), where Q is a user-defined parameter, represents the total delay generated by data transmission. Next, I max rounds of loop iterations are carried out. The initial solution selection in each round depends on a reinforcement flag (initialized to the on state), and then S203-2, S203-3, and S203-4 are executed in sequence.
[0106] S203-2. For any sequence π k (i), where i represents the round of iteration, first modify its sequence through R swap operations to obtain The swap operation follows the following rules: First, randomly select an instance set r, and then select an instance set s ≠ r to swap and There are two strategies for the selection of the instance set s: (1) Select the instance set s that makes the largest; (2) Select the instance set s with probability . Each time, the first strategy is selected with probability q, and the second strategy is selected with probability 1 - q. After the swap operation, perform local search on to obtain
[0107] S203-3. Determine the initial value π of the next iteration according to the reinforcement status flag k (i + 1) for any sequence π k (i), if the reinforcement flag is in the enabled state, π k (i)+1) is the optimal solution between π k (i) and otherwise, if for any π k (i + 1) all satisfy π k (i + 1) = π k (i), then update the reinforcement flag to the disabled state. If there exists such that update π * to and enable the reinforcement flag. Then, update the pheromone trail matrix T. First, weaken all pheromone trails by setting τ ij =(1 - α 1 )τ ij , and secondly, reinforce the pheromone trails only according to the optimal solution π *
[0108] S203-4. If no improvement to the current optimal solution is detected in the most recent S (S ≤ I max ) iterations, activate the diversification mechanism. The diversification process includes clearing all information in the pheromone trail matrix by re-initializing it, taking the optimal solution found so far as the current solution of 1 ant, and randomly generating new current solutions for m - 1 ants.
[0109] where the parameter settings are as follows: R = |V| / 3, α 1 = α 2 = 0.1, Q = 100, S = |V| / 2, q = 0.9, m = 10.
[0110] Perform performance verification on the pre-deployment method of the sharing-based service model in the edge computing network provided by the present invention. The experimental results are as Figure 5 As shown, the method proposed by the present invention is DCFF + NLS + HAS. In the experiment, the present invention is compared with three other methods, namely CFF + GREEDY, DSP + NN, and RQAP. The difference between CFF + GREEDY and the method of the present invention is that when determining the deployment quantity of each sub-service model and the sub-service model sharing scheme, the service model and the sub-service model instances are not sorted, and then the greedy algorithm is used to map each sub-service model instance to the edge server that generates the minimum transmission delay. DSP + NN is a method based on dynamic programming and the nearest neighbor algorithm, which does not consider the sharing of sub-service models. RQAP is a method based on Markov chain and dynamic programming, which selects the deployment scheme with the minimum cost for each service model and considers reusing the existing sub-service model instances during the deployment process. Figure 5 The left figure in [Figure] shows the deployment resource consumption generated by each method to complete all requests under the change of the number of requests. It can be seen that the method proposed by the present invention consumes the least resources. Figure 5 The right figure in [Figure] shows the total delay generated by each method to complete all requests under the change of the number of requests. It can be seen that the performance of the method of the present invention is similar to that of RQAP and both are superior to the other two methods, but the method proposed by the present invention consumes fewer resources than RQAP.
Claims
1. A method for pre-deploying a shared service model in an edge computing network, characterized in that: The steps include: S1. System modeling for the pre-deployment problem of shared service models in edge computing networks. The optimization goal is to minimize the weighted sum of the service model deployment cost and the data transmission delay between sub-service models. S2. Decompose the pre-deployment problem of the shared service model in the edge computing network into three sub-problems and solve them separately, including: S201. The first sub-problem is to determine the number of deployments of each sub-service model and the sub-service model sharing scheme, which is modeled as a packing problem with the optimization goal of minimizing the model deployment cost. A descending chain first fit algorithm DCFF is proposed, and finally the number of sub-service model instances required to be deployed and the sub-service model sharing scheme are obtained. An undirected graph is constructed, with all sub-service model instances to be deployed as vertices, and the amount of data required to be transmitted between any two consecutive sub-service models belonging to the same service model is the edge weight of the corresponding two vertices. S202. The second sub-problem is to determine the grouping of sub-service models. Each group of sub-service models is deployed on the same edge server. It is modeled as a maximum k-uncut problem with capacity constraints to maximize the sum of the connection weights within the group. A set partitioning algorithm NLS based on neighborhood search is proposed, and finally k groups of sub-service model instances are obtained. Each group of sub-service model instances is deployed on the same server. An undirected graph with k vertices is constructed, where each vertex corresponds to a virtual edge server, and the connection weight between any two vertices is the amount of data required to be transmitted between the sub-service models deployed on the two virtual edge servers. S203. The third sub-problem determines the service model deployment scheme for each edge server in the network, that is, k groups of sub-service model instances are mapped to k edge servers, which is modeled as a quadratic allocation problem with the goal of minimizing data transmission delay, and a meta-heuristic optimization method HAS based on a hybrid ant colony algorithm is proposed.
2. A method for pre-deploying a service model based on sharing in an edge computing network according to claim 1, characterized in that: Step S1 includes the following sub-steps: S101. Build a system model; model the edge computing network as an undirected graph Each edge server The amount of resources provided for the pre-deployed service model is C v , any two edge servers The communication bandwidth between u,v , the m types of computing services provided by the network to the terminal device are represented as a service model set Each service model Represented as a triple Among them, p k is the utilization rate of the service model, is an ordered sequence of sub-service models that make up the service model. Assume that there is no duplicate sub-service model in the sequence. The preceding sub-service model The output of The input of All connections form a collection Subservice Model The amount of output data, that is, the sub-service model The amount of input data is expressed as Suppose the set of all sub-service models is Right now Deploy each subservice model The amount of resources required is expressed as C(F i ); S102. Build a pre-deployment system architecture based on a shared service model; set up edge servers Deploy M vi Sub-service model For example, Represents the instance set, that is, set up Represents the service model S k Use Case Otherwise, it is 0; for each sub-service model instance pre-deployed on the edge server v It can be shared by multiple service models without conflict if and only if the sum of the usage rates of these service models is not greater than l, that is, When the sum of the usage rates of different service models sharing the same sub-model instance is no greater than 1, and their sub-service model sequences differ in composition and order, the possibility of them competing for the computing resources of the same server at the same time can be reduced; S103, constructing a deployment cost model and a data transmission delay model; each edge server The deployment cost is defined as the set of deployed sub-service model instances The amount of resources required and Therefore, the total deployment cost of the edge computing network is When a service model can share sub-service model instances without conflict, the computational latency of this service model is not affected by the deployment strategy. Therefore, the latency model no longer includes the computational latency of the sub-service model, but only includes the data transmission latency between sub-service models. The latency model is represented as the sum of the latency caused by data transmission. Its transmission delay is expressed as in express Mapped to e(u, v), otherwise 0; the total data transmission delay of the edge computing network is S104. Formal definition of the problem: The goal is to propose a shared-based service model pre-deployment and scheduling strategy to minimize the weighted sum of the deployment cost of the service model and the data transmission delay between sub-service models.
3. The method for pre-deploying a service model based on sharing in an edge computing network according to claim 2, characterized in that: In step S104, the weighted sum of the deployment cost of the minimization service model and the data transmission delay between the sub-service models is specifically: enter: Edge computing network topology; Service model collection; A collection of subservice models; The amount of edge server resources and the bandwidth capacity between edge servers; The amount of resources required by the sub-service model; Output: Deployment decisions for all sub-service models; Sub-model instance usage decisions for all service models; Routing decisions for all service models; Optimization goal: Among them, C1 ensures that the resource requirements of the sub-model deployed on the edge server do not exceed the upper limit that the edge server can provide, and C2 ensures that any service model S k Any sub-service model can only be served by one sub-model instance in the network, C3 ensures the traffic conservation law, and C4 limits the sum of the usage rates of service models sharing the same sub-model instance to no more than 1.
4. The method for pre-deploying a service model based on sharing in an edge computing network according to claim 1, characterized in that: Step S201 includes the following sub-steps: S201-1. Create an instance diagram Initially empty and any Create an instance set Insi, which is initially empty. i Always sort by instance remaining capacity; S201-2. For each model In the conflict-free sharing architecture, the capacity of each sub-service model instance is 1. The required capacity for the instance pair is pk; Create subservice model instances and for each subservice model Determine the associated instances so that the sum of the sub-model capacity requirements of the instance services is no greater than 1. The connection weight between instances is the sum of the amount of data transmitted between the two instances. Finally, a complete instance graph is obtained.
5. The method for pre-deploying a service model based on sharing in an edge computing network according to claim 1, characterized in that: Step S202 includes the following sub-steps: S202-1. Determine the minimum number of instances n that can be deployed on each edge server m And the number of groups k, each group represents an instance set, and eventually all instances in the same instance set will be deployed on the same edge server; S202-2, Vertex Set The random initial division is into k sets, corresponding to the set of sub-service model instances deployed by k virtual edge servers, denoted as in, S202-3, iterative optimization set partitioning, for any If u I Exchange to collection v I Exchange to If the amount of data transferred between sets of instances can be reduced, the swap is performed until no more vertex pairs can be optimized; S202-4, Update The connection between them gets the k-partition graph 6. The method for pre-deploying a service model based on sharing in an edge computing network according to claim 1, characterized in that: Step S203 includes the following sub-steps: S203-1. Randomly generate m initialization sequences π 1 (1),…,γ m (1), each sequence They all represent a sequence of edge servers, which is a feasible solution for deployment; for π 1 (1),…,π m (1) Perform local search and retain the optimal solution π * ,initialization The pheromone trajectory matrix T is then I max In rounds of iterations, the initial solution selection of each round of iteration depends on a strengthening flag, which is initialized to an on state, and then S203-2, S203-3, and S203-4 are executed in sequence; S203-2. For any sequence π k (i), i represents the number of iterations. First, the sequence is modified by R swap operations to obtain After the swap operation is completed, Perform a local search to obtain S203-3, according to the strengthening status mark in π k (i) with Determine the initial value π of the next iteration k (i+1); if for any π k (i+1) all satisfy π k (i+1)=π k (i), then the update enhancement flag is turned off. If the optimal solution π * Changes occur, the reinforcement flag is turned on, and then the pheromone trajectory matrix T is updated; S203-4, if in the nearest S (S ≤ I max ) iterations, if no improvement on the current optimal solution is detected, the diversification mechanism is activated. The diversification process includes clearing all information in the pheromone trails by reinitializing the pheromone trail matrix, taking the optimal solution found so far as the current solution for one ant, and randomly generating new current solutions for m-1 ants.
7. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.