Calculation network storage resource collaborative allocation method based on optimization and machine learning dual-drive

By adopting a dual-driven resource allocation method based on optimization and machine learning in computing power network, using DRL model and optimization problems to decompose complex optimization problems, the problem of difficult to guarantee real-time resource allocation in dynamic environments is solved, and efficient and real-time resource allocation and system performance improvement is achieved.

CN119945995AActive Publication Date: 2025-05-06UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510103432.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the real-time nature of resource allocation in a dynamic environment, especially in computing power networks. With the increase in network size and number of users, the overhead of dynamic planning and iterative optimization of scores has increased significantly, resulting in a decrease in resource allocation efficiency.

Method used

The collaborative allocation method of computing network storage resources based on optimization and machine learning is adopted. Through the DRL model, complex multi-stage random optimization problems are decomposed to realize joint optimization of service function deployment, static object cache and traffic planning.

Benefits of technology

This method significantly improves the real-time and efficiency of resource allocation, reduces algorithm complexity, and enables the solution to be scaled to larger networks and higher user requests, improving user request satisfaction rate and system resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945995A_ABST
    Figure CN119945995A_ABST
Patent Text Reader

Abstract

The invention discloses an optimization and machine learning dual-drive-based computing network storage resource collaborative allocation method, which comprises the following steps of: obtaining all user requests to be subjected to resource allocation, and storing the user requests to a request set; inputting the current network resource state, the algorithm decision state and the user state of the service function m of the service request phi which is not traversed in the request set into a trained DRL model to obtain the service function m of the service request phi and the benefit weights wf and ws of the service function m of the service request phi and the required static objects which are respectively placed in each node; respectively splicing wf and ws to weight matrixes Wf and Ws of an algorithm decision state, judging whether m is equal to the total number of service functions of the service request phi, if so, entering the next step, otherwise, updating m to be equal to m + 1, and returning to the model identification step; and judging whether the service request phi is equal to H, if so, obtaining a resource allocation strategy by solving an ILP optimization problem and an LP optimization problem, otherwise, updating phi = phi + 1, and returning to the model identification step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to network resource allocation technology, and in particular to a method for collaboratively allocating computing, network and storage resources based on dual drive of optimization and machine learning. Background Art

[0002] The computing network coordinates the computing and storage resources of distributed cloud, fog, edge and other heterogeneous nodes and the network resources of wide area networks to form a new generation of information infrastructure, which greatly promotes the widespread application of data-intensive augmented information (D-AgI) technology in various emerging business scenarios. D-AgI technology integrates, processes and enhances various real-time data streams to provide users with more personalized and customized services and interaction methods. It not only covers immersive experiences (such as extended reality and holographic communication), but also goes deep into fields such as artificial intelligence generated content (AIGC) and large language models (LLM), broadening the possibility of user interaction with the digital environment and effectively blurring the boundaries between the real and virtual worlds, thereby leading users into an unprecedented digital experience dimension. Future application scenarios such as metaverse, panoramic learning, remote surgery, and smart factories triggered by D-AgI will profoundly change the way people learn, work and entertain in the future.

[0003] In order to meet the high demand of D-AgI for multi-dimensional resources, some works suggest using distributed node resources in multi-layer computing networks to form a service chain, and improve system efficiency by distributing the service functions and storage resources required by the service on multiple nodes and combining them with the scheduling of user request traffic. Among them, the most advanced solution is to use dynamic programming to make decisions on traffic scheduling and update cache placement through multiple rounds of scoring results.

[0004] As the network size and number of users increase, the overhead of dynamic programming and scoring iterative optimization increases significantly. For example, in some work, the complexity of dynamic programming is O(V 3 +V 2 *N), which makes it difficult to ensure the real-time resource allocation in a dynamic environment. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the computing, network and storage resource collaborative allocation method based on dual drive of optimization and machine learning provided by the present invention solves the problem that the prior art is difficult to ensure the real-time resource allocation in a dynamic environment.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0007] A method for collaborative allocation of computing, network and storage resources based on optimization and machine learning dual drive is provided, which comprises the following steps:

[0008] S1. Obtain all user requests to be allocated resources and store them in a request set; each user request corresponds to a service graph, including multiple service functions and static objects executed sequentially;

[0009] S2. Input the current network resource status, algorithm decision status, and user status of the service function m of the service request φ that has not been traversed in the request set into the trained DRL model, and obtain the service function m of the service request φ and the benefit weight w of the static objects required by each node. f 、w s ;

[0010] S3. Set the benefit weight w f 、w s The weight matrix W is spliced ​​to the algorithm decision state respectively f and W s , and determine whether m is equal to the total number of service functions M of the service request φ φ If yes, go to step S4, otherwise update m=m+1 and return to step S2, 1≤m≤M φ ;

[0011] S4, determine whether the service request φ is equal to H, if so, proceed to step S5, otherwise update φ=φ+1, and return to step S2; 1≤φ≤H, H is the total number of service requests in the request set;

[0012] S5. Solve the ILP optimization problem to obtain the allocation strategies x and y of all service functions and required static objects on nodes for all service requests. The ILP optimization problem is to maximize the weight matrix W. f and W s The weighted sum of the allocation strategies x and y;

[0013] S6. Solve the LP optimization problem to obtain the user real-time traffic a on the feasible path σ on the service request φ under the allocation strategy x, y in time slot t (φ,σ) (t), using allocation strategies x, y and user real-time traffic a (φ,σ) (t) As the final resource allocation strategy, the LP optimization problem is to minimize the short-term queue backlog.

[0014] Furthermore, the user status of the service function m of the service request φ includes the request amount of the service request φ The four-tuple of the starting point src, the end point dst and the service request

[0015] The static object type required by the service function m to serve the request φ; is the aggregation rate of service function m serving request φ, The workload of service function m serving request φ, is the reduction factor of the service function m for the service request φ; is the sum of the user real-time traffic of all feasible paths σ on the service request φ at time slot t, and its expression is:

[0016]

[0017] Where σ is the feasible path on the service request φ at time slot t; The hierarchical graph G constructed for the service request φ φ In the example, the K shortest path algorithm is used to find the shortest path between the source and the destination, and the set of K feasible paths that meet the service request constraints is obtained.

[0018] Furthermore, maximize the weight matrix W f and W s The expression of the weighted sum of the allocation strategies x and y is:

[0019]

[0020] st

[0021]

[0022] Among them, max is the maximum value; F k is the size of the static object k; K is the set of all static objects; x i,k ∈{0,1} is whether to cache static object k in node i; y i,φ,m (t) is whether the service function m of the service request φ is deployed at node i at time slot t; S i (t) is the storage resource available to node i at time slot t; x is the total number of all x i,k The allocation strategy composed of i,φ,m (t) is the allocation strategy composed of; ⊙ is the dot product.

[0023] Furthermore, the expression for minimizing the short-term queue backlog is:

[0024]

[0025] st

[0026]

[0027] Among them, min is the minimum value; and are the virtual queues of node i and link ij at time slot t respectively; and are the accumulated traffic of node i and link ij in the physical topology at time slot t; The workload of the service function serving the request φ; is the link flow from node i at the mth layer to node i at the m+1th layer in the hierarchical graph of user request φ, is the link flow from node i at the mth layer to node j at the mth layer in the hierarchical graph of user request φ; C i (t) is the computation frequency of node i in time slot t; is the traffic on link ij associated with the static object k used in the service request; C ij (t) is the transmission bandwidth of link ij in time slot t; The hierarchical graph G constructed for the service request φ φ In , the K shortest path algorithm is used to find the shortest path between the source and the sink to obtain the set of K feasible paths that meet the service request constraints; σ is the feasible path on the service request φ at time slot t; The weight used to calculate the accumulated traffic of the physical topology; To allow the traffic of service request φ to flow through σ, the computational load borne by node i.

[0028] Further, and The expressions are:

[0029]

[0030] in, and are the cumulative scaling factors of the mth and m-1th service functions in the service request φ, respectively; Indicates that when link i m i m+1 It is 1 when it belongs to the feasible path σ, and 0 when it does not belong to σ; is the reduction factor of the service function m-1 for the service request φ.

[0031] Furthermore, the virtual queue and The expressions are:

[0032]

[0033] in, and are the virtual queues of node i and link ij at time slot t-1 respectively; C i (t-1) is the calculation frequency of node i in time slot t-1; C ij (t-1) is the transmission bandwidth of link ij in time slot t-1; and are the accumulated traffic of node i and link ij in the physical topology at time slot t-1; [·] + Indicates that the result of the equation in this symbol must be greater than or equal to 0, otherwise it is forcibly set to 0.

[0034] Furthermore, the flow The expression is:

[0035]

[0036]

[0037] Among them, SP(o k ,l) is from the super source point o k The shortest path to node l; is the traffic on node l associated with the use of static object k in service requests at time slot t; is the traffic on node i associated with the use of static object k in service requests at time slot t; is the reduction factor of the service function m for the service request φ; For the service request φ in time slot t, after being processed by the service function, it is sent to link i m j m Traffic on The static object type required by the service function m to serve the request φ; d φ is the end point of the user request φ; i m is the node i at the mth layer of the hierarchical graph.

[0038] Furthermore, the DRL model is a model designed based on the Actor-Critic architecture; the Critic network evaluates the benefit weight w generated by the Actor network by estimating the value function f 、w s , the value function is determined by the parameters of the Critic network. The parameters of the Critic network are updated according to the reward of reinforcement learning. The expression of reward is:

[0039]

[0040] Among them, Reward is the model reward; Value is the target value obtained by solving the LP optimization problem, that is, C is a constant.

[0041] Furthermore, the output of the DRL model is 2×|V| values, and the first|V| values ​​and the last|V| values ​​of the 2×|V| values ​​are used as the benefit weights w f and w s , |V| is the number of nodes in the entire network combined with the number of nodes in V.

[0042] Furthermore, user requests generate content for artificial intelligence, large language models that require the transmission of user-provided multimodal data, or visual data in virtual reality, augmented reality, holographic communications, and the metaverse.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. This solution designs DRL as a multi-stage output, that is, an inference output is performed at each stage of each user request, and each output result is defined as a weight, which represents the contribution value of each integer decision variable to the optimization target; then, a simple ILP with constrained maximization of the sum of weights is solved to output the final deployment and caching actions. This design will greatly reduce the state and action space of DRL, improve the algorithm operation efficiency, and ensure that the final output satisfies the constraints and globally optimized actions.

[0045] 2. This solution uses the dual-granularity solution framework of Lyapunov optimization and deep reinforcement learning (DRL) to decompose complex multi-stage random optimization problems into easy-to-solve single-stage sub-problems, namely ILP optimization problems and LP optimization problems, reducing the difficulty of solving the scale and providing the feasibility of real-time solution, thereby ensuring real-time and efficient resource allocation, while also ensuring the long-term stability of the data queue (achieved by LP optimization problems). The combination of DRL and ILP can also greatly reduce the complexity of the algorithm, making the solution scalable to larger-scale networks and higher user request volumes, and improving the versatility and adaptability of the algorithm.

[0046] 3. This solution comprehensively considers service function deployment, static object caching and traffic planning, realizes the joint optimization of transmission, computing and storage resources, and significantly improves the user request satisfaction rate and the utilization efficiency of system resources. Based on deep reinforcement learning, this solution can quickly adapt to changes and maintain efficient operation under highly dynamic user requests and network resource conditions, ensuring higher user traffic carrying capacity and better queue backlog control. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Request service graph for users.

[0048] Figure 2 It is an auxiliary graph constructed based on the service graph in the user request; (a) is the original graph structure, (b) is the layered graph corresponding to the user request, and (c) is the super source point graph.

[0049] Figure 3 This is a flow chart of the collaborative allocation method of computing, network and storage resources driven by both optimization and machine learning.

[0050] Figure 4This is a block diagram of the principle of the collaborative allocation method of computing, network and storage resources driven by optimization and machine learning.

[0051] Figure 5 This is a block diagram of the Actor-Critic architecture.

[0052] Figure 6 Schematic diagram of queue backlog in dynamic network scenarios; (a) and (b) are respectively schematic diagrams of queue backlog for node processing and link transmission in cloud-edge heterogeneous network resource scenarios, (c) and (d) are respectively schematic diagrams of queue backlog for node processing and link transmission in edge network resource scenarios, (e) and (f) are respectively schematic diagrams of queue backlog for cloud-edge network resource scenarios.

[0053] Figure 7 Schematic diagram of resource utilization in a dynamic network scenario; (a) schematic diagram of node resource utilization; (b) schematic diagram of link resource utilization. DETAILED DESCRIPTION

[0054] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0055] In order to facilitate the understanding of this solution, before introducing the collaborative allocation method of computing network storage resources, the multi-layer heterogeneous network model, user request, layered graph and super source point graph involved in this solution are introduced:

[0056] Multi-layer heterogeneous network model

[0057] This solution models the multi-layer network of transmission, computing, and storage as an undirected graph G = (V, E), where V represents the set of all network nodes and E represents the set of all physical links connecting network nodes. ― (i) and δ + (i) are the sets of neighbor nodes with traffic flowing into and out of node i. The multi-layer heterogeneous network model of this scheme assumes that each node has computing and storage capabilities, and can perform various service functions and cache various static objects under the condition that the amount of resources allows. For each node i∈V, use C i (t) represents the computation frequency in time slot t, and its physical meaning is the maximum number of executable instructions in a time slot; S i (t) represents the storage capacity of node i, that is, the maximum total volume of static objects that can be cached at this point.

[0058] For each link ij∈E, use Cij (t) represents the transmission bandwidth of link ij, that is, the maximum total transmission volume in the time slot. Define K as the set of all static objects; F k is the size of the static object k, k∈K. Note that in this model, C i , S i and C ij All of them are time-varying, which is based on the following considerations: first, the centralized or distributed management system of resources can dynamically and flexibly adjust the amount of resources to cope with fluctuations in user demand or maintain costs; second, due to equipment aging, failure, and upgrades, the amount of available resources may also change. Therefore, the present invention considers a network system with a certain degree of time-varying resource volume.

[0059] This model defines the cache vector x as one of the decision variables, indicating whether to cache the static object k at node i:

[0060] x={x i,k ∈{0,1}:i∈V,k∈K}

[0061] A set of feasible cache vectors needs to satisfy the following storage resource constraints:

[0062]

[0063] A static object can be stored in multiple network nodes, and the storage resource constraint indicates that the total volume of static objects that can be stored in a network node cannot exceed the storage capacity of that point.

[0064] User Request

[0065] This scheme defines that all user requests constitute a request set Φ, and all service functions required by each user request φ∈Φ are represented as service chains: a service chain has M φ In each stage, the user's real-time data must pass through the corresponding service function of each stage in sequence and produce the usable results of the next stage. At the same time, in order to process the real-time data of each unit of users, the service function of each stage also needs to accept a static object as input, which can be reused by different user requests and needs to be pre-cached in the network. The above process can be represented by a directed acyclic graph (DAG), see Figure 1 For every service function m∈M φ , this scheme uses a four-tuple To record service resource requirements:

[0066] The static object type required by the service function m to serve the request φ; is the aggregation rate of service function m serving request φ, The workload of service function m serving request φ, is the reduction factor of the service function m of the service request φ. Since the service function processes the data, the amount of data after processing may change. The reduction factor is the ratio of the data before and after processing.

[0067] The value of the above four-tuple is related to the user request φ itself, and has nothing to do with how it is deployed in the network or the time slot. Each user request has a starting point and an end point, indicating the source and destination of the real-time data flow. At each time slot t, the user traffic arrives in accordance with the Poisson distribution of λ, denoted as

[0068] Considering the characteristics of D-AgI chain service and static object selection, it is often necessary to Figure 2 Traffic planning is performed on the service graph shown in (a), which is very complex. Figure 2 (b) and Figure 2 The two graph structures and corresponding graph theory methods shown in (c) are used to assist modeling, simplify traffic planning ideas, and obtain feasible paths for traffic.

[0069] Hierarchical graph: Since user requests need to be processed and transmitted in different computing nodes in the form of service chains, this structure is complex and prone to loops, which is not conducive to the execution of graph algorithms. Therefore, a hierarchical graph is introduced to simplify it. Assume that the deployment of service functions is known, and the length of the service chain φ is M φ , an M can be constructed for each user request (corresponding to a service chain) φ Layered graph G φ . G φ Each layer in the network replicates the original physical network topology for data transmission for a specific service function. i If the i-th service function is deployed, the node is connected to the node n in the i+1th layer. i+1 Flowing through such a cross-layer link indicates that the data has been processed by service function i. φ On the Mth layer, user traffic is input from the corresponding source point on the first layer and φ The present invention uses the K shortest path algorithm to find the shortest path between source and sink points, and can obtain K feasible paths that meet the service chain constraints, and let their set be At the same time, considering that the amount of data will be scaled after being processed by each layer of service functions, the present invention further defines a cumulative scaling factor:

[0070]

[0071] In order to facilitate the understanding of the layered diagram, combined Figure 2Take (b) for example: Assuming the shortest path from A1 to C3 is: A1-C1-C2-D2-D3-C3, it means that in the original graph, the user traffic planning is: AC(F1)-D(F2)–C(F3). It can be seen that the traffic passes through the C node twice, which is why a hierarchical graph is used to simplify path calculation and avoid loops.

[0072] Super source graph: Assuming that the deployment location of the static object k is known, a new super source o is added to the original physical network topology. k , connect all nodes deployed with k∈K and set the weight to 0. After the construction is completed, the Dijkstra algorithm can be used to find o k A shortest path SP(o k ,j)={o k ,i,…,j}, indicating that transmitting static object k from node i to node j is the shortest transmission path.

[0073] For example: In the original graph (a), both nodes A and C deploy static object 1 (green). An auxiliary graph (c) is constructed, and the shortest paths from O1 to B and D are calculated to be O1-AB and O1-CD respectively. It can be seen that if B needs the static object O1, it should be called from node A, and D needs to be called from node C.

[0074] refer to Figure 3 , Figure 3 The paper shows a method for collaborative allocation of computing, network and storage resources based on optimization and machine learning. Figure 4 The principle block diagram of the method for collaborative allocation of computing, network and storage resources based on optimization and machine learning is shown; Figure 3 and Figure 4 As shown, the method S includes steps S1 to S6.

[0075] In step S1, all user requests to be allocated resources are obtained and stored in a request set; each user request corresponds to a service graph, including multiple service functions and static objects executed sequentially; user requests are content generated by artificial intelligence, multimodal data provided by users that need to be transmitted in large language models, or visual data of virtual reality, extended reality, holographic communication, and metaverse.

[0076] In step S2, the current network resource state and algorithm decision state and the user state of the service function m of the service request φ that has not been traversed in the request set are input into the trained DRL model to obtain the benefit weight w of the service function m of the service request φ and its required static objects placed on each node. f 、w sSpecifically, the output of the DRL model is 2×|V| values, and the first|V| values ​​and the last|V| values ​​of the 2×|V| values ​​are used as the benefit weights w f and w s , |V| is the number of nodes in the entire network combined with the number of nodes in V.

[0077] When implemented, the scheme preferably selects the user status of the service function m of the service request φ including the request amount of the service request φ The four-tuple of the starting point src, the end point dst and the service request is the sum of the user real-time traffic of all feasible paths σ on the service request φ at time slot t, and its expression is:

[0078]

[0079] Where σ is the feasible path on the service request φ at time slot t; The hierarchical graph G constructed for the service request φ φ In the example, the K shortest path algorithm is used to find the shortest path between the source and the destination, and the set of K feasible paths that meet the service request constraints is obtained.

[0080] In step S3, the benefit weight w f 、w s The weight matrix W is spliced ​​to the algorithm decision state respectively f and W s , and determine whether m is equal to the total number of service functions M of the service request φ φ If yes, go to step S4, otherwise update m=m+1 and return to step S2, 1≤m≤M φ ;

[0081] In step S4, determine whether the service request φ is equal to H. If so, proceed to step S5. Otherwise, update φ=φ+1 and return to step S2; 1≤φ≤H, H is the total number of service requests in the request set;

[0082] In step S5, the ILP optimization problem is solved to obtain the allocation strategies x and y of all service functions and required static objects on nodes for all service requests. The ILP optimization problem is to maximize the weight matrix W f and W s The weighted sum of the allocation strategies x and y.

[0083] When implemented, this scheme preferably maximizes the weight matrix W f and W s The expression of the weighted sum of the allocation strategies x and y is:

[0084]

[0085] st

[0086]

[0087] Among them, max is the maximum value; F k is the size of the static object k; K is the set of all static objects; x i,k ∈{0,1} is whether to cache static object k in node i; y i,φ,m (t) is whether the service function m of the service request φ is deployed at node i at time slot t; S i (t) is the storage resource available to node i at time slot t; x is the total number of all x i,k The allocation strategy composed of i,φ,m (t) is the allocation strategy composed of; ⊙ is the dot product.

[0088] In step S6, the LP optimization problem is solved to obtain the user real-time traffic a of the feasible path σ on the service request φ under the allocation strategy x, y at time slot t (φ,σ) (t), using allocation strategies x, y and user real-time traffic a (φ,σ) (t) As the final resource allocation strategy, the LP optimization problem is to minimize the short-term queue backlog.

[0089] In one embodiment of the present invention, the expression for minimizing the short-term queue backlog is:

[0090]

[0091] st

[0092]

[0093]

[0094] Among them, min is the minimum value; and are the virtual queues of node i and link ij at time slot t respectively; and are the accumulated traffic of node i and link ij in the physical topology at time slot t; The workload of the service function serving the request φ; is the link flow from node i at the mth layer to node i at the m+1th layer in the hierarchical graph of user request φ, is the link flow from node i at the mth layer to node j at the mth layer in the hierarchical graph of user request φ; C i (t) is the computation frequency of node i in time slot t; is the traffic on link ij associated with the static object k used in the service request; C ij(t) is the transmission bandwidth of link ij in time slot t; The hierarchical graph G constructed for the service request φ φ In , the K shortest path algorithm is used to find the shortest path between the source and the sink to obtain the set of K feasible paths that meet the service request constraints; σ is the feasible path on the service request φ at time slot t; The weight used to calculate the accumulated traffic of the physical topology; To allow the traffic of service request φ to flow through σ, the computational load borne by node i.

[0095] In the process of minimizing the short-term queue backlog, and The expressions are:

[0096]

[0097] in, and are the cumulative scaling factors of the mth and m-1th service functions in the service request φ, respectively; Indicates that when link i m i m+1 It is 1 when it belongs to the feasible path σ, and 0 when it does not belong to σ; is the reduction factor of the service function m-1 for the service request φ.

[0098] Virtual Queues and The expressions are:

[0099]

[0100]

[0101] in, and are the virtual queues of node i and link ij at time slot t-1 respectively; C i (t-1) is the calculation frequency of node i in time slot t-1; C ij (t-1) is the transmission bandwidth of link ij in time slot t-1; and are the accumulated traffic of node i and link ij in the physical topology at time slot t-1; [·] + Indicates that the result of the equation in this symbol must be greater than or equal to 0, otherwise it is forcibly set to 0.

[0102] flow The expression is:

[0103]

[0104] Among them, SP(ok ,l) is from the super source point o k The shortest path to node l; is the traffic on node l associated with the use of static object k in service requests at time slot t; is the traffic on node i associated with the use of static object k in service requests at time slot t; is the reduction factor of the service function m for the service request φ; For the service request φ in time slot t, after being processed by the service function, it is sent to link i m j m Traffic on The static object type required by the service function m to serve the request φ; d φ is the end point of the user request φ; i m is the node i at the mth layer of the hierarchical graph.

[0105] The DRL model of this solution is a model designed based on the Actor-Critic architecture; Figure 5 As shown in the figure, the Actor-Critic architecture increases stability by introducing a residual network module. Among them, the Actor network generates action probabilities based on the current state input. It uses a series of residual blocks, each of which contains two linear transformations and a ReLU activation function, as well as batch normalization to stabilize the learning process, and optionally uses dropout for regularization. This architecture helps alleviate the gradient vanishing problem during training by introducing short-circuit connections, enhancing the network's ability to learn deep representations. The final output layer uses a Sigmoid activation function to ensure that the output value is in the range of [0,1] to ensure that it can be interpreted as a probability. The Critic network evaluates the proposed action by estimating the value function, which represents the expected return for a given state-action pair. Specifically:

[0106] The Critic network evaluates the benefit weight w generated by the Actor network by estimating the value function f 、w s , the value function is determined by the parameters of the Critic network. The parameters of the Critic network are updated according to the reward of reinforcement learning. The expression of reward is:

[0107]

[0108] Among them, Reward is the model reward; Value is the target value obtained by solving the LP optimization problem, that is, C is a constant.

[0109] Similar to the Actor network, the Critic network uses a series of residual blocks to process the concatenated input of the state and action, providing a powerful feature extraction mechanism. The output is a single value from the linear layer, representing the value of the action under the current policy instructed by the Actor network.

[0110] In order to verify the effect of this solution, the following is a detailed description with specific examples:

[0111] Experimental scenario: In terms of network topology selection, the present invention adopts NSFNET with 14 nodes and USBNET with 24 nodes; in terms of network resource quantity, the present invention sets the node calculation frequency C i , node storage capacity S i and link transmission bandwidth C ij The scope includes resource configuration of three network scenarios: edge, cloud-edge heterogeneous and cloud-end; on the user request service graph, the present invention sets the service graph quadruple Range and volume of static objects F k range, and randomly generate a service graph for each user from it. It is worth noting that the present invention uses Zipf distribution to allow overlapping static object requirements between different users, making it possible to utilize the sharing of static objects; Regarding the user request volume, the user request volume in the experiment obeys the Poisson distribution of parameter λ, and the starting point and end point of the traffic are randomly generated; the number of user requests is the size of the request set |Φ|; the number of service functions for each request is M φ ; The setting of simulation parameters can refer to Table 1.

[0112] Table 1 Simulation parameter settings

[0113]

[0114] The algorithm proposed in the present invention (hereinafter referred to as Bridge algorithm) takes each stage of each user request as input and determines the final action through multiple outputs.

[0115] Comparison Algorithms: For performance comparison, this paper considers two benchmark methods:

[0116] Greedy Random Algorithm: (1) Sort static objects in descending order by size, and then evenly distribute them to nodes, ensuring that each static object is cached at least once. Then, randomly select a node and a static object, check whether they can be cached, and repeat this process until all nodes can no longer cache; (2) Find a shortest path between the starting point and the end point of a specific user request, and map the service functions to the nodes on the path equidistantly. Then, on the premise of ensuring that each service function is deployed at least once, randomly select a mapping list, add, delete, or replace the mapping position of a service function in it, and iterate this process 1000 times. It should be noted that this algorithm performs the deployment of service functions and the caching of static objects separately, and does not consider joint optimization.

[0117] DRL algorithm: In the DRL comparison algorithm designed based on LyDROO, all stages of all user requests are directly taken as input, and all deployment and caching actions are output at one time. This method has a larger state space and action space, and the convergence difficulty is increased. Therefore, the DRL algorithm limits the model output to a continuous value of [0, 1], and uses a quantizer to quantize the output 10 times to obtain 10 groups of feasible 0-1 integer action plans. The 10 groups of plans are applied to LP solution respectively, and the solution with the best solution result is selected for training. The algorithm jointly optimizes the deployment of service functions and the caching of static objects. At the same time, the use of quantizers speeds up model training and convergence, improves the quality of algorithm solution, but also introduces high time overhead. In addition, it is necessary to introduce a penalty term for violating constraints in the reward function so that the DRL algorithm produces output that meets the constraints of the optimization problem.

[0118] Performance indicators:

[0119] Average queue backlog: The present invention uses the average value of virtual queue backlog to reflect the stability of the system, including node processing queue backlog and link transmission queue backlog. Queue backlog shows the effectiveness of the current resource allocation algorithm. Less backlog means faster processing and routing of data packets, thereby improving overall network performance.

[0120] Average resource utilization: By counting the utilization of computing power and transmission resources in the entire network by the resource allocation scheme, it can help to explain whether the resource allocation algorithm is reasonable. For example, if the average queue backlog is large, but the average resource utilization is not high, it means that the algorithm has caused unnecessary hotspot congestion and limited the upper limit of the traffic that can be accommodated.

[0121] The greedy-random algorithm, the DRL algorithm and the Bridge algorithm of the present invention are all simulated under the simulation conditions shown in Table 1. The simulation results of these three methods can be referred to Figure 6 and Figure 7 :

[0122] Figure 6 The performance of the Bridge algorithm and its two benchmark methods in different network resource scenarios is demonstrated. Here, the user quantity |Φ| is set to 10, and all user requests are randomly updated every 10 time slots; in each time slot, the network resource quantity will randomly change by ±10%, simulating a highly dynamic network. It is worth mentioning that considering that the system will crash when the network queue increases to a certain extent, the present invention sets 1300 as the upper limit of the queue backlog. Exceeding this value indicates that the resource allocation scheme is inappropriate or the current network resource quantity cannot bear the user request volume. Figure 6 The average data queue backlog changes within 3000 runtime slots in three resource scenarios are plotted in Figure 1. At the same time, the average resource utilization of each algorithm in the three scenarios is plotted in Figure 2. Figure 7 In the example, L and H are used to distinguish the performance of the same algorithm under different λ.

[0123] from Figure 6 As can be seen from (a) and (b), the node processing queue backlog of the Greedy Random algorithm is very low, but at the expense of a high backlog in the link transmission queue. When λ = 55Mbps, the transmission queue frequently reaches the upper limit of the backlog; when λ is reduced to 45Mbps, the transmission queue backlog reaches the upper limit, but there is still an extremely high backlog (750Mb on average). This is because the greedy strategy of the Greedy Random algorithm only considers the allocation of resources requested by each user along the shortest path, ignoring the global load balancing and the joint optimization of transmission, computing, and storage resources, which can easily cause unbalanced resource allocation and local congestion.

[0124] pass Figure 7 It can be observed from (a) and (b) that the Greedy Random algorithm simultaneously presents the highest link resource utilization and the lowest node resource utilization, which supports the rationality of the above analysis. In contrast, the DRL and Bridge algorithms with global optimization considerations can also stabilize the queue backlog when λ=55Mbps. Among them, the processing queue backlog of Bridge is slightly higher than that of the DRL algorithm, but the transmission queue backlog is very slight (less than 100Mb). The transmission queue backlog of the DRL algorithm is 300Mb on average, and has certain fluctuations, up to 750Mb. From Figure 7 It can be seen that Bridge(L) has a higher node computing power utilization and a lower link resource utilization than DRL. To a certain extent, it transfers the link load to the node, resulting in a slight increase in the processing queue backlog and a significant decrease in the transmission queue backlog. This shows that Bridge has learned how to jointly optimize the transmission, computing, and storage resources.

[0125] Finally, the present invention attempts to increase the λ of DRL and Bridge to see whether it can carry higher user request traffic. After a long period of training, the DRL algorithm is still difficult to converge, either the output cannot meet the constraints, or the queue backlog will reach the upper limit. This is because the optimization strategy is more difficult to explore under limited resources and high load. The Bridge algorithm can converge quickly even when λ=70Mbps. It can be predicted that as the load increases, the processing queue backlog of the Bridge algorithm continues to rise to an average of 450Mb compared to λ=55Mbps, and the backlog of the transmission queue is similar to that of the DRL algorithm when λ=55Mbps. On the one hand, thanks to the conversion of multi-stage actions and the constraint guarantee of the ILP solver, the Bridge algorithm is significantly better than the DRL algorithm in convergence performance, making it possible to carry more user traffic; on the other hand, the performance of Bridge in the transmission queue backlog when λ=70Mbps once again reflects the superiority of the Bridge resource allocation strategy.

[0126] exist Figure 6 In (c) and (d), we can see that as the amount of network resources decreases, the Greedy Random algorithm, which originally had a stable transmission queue when λ=45Mbps in the cloud-edge network, has a backlog that increases and reaches the upper limit in the edge network. When λ drops to 40Mbps, the backlog is still very high, near the upper limit. In contrast, the DRL and Bridge algorithms can still maintain a high load. Compared with the cloud-edge heterogeneous scenario, the queue backlogs for processing and transmission have increased slightly, and are more volatile, but can still be stabilized in an acceptable backlog range.

[0127] And in Figure 6 In (e) and (f), as the amount of resources increases, the queue backlog of each algorithm is significantly improved, and the queue backlog of the DRL and Bridge algorithms does not exceed 100Mb. The Greedy Random algorithm can already support a load of λ=55Mbps, and the queue backlog is not high. However, it still cannot withstand a load of λ=60Mbps. At the same time, the DRL algorithm can also converge under a load of λ=60Mbps, because the constraint is easier to satisfy. It is worth noting that in the three scenarios, the node and link resource utilization characteristics of each algorithm are similar, that is, for Figure 6 The analysis of (a) and (b) in the above also applies to Figure 6 (c and (d) and Figure 6 (e) and (f) were analyzed.

Claims

1. A method for collaborative allocation of computing, network and storage resources based on optimization and machine learning, characterized in that: Includes steps: S1. Obtain all user requests to be allocated resources and store them in a request set; each user request corresponds to a service graph, including multiple service functions and static objects executed sequentially; S2. Input the current network resource status, algorithm decision status, and user status of the service function m of the service request φ that has not been traversed in the request set into the trained DRL model, and obtain the service function m of the service request φ and the benefit weight w of the static objects required by each node. f 、w s ; S3. Set the benefit weight w f 、w s The weight matrix W is spliced ​​to the algorithm decision state respectively f and W s , and determine whether m is equal to the total number of service functions M of the service request φ φ If yes, go to step S4, otherwise update m=m+1 and return to step S2, 1≤m≤M φ ; S4, determine whether the service request φ is equal to H, if so, proceed to step S5, otherwise update φ=φ+1, and return to step S2; 1≤φ≤H, H is the total number of service requests in the request set; S5. Solve the ILP optimization problem to obtain the allocation strategies x and y of all service functions and required static objects on nodes for all service requests. The ILP optimization problem is to maximize the weight matrix W. f and W s The weighted sum of the allocation strategies x and y; S6. Solve the LP optimization problem to obtain the user real-time traffic a on the feasible path σ on the service request φ under the allocation strategy x, y in time slot t (φ,σ) (t), using allocation strategies x, y and user real-time traffic a (φ,σ) (t) As the final resource allocation strategy, the LP optimization problem is to minimize the short-term queue backlog.

2. The method for collaborative allocation of computing, network and storage resources according to claim 1, characterized in that: The user state of the service function m of the service request φ includes the request amount of the service request φ The four-tuple of the starting point src, the end point dst and the service request The static object type required by the service function m to serve the request φ; is the aggregation rate of service function m serving request φ, The workload of service function m serving request φ, is the reduction factor of the service function m for the service request φ; is the sum of the user real-time traffic of all feasible paths σ on the service request φ at time slot t, and its expression is: Where σ is the feasible path on the service request φ at time slot t; The hierarchical graph G constructed for the service request φ φ In the example, the K shortest path algorithm is used to find the shortest path between the source and the destination, and the set of K feasible paths that meet the service request constraints is obtained.

3. The method for collaborative allocation of computing, network and storage resources according to claim 1 or 2, characterized in that: Maximize the weight matrix W f and W s The expression of the weighted sum of the allocation strategies x and y is: st Among them, max is the maximum value; F k is the size of the static object k; K is the set of all static objects; x i,k ∈{0,1} is whether to cache static object k in node i; y i,φ,m (t) is whether the service function m of the service request φ is deployed at node i at time slot t; S i (t) is the storage resource available to node i at time slot t; x is the total number of all x i,k The allocation strategy composed of i,φ,m (t) is the allocation strategy composed of; ⊙ is the dot product.

4. The method for collaborative allocation of computing, network and storage resources according to claim 3, characterized in that: The expression for minimizing the short-term queue backlog is: st Among them, min is the minimum value; and are the virtual queues of node i and link ij at time slot t respectively; and are the accumulated traffic of node i and link ij in the physical topology at time slot t; The workload of the service function serving the request φ; is the link flow from node i at the mth layer to node i at the m+1th layer in the hierarchical graph of user request φ, is the link flow from node i at the mth layer to node j at the mth layer in the hierarchical graph of user request φ; C i (t) is the computation frequency of node i in time slot t; is the traffic on link ij associated with the static object k used in the service request; C ij (t) is the transmission bandwidth of link ij in time slot t; The hierarchical graph G constructed for the service request φ φ In , the K shortest path algorithm is used to find the shortest path between the source and the sink to obtain the set of K feasible paths that meet the service request constraints; σ is the feasible path on the service request φ at time slot t; The weight used to calculate the accumulated traffic of the physical topology; To allow the traffic of service request φ to flow through σ, the computational load borne by node i.

5. The method for collaborative allocation of computing, network and storage resources according to claim 4, characterized in that: and The expressions are: in, and are the cumulative scaling factors of the mth and m-1th service functions in the service request φ, respectively; Indicates that when link i m i m+1 It is 1 when it belongs to the feasible path σ, and 0 when it does not belong to σ; is the reduction factor of the service function m-1 for the service request φ.

6. The method for collaborative allocation of computing, network and storage resources according to claim 4, characterized in that: Virtual Queues and The expressions are: in, and are the virtual queues of node i and link ij at time slot t-1 respectively; C i (t-1) is the calculation frequency of node i in time slot t-1; C ij (t-1) is the transmission bandwidth of link ij in time slot t-1; and are the accumulated traffic of node i and link ij in the physical topology at time slot t-1; [·] + Indicates that the result of the equation in this symbol must be greater than or equal to 0, otherwise it is forced to 0.

7. The method for collaborative allocation of computing, network and storage resources according to claim 4, characterized in that: flow The expression of (t) is: Among them, SP(o k ,l) is from the super source point o k The shortest path to node l; is the traffic on node l associated with the use of static object k in service requests at time slot t; is the traffic on node i associated with the use of static object k in service requests at time slot t; is the reduction factor of the service function m for the service request φ; For the service request φ in time slot t, after being processed by the service function, it is sent to link i m j m Traffic on The static object type required by the service function m to serve the request φ; d φ is the end point of the user request φ; i m is the node i at the mth layer of the hierarchical graph.

8. The method for collaborative allocation of computing, network and storage resources according to claim 4, characterized in that: The DRL model is a model designed based on the Actor-Critic architecture; the Critic network evaluates the benefit weight w generated by the Actor network by estimating the value function f 、w s , the value function is determined by the parameters of the Critic network, and the parameters of the Critic network are updated according to the reward Reward of reinforcement learning. The expression of reward Reward is: Among them, Reward is the model reward; Value is the target value obtained by solving the LP optimization problem, that is, C is a constant.

9. The method for collaborative allocation of computing, network and storage resources according to any one of claims 1 to 8, characterized in that: The output of the DRL model is 2×|V| values, and the first |V| values ​​and the last |V| values ​​of the 2×|V| values ​​are respectively used as the benefit weights w f and w s , |V| is the number of nodes in the entire network combined with the number of nodes in V.

10. The method for collaborative allocation of computing, network and storage resources according to any one of claims 1 to 8, characterized in that: The user request is for artificial intelligence generated content, multimodal data provided by the user that needs to be transmitted in a large language model, or visual data in virtual reality, augmented reality, holographic communication, and the metaverse.

Citation Information

Patent Citations

  • Resource allocation method for computing power network service function chain

    CN114710791A

  • Cloud edge resource collaborative scheduling method and system for computing network integration

    CN117950860A

  • Edge-end collaborative intelligent unloading method

    CN118394512A

  • Method for task offloading based on power control and resource allocation in industrial internet of things

    US20220377137A1