DRL-based SFC robust deployment method under dynamic computing power network
By introducing probabilistic constraint optimization and deep reinforcement learning into the dynamic computing network, the instability of SFC deployment caused by computing resource fluctuations is solved, achieving efficient and robust deployment in dynamic environments, reducing resource constraint violation rate and communication latency, and improving the system's adaptability and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-17
AI Technical Summary
In dynamic computing networks, existing technologies struggle to achieve stable deployment of SFC (System-Functional Computing) when faced with highly heterogeneous and dynamically fluctuating computing resources, leading to performance fluctuations and service interruption risks. Furthermore, traditional methods incur excessive computational overhead in large-scale networks, making it difficult to achieve efficient and robust deployment.
A robust deployment method for SFC based on deep reinforcement learning (DRL) is adopted. By constructing a probabilistic constraint optimization model and a hybrid Newton-bisection algorithm, combined with deep reinforcement learning (DRL) and convolutional neural networks (CNN), the network topology and resource status are perceived in real time, and adaptive VNF deployment and routing paths are generated to achieve robust deployment against resource fluctuations.
It significantly reduces resource constraint violation rates in dynamic resource environments, improves service quality and throughput, reduces communication latency and resource costs, adapts to large-scale network environments, and enhances system resilience and efficiency.
Smart Images

Figure CN121887876A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network communication technology, specifically relating to a robust SFC deployment method based on DRL in a dynamic computing network. Background Technology
[0002] Service Function Chain (SFC) deployment, as a core research topic in Network Functions Virtualization (NFV) environments, aims to address the diverse and latency-sensitive service delivery challenges. Currently, research on SFC deployment in Computing Networks (CPNs) can be mainly divided into deployment in heterogeneous multi-domain environments, deployment in Mobile Edge Computing (MEC) environments, and robust deployment for uncertain service demands.
[0003] To improve SFC deployment efficiency, some existing studies aim to minimize deployment costs and balance network load by optimizing VNF placement, reducing cross-domain overhead, and utilizing deep learning techniques for intelligent orchestration. However, most of these methods focus on static network environments or only consider random fluctuations in service demand (such as traffic load uncertainty), often neglecting the dynamic fluctuations in resource conditions within the underlying computing network infrastructure.
[0004] In a computing network environment, computing resources are highly heterogeneous and distributed. The available capacity of the resource pool often fluctuates randomly due to resource contention caused by hardware maintenance, system upgrades, and coexisting background tasks (such as online transactions and payment processing). This uncertainty in infrastructure resources can lead to significant performance fluctuations and significantly increase the risk of service interruptions. For ultra-low latency applications such as holographic communication and digital twins in the 6G era, ignoring the dynamic changes in resource availability will make it difficult to ensure the continuous and reliable operation of SFC (Sustainable Fiber Communication).
[0005] While robust optimization techniques have been partially introduced to enhance network resilience, significant research gaps remain regarding how to accurately quantify the distribution uncertainty of computing resource pools in dynamic computing power network scenarios and implement scalable adaptive deployment strategies in large-scale heterogeneous networks. Furthermore, traditional integer linear programming (ILP) solutions face scalability bottlenecks in large-scale networks due to excessive computational overhead and NP-hard characteristics. Therefore, developing a robust SFC deployment scheme that can cope with real-time changes in computing resource availability while ensuring service efficiency under high load is a critical issue that urgently needs to be addressed in the current evolution of computing power networks. Summary of the Invention
[0006] Purpose of the invention: In order to minimize the communication latency and resource usage cost of SFC in a dynamic computing network environment, while maximizing service throughput and strictly ensuring the stability of resource constraints, and to meet the stringent requirements of latency-sensitive services in the 6G era, this invention provides a robust deployment method of SFC based on DRL in a dynamic computing network.
[0007] Technical Solution: The present invention provides a robust deployment method for SFC based on DRL in a dynamic computing network. The method is based on the user submitting SFC service request information to the computing network controller. The computing network controller analyzes the service demand characteristics and QoS constraints based on the service request information, including computing resource requirements, bandwidth resource requirements, VNF sequence structure, and maximum tolerable latency requirements. At the same time, combined with the computing network topology, the dynamic fluctuation status of the distributed resource pool, and the network link operation status, a robust optimization mechanism combining probabilistic constraints and deep reinforcement learning technology is used to optimize the node allocation and virtual link mapping of SFC. The method includes the following steps: (1) Construct a computing power network and its corresponding system model, which includes a computing power controller, a distributed computing resource pool and network links. The computing power controller senses the dynamic heterogeneous characteristics of the computing power network in real time, obtains the computing power, link bandwidth capacity and propagation delay physical parameters of the resource pool, and receives an SFC request set composed of source node, target node, VNF sequence and QoS constraints, and formalizes it into a deployment task to be processed. (2) In view of the random fluctuation characteristics of the distributed computing resource pool, a reference distribution is extracted from historical data and prediction information. The uncertainty set between the actual resource distribution and the reference distribution is quantified by KL divergence. A fault tolerance threshold is introduced to adjust the degree of conservatism. The hard resource capacity constraint is reconstructed into a probability constraint model to realize the mathematical measurement of resource fluctuation risk. (3) The probabilistic constraint optimization problem is transformed by the Lagrange multiplier method and the KKT optimality condition. A hybrid algorithm combining the Newton iteration method and the bisection method is used to search for and determine the robust SFC deployment decision threshold that satisfies the fault tolerance upper limit. The robust SFC deployment decision threshold serves as a safety boundary for resource allocation, mapping the underlying random and uncertain constraints to the upper-layer computable robust criteria. (4) Based on the SFC deployment decision threshold, construct a multi-objective optimization model with the comprehensive objectives of minimizing end-to-end communication latency, reducing resource usage costs, and maximizing total service throughput; The objective optimization model follows VNF placement constraints, traffic conservation constraints, and resource capacity robustness constraints to ensure that the generated deployment strategy can withstand resource fluctuations in the worst case while meeting stringent QoS requirements. (5) The SDPC algorithm based on the approximate policy optimization framework and convolutional neural network is called to solve the problem. The high-dimensional network topology and resource features are extracted by the convolutional neural network. The policy network generates VNF deployment actions and combines the weighted Dijkstra algorithm to complete the routing path mapping. Finally, a robust SFC deployment scheme with adaptive capabilities is output.
[0008] Furthermore, the system architecture described in step (1) consists of a computing power controller and several distributed computing resource pools. Multiple routers and physical links The resulting computing network (CPN). Resource pool. The computing power is denoted as ,link The bandwidth capacity is denoted as ,link The propagation delay is recorded as .node Indicates link The initial resource pool or router, and Indicates link The terminated resource pool or router. Node The cost of a unit of computational resources is defined as follows: ,link The cost of a unit of bandwidth resource is defined as .
[0009] In the invention, This represents the set of SFC requests. Each SFC request... By quintuple Description, in which and These represent the source node and the target node, respectively. (SFC request) (SFC) The traffic demand is denoted as , Indicates SFC request Maximum tolerable latency. SFC request. The topology is composed of Given, among which Indicates SFC request VNF set, SFC A set of virtual links. An SFC consists of different VNF instances. Assume there exists... There are VNF types, and their sets are denoted as . .variable VNF The type.
[0010] Using the CPN model and the SFC request model, an SFC deployment decision matrix is defined. and as follows: (1) (2) binary variables Indicates virtual network function Whether deployed in the computing resource pool Above, and binary variables Indicates virtual link Is it mapped to a physical link? The success of SFC deployment depends on meeting resource availability and Quality of Service (QoS) requirements. Among these variables... Request on behalf of SFC The deployment status indicates whether the deployment was successful.
[0011] (3) (4) (5) Virtual Network Function The computing resources required to process one unit of flow are denoted as . The delay required for this function to process one unit of traffic is denoted as . .
[0012] The core requirement of CPN is to determine the time period during CPN operation. Within this framework, the deployment strategy for SFC must meet the resource constraints of the CPN while ensuring the Quality of Service (QoS) of the SFC. These requirements are stated below: (6) (7) (8) Constraint (6) ensures that the computing resources used within the resource pool do not exceed the available capacity. Constraint (7) stipulates that bandwidth consumption must be within the link bandwidth limit to prevent overuse. Constraint (8) guarantees the quality of service communication, ensuring that all service requirements are met without performance degradation.
[0013] use Indicates SFC request The cost of using computing resources is defined as follows: (9) The resource pool in a CPN comprises servers from various vendors. Therefore, any changes or updates to these servers will affect available computing resources. Furthermore, the CPN not only manages new tasks but also strives to optimize resource utilization and profitability by serving a diverse user base. This includes supporting online gaming, providing email services to businesses, and facilitating third-party services such as payment processing. Therefore, the computing resources in the resource pool can be viewed as random variables with unknown probability distributions. Thus, establishing stochastic models to address resource uncertainty is crucial.
[0014] First, a formalized expression of the SFC deployment problem in the Dynamic Computing Power Network (SDCPN) is given: (10)
[0015] Equation (10) integrates optimization objectives, including minimizing communication latency, reducing resource usage costs, and maximizing the total throughput of all SFCs. yes The vector, It means The vector is described. The deployment constraints of SFC are outlined, and the uncertainty of computing resources within the resource pool is considered. Subsequently, an uncertainty model to describe the volatility of the resource pool's computing power will be presented.
[0016] Characterizing the computational power of a resource pool is often challenging. In the optimization model of this invention, the random variables in the optimization model... The calculations are both tedious and difficult to perform. Furthermore, in practice, it is often impossible to know precisely the... The inherent distributional characteristics of random variables mean that solutions based on assumed distributions may lack rationality. Traditional methods describe the volatility of random variables using variance or second moment, but these indicators may not adequately characterize their intrinsic distributional properties.
[0017] This invention innovatively extracts a reference distribution from historical data and forecast information, rather than relying on moment statistics to capture distribution characteristics. Given that the computing power of a resource pool fluctuates over time and is difficult to describe, this invention uses an empirical distribution as an effective reference, allowing the actual distribution to fluctuate around it. For example, assuming the computing power distribution of the resource pool... Around the known distribution The fluctuations, the distribution of which can be obtained through prediction and derivation from long-term field measurement data.
[0018] Its reference distribution The difference between them can be quantified using a probabilistic distance metric called KL divergence. KL divergence measures the difference between the distributions of two probabilities. Let these two distributions be... and .generally, This represents the true distribution obtained through accurate modeling, while This is based on theoretical assumptions and simplified approximations. The KL divergence between two continuous distributions is defined as follows: (11) in Let the integration domain be . When the distribution and When they are close to each other, the distance metric approaches zero. By adjusting the KL divergence, the distribution uncertainty set is defined as follows: (12) in This represents the distance limit, a value that can be obtained through empirical data or real-time measurements. It indicates the degree of variation in computing power.
[0019] Considering the distribution of computational capacity Its reference distribution is Distance limit is Then calculate the capacity distribution The following constraints must be met: (13) (14) Equation (14) above expresses the fundamental property that the total integral of the probability density function over its domain is equal to 1. Combining (13) and (14), the constraints in (6) can now be restated to solve the optimization problem defined in (10) more efficiently.
[0020] Furthermore, the opportunity constraint optimization in step (2) is as follows: For equation (6) above, the computational capability constraint can be expressed as: .
[0021] In practical applications, decision-making criteria need to reasonably set decision variables. To satisfy the condition constraints in (6), a parameter can be introduced. Adjusting the level of conservatism transforms the above conditions into probabilistic constraints: (15) Here, This represents the fault tolerance threshold of the computing power network (CPN), quantifying the maximum permissible probability that available computing power cannot meet the required level. Based on this, the condition can be restated as: , (16) Equivalent to: (17) definition For robust SFC deployment decisions, this value is equivalent to a time slot. internal nodes The computational resource consumption is calculated. Therefore, the following auxiliary function is introduced: (18) The left side of inequality (17) can be transformed into an optimization problem: , (13), (14), (19) definition This represents the probability of resource violation in the worst-case scenario. From this, we can obtain the worst-case mapping. This mapping will enable robust SFC deployment decisions. Mapped to .
[0022] (20)
[0023] Furthermore, the robust SFC deployment decision threshold is determined in step (3) as follows: Because there are random variables in the constraints The SFC deployment problem cannot be solved directly. Therefore, the subproblem aims to find a robust SFC deployment decision threshold. The CPU constraint (6) is transformed into a solvable form.
[0024] Theorem 1: Problem formula (19) is a convex optimization problem.
[0025] The proof of this theorem is as follows: Rewrite formula (19) as: (twenty one)
[0026] Note the relationship between the objective function and the equality constraint function. It exhibits a linear variation. Furthermore, the convexity of the inequality constraint function can be proven.
[0027] Lemma 1: If the function If it is a convex function, then its perspective mapping satisfy: (twenty two) Its domain is: (twenty three) The mapping maintains convexity.
[0028] If function If it is a convex function, then its perspective function is... It is also a convex function. This conclusion can be proven in various ways, such as directly verifying the defined inequality, or... The above uses a top-down view and perspective mapping.
[0029] Considering Convex functions defined on Its perspective function is: (twenty four) And in The top is convex. Function express and The relative entropy between them. Therefore, the distribution and The KL divergence between them is It is convex. This demonstrates that the inequality constraint applies to the distribution. It has convexity.
[0030] Using Theorem 1 and the Slater condition, we proved that problem (19) has strong duality. The worst-case failure probability was derived using the Lagrange method. as follows: (25) in and These are the Lagrange multipliers related to the constraints of problem (19). Let, (26) about The derivative can be derived as: (27) By applying the KKT optimality condition, we can obtain: (28) (29) (30) (31) From formula (28), the optimal distribution function can be expressed as: (32) The two variables in formula (32) A reasonable choice must be made to satisfy conditions (29)-(31). Specifically, the following results are obtained: Theorem 2: Point Pairs The following nonlinear equations are satisfied: (33) (34)
[0031] and Represents the reference distribution The cumulative distribution function.
[0032] The proof of Theorem 2 can be obtained directly by substituting (32) into equations (29) to (31). However, deriving explicit solutions from equations (33) and (34) remains challenging. Unlike the standard Newton iteration, which relies solely on local gradient information and is susceptible to instability due to initial estimation, the method in this invention incorporates a bisection strategy during the iteration process. This Newton-bisection hybrid method combines the fast convergence of Newton's method with the global robustness of the bisection method, thereby simultaneously improving convergence reliability and numerical stability within the feasible region.
[0033] After finding the solutions to (33) and (34) in Theorem 2, the worst-case failure probability can be derived from formulas (28) and (31) as follows: (35) Next, we determine the decision threshold for robust SFC deployment. , making This involves solving... The inverse function of , which cannot be directly derived from equation (35). However, The following properties can provide a basis for designing search methods.
[0034] Theorem 3: Failure Probability in the Worst Case With robust SFC allocation decision Monotonically increasing.
[0035] The derivation of Theorem 3 is direct, because
[0036] Although the above equation cannot be solved directly, Inspired by the monotonicity, a binary search method is used to satisfy... The solution. This method requires the interval... Search within, among which This is an empirical constant, ensuring that it satisfies... .
[0037] To obtain the robust SFC deployment decision threshold under constraint (6) Then, (6) is transformed into the following constraint: (36) Further, step (4) determines the robust SFC deployment decision threshold as follows: After determining the robust SFC deployment decision threshold, the SFC deployment process will proceed. The robustness problem is defined as follows: (37) Because the units of the objective functions are different, weights are used. , and To adjust or balance the various objectives.
[0038] Furthermore, the SFC deployment algorithm based on DRL in step (5) is as follows: This example presents the SFC deployment algorithm SDPC based on Deep Reinforcement Learning (DRL). This algorithm integrates the PPO framework with Convolutional Neural Networks (CNNs) to learn a near-optimal deployment strategy in a data-driven manner. This example proposes the SDPC algorithm, which uses network topology... Service request set As input, the algorithm outputs an efficient SFC deployment scheme that includes node resource allocation and path selection. During multiple training iterations, the algorithm processes service requests sequentially, and the output probability distribution from the policy network... The deployment actions of the sampling virtual network function (VNF) are limited to the action space.
[0039] Each training iteration incurs latency, resource utilization, and throughput. The learning objective utilizes the truncation of the PPO (Procedure-Based Error) loss function, combining policy loss, value function error, and entropy regularization to balance exploration and exploitation while ensuring stable convergence. Indicates the current strategy Compared with previous strategies In state Next action The probability ratio. This represents the advantage function, used to quantify the relative return of a specific action relative to the expected value of the current strategy. This represents the estimated value of the current state, while Corresponding to the value of the real state, it is obtained through time difference learning. Parameters A truncation threshold for policy update magnitude is defined to ensure a stable and conservative optimization process. The gradients of this loss function with respect to the policy network, value network, and shared convolutional neural network parameters are calculated via backpropagation, and a joint update is performed using the Adam optimizer, achieving efficient end-to-end training within the SDPC framework.
[0040] (38) (39) The robust optimization method employs an outer binary search approach to iteratively narrow the search interval to ensure global convergence, while the inner Newton iteration enables fast local convergence within a finite number of steps. The overall computational complexity is primarily influenced by the parameter dimensionality. The impact can be approximated as ,in and These represent the number of iterations for the bisection method and Newton's method, respectively.
[0041] The computational complexity of an integer linear programming solver is mainly determined by the number of decision variables, and its expression is: ,in Indicates the size of the set of decision variables. This represents the number of virtual network functions corresponding to each decision variable. Indicates the number of network nodes. This represents the number of links. Given the NP-hard nature of Integer Linear Programming (ILP), its worst-case computational complexity increases exponentially. It grows exponentially, formally expressed as This greatly limits the applicability of this method in large-scale scenarios.
[0042] The complexity of the SDPC algorithm proposed in this invention mainly stems from three components: feature extraction based on a convolutional neural network (CNN), forward and backward propagation in the policy and value networks, and route calculation. Specifically, the computational cost of a CNN composed of multiple two-dimensional convolutional layers increases exponentially, and its complexity depends on the number of convolutional layers. Kernel size Number of channels and input feature map size Approximately ,in This corresponds to the complexity of a fully connected layer. The policy network selects actions by sampling from the output probability distribution, resulting in relatively low computational overhead. The routing path is calculated using a weighted Dijkstra algorithm, with a complexity of O(n log n). Considering that training has iterative properties and the total number of steps is... The overall training complexity is .
[0043] Furthermore, it is pointed out that although the robust optimization problem (Equation 37) can be solved using integer linear programming (ILP) with the help of commercial solvers such as Gurobi, this method suffers from limited scalability and excessive computational overhead in large-scale networks. To address this challenge, a SFC deployment algorithm SDPC based on deep reinforcement learning (DRL) is proposed. This algorithm integrates the approximate policy optimization (PPO) framework with convolutional neural networks (CNN) to learn near-optimal deployment policies in a data-driven manner.
[0044] Unlike directly applying PPO to the SFC deployment problem, several innovations for CPN are introduced: first, a state representation is constructed to capture the spatiotemporal fluctuations of computing resource availability and VNF deployment constraints; second, a reward function is designed to balance service reliability and resource efficiency under CPN conditions.
[0045] Compared to the ILP solver, SDPC significantly reduces computational complexity and is better suited to dynamic, large-scale CPN environments.
[0046] Beneficial Effects: This invention presents a robust SFC deployment method based on DRL in dynamic computing networks (SDCPN). This method addresses the problem of deploying SFCs in dynamic computing networks, where resource availability within the computing power pool fluctuates randomly. By introducing uncertainty sets and chance constraints, the problem is reformulated as a robust optimization problem, aiming to maximize overall throughput while minimizing communication latency and resource costs. A robust optimization algorithm combining Newton's iteration method and bisection is proposed to determine the threshold of the computing resource pool, followed by solving the problem using integer linear programming. To improve scalability and applicability in large-scale networks, a deep stochastic learning-based SDPC algorithm is developed. Finally, experimental evaluation shows that the robust optimization scheme significantly reduces resource constraint violation rates while achieving performance comparable to deterministic SFC deployment methods, thus ensuring enhanced robustness of the system in uncertain environments. Furthermore, the proposed SDPC algorithm achieves near-optimal performance in small-scale networks and surpasses state-of-the-art methods in large-scale network scenarios. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating the method described in this invention; Figure 2 This is a schematic diagram illustrating the service provision in the computing power network in an embodiment of the present invention; Figure 3 This is a diagram of the SDPC framework for SFC deployment based on PPO-CNN in an embodiment of the present invention; Figure 4 shows the convergence diagram of the algorithm in a small-scale network in the embodiment of the present invention, wherein: Figure 4(a) is a convergence diagram comparing the training loss of the algorithm, and Figure 4(b) is a convergence diagram comparing the value loss of the algorithm; Figure 5 This is a comparison chart of the average reception rate of SFC under different traffic demands in a small-scale network according to an embodiment of the present invention; Figure 6 shows the performance evaluation of the SFC algorithm under different traffic requirements in a small-scale network in the embodiment of the present invention, wherein: Figure 6 (a) is the average optimization objective function value of SFCs under different traffic requirements, Figure 6 (b) is the average end-to-end delay of SFCs under different traffic requirements, and Figure 6 (c) is the average resource usage cost of SFCs under different traffic requirements. Figure 7 This refers to the average resource utilization rate of SFC under different traffic demands in a small-scale network in this embodiment of the invention. Figure 8 This represents the average resource constraint violation rate of SFCs under different traffic demands in a small-scale network according to the embodiments of the present invention. Figure 9 shows the convergence diagram of the algorithm under a large-scale network in the embodiment of the present invention, wherein: Figure 9(a) is a convergence diagram comparing the training loss of the algorithm, and Figure 9(b) is a convergence diagram comparing the value loss of the algorithm; Figure 10 This refers to the average reception rate under different numbers of SFCs in a large-scale network in this embodiment of the invention; Figure 11 shows the algorithm performance evaluation under a large-scale network with different numbers of SFCs in the embodiment of the present invention, wherein: Figure 11 (a) is the average optimization objective function value of SFCs under different traffic requirements, Figure 11 (b) is the average end-to-end delay of SFCs under different traffic requirements, and Figure 11 (c) is the average resource usage cost of SFCs under different traffic requirements. Detailed Implementation
[0048] To illustrate the technical solutions disclosed in this invention in detail, the invention will be further described below with reference to the accompanying drawings and embodiments.
[0049] This invention provides a robust deployment method for SFC based on DRL in a dynamic computing network. It mainly introduces an uncertainty quantification model and a robust optimization mechanism to minimize communication latency and resource costs while maximizing service throughput in a computing network (CPN) with dynamically fluctuating resources, thereby ensuring the quality of service (QoS) under uncertain environments.
[0050] With the advent of the 6G era, emerging services such as holographic communication, digital twins, and virtual universes place stringent demands on network latency and computing power. Computing Network (CPN), as a new architecture integrating diverse computing and network infrastructure, deconstructs services into Service Function Chains (SFCs) composed of a series of Virtual Network Functions (VNFs) through Network Function Virtualization (NFV) technology. However, in cloud-native environments, computing resources are typically provided by multiple entities such as cloud service providers and edge service providers, exhibiting significant heterogeneity, distributed characteristics, and high dynamic uncertainty.
[0051] In actual operation, the available capacity of computing resource pools (such as cloud and edge nodes) often fluctuates in real time due to hardware upgrades, routine maintenance, operational status switching, and resource contention for background tasks (such as online payments and email services). This uncertainty in resource availability poses a significant challenge to the stable deployment of SFC: overload deployment may lead to performance degradation or even service interruption, while overly conservative deployment can increase communication latency between geographically distant nodes. Previous SFC deployment studies have mostly assumed that resource conditions are static or predictable, lacking robust mechanisms to cope with rapid fluctuations in infrastructure resources. Therefore, how to achieve robust deployment that balances stability and efficiency in a computing power network with dynamically changing resources is a critical problem that urgently needs to be solved.
[0052] To address the aforementioned challenges, this invention considers the stochastic fluctuations in the computing power of resource pools and constructs a robust resource optimization model by introducing a reference distribution and an uncertainty set. Unlike traditional methods that rely solely on simple statistics such as variance, this invention extracts an empirical reference distribution from historical trajectories and predicted data, and combines this with probabilistic distance metrics (such as KL divergence) to characterize the deviation in the actual resource distribution. This method can more accurately capture the inherent distribution characteristics of resource pools, providing data support for subsequent robust decision-making.
[0053] Building upon this foundation, this invention proposes a robust optimization framework under probabilistic constraints, reconstructing the complex SFC deployment problem into a robust optimization model. To address the solution challenges caused by random variables, this scheme develops a hybrid Newton-bisection algorithm, iteratively calculating a resource threshold that ensures reliable service provision. This threshold serves as a "safety boundary," guiding the system to make robust deployment decisions that resist resource fluctuations while meeting fault tolerance limits.
[0054] Considering the high computational cost and poor scalability of traditional optimization solvers (such as ILP) in large-scale network environments, this invention further proposes an SDPC architecture based on deep reinforcement learning. This scheme integrates the Approximate Policy Optimization (PPO) framework with a Convolutional Neural Network (CNN) to adaptively learn the optimal deployment strategy through a data-driven approach. Specifically, the architecture extracts the spatiotemporal features of network topology and resource states using CNN, generates the probability distribution of VNF deployment actions using an Actor-Critic network, and determines the routing scheme by combining a weighted path selection algorithm. This mechanism not only responds to environmental changes in real time but also exhibits superior scalability in large-scale dynamic scenarios.
[0055] This invention provides a robust SFC deployment method based on DRL in dynamic computing networks. The computing controller perceives the network topology and SFC requests, constructs a resource uncertainty model, and combines robust threshold calculation and reinforcement learning decision-making to collaboratively complete node allocation and link mapping in VNFs. Simulation experiments show that the proposed scheme can significantly reduce resource constraint violation rate in resource-constrained and highly volatile environments, and exhibits near-optimal performance in terms of service reception rate, end-to-end latency, and resource usage cost, effectively improving the resilience and efficiency of computing networks under uncertain conditions.
[0056] The method described in this invention enables robust deployment of Service Function Chains (SFCs) and intelligent decision-making regarding resource availability thresholds in a computing power network. This implementation mainly consists of a Computing Power Network (CPN) controller, a computing power resource scheduling and allocation module, a network resource scheduling and allocation module, an underlying network topology architecture, and diverse service requests. The CPN controller, as the decision-making center of the entire system, is responsible for real-time perception of the computing power network topology and resource distribution status. The computing power resource scheduling and allocation module manages geographically dispersed and heterogeneous computing power resource pools. The network resource scheduling and allocation module collaborates with the computing power resources to allocate optimal transmission paths to adjacent Virtual Network Functions (VNFs) within the SFC. The underlying network topology architecture consists of a series of routers, switches, and physical links with computing capabilities. Under the unified scheduling of the controller, this architecture provides underlying physical support for the SFC. Diverse service requests represent user-side service inputs; each request consists of a source node, a target node, a specific sequence of VNFs, and QoS constraints such as maximum tolerable latency. Users initiate requests to the controller through access points, and the system, through an automated decision-making process, matches the optimal computing power and network resources for them in a complex resource fluctuation environment.
[0057] The main implementation process of the method described in this invention is as follows: Figure 1 As shown, based on the above technical solution, further detailed description is provided in the embodiments, specifically including the following steps: (1) Establish a computing power network system architecture This invention constructs a system model consisting of a computing network controller, a distributed computing resource pool, and network links. The controller perceives the dynamic heterogeneous characteristics of the computing network in real time, acquiring physical parameters such as the computing power of the resource pool, link bandwidth capacity, and propagation delay. Simultaneously, it receives a set of SFC requests consisting of source nodes, target nodes, VNF sequences, and QoS constraints, and formalizes them into deployment tasks to be processed.
[0058] (2) Opportunity Constraint Optimization
[0059] This invention addresses the random fluctuations inherent in computing resource pools during actual operation by extracting a reference distribution from historical data and predictive information, rather than relying on simple moment statistics. It utilizes the KL divergence metric to quantify the deviation between the actual resource distribution and the reference distribution, and defines a set of distribution uncertainty accordingly. By introducing a fault-tolerance threshold to adjust the degree of conservatism, the original rigid resource capacity constraint is reconstructed into a probabilistic constraint model, thereby achieving a mathematical measure of resource fluctuation risk.
[0060] (3) Determine the robust SFC deployment decision threshold
[0061] This invention transforms the aforementioned probabilistic constraint optimization problem using the Lagrange multiplier method and the KKT optimality condition. A hybrid algorithm combining Newton's iteration method and the bisection method is employed to search for and determine a robust SFC deployment decision threshold that satisfies the fault tolerance limit. This threshold serves as a safety boundary for resource allocation, mapping the complex, underlying random uncertainties to a computationally achievable robust criterion at the upper level, thus solving the problem of the difficulty in directly solving for random variables.
[0062] (4) Define the main problem of the required robustness
[0063] After obtaining the robust decision threshold, this invention constructs a multi-objective optimization model with the comprehensive goals of minimizing end-to-end communication latency, reducing resource usage costs, and maximizing total service throughput. This model integrates VNF placement constraints, traffic conservation constraints, and resource capacity robustness constraints, ensuring that the generated deployment strategy meets stringent QoS requirements while resisting resource fluctuations in the worst-case scenario.
[0064] (5) SFC deployment algorithm based on DRL
[0065] To improve solution efficiency in large-scale dynamic network environments, this invention employs the SDPC algorithm based on the Approximate Policy Optimization (PPO) framework and Convolutional Neural Network (CNN). The CNN extracts high-dimensional network topology and resource features, the policy network generates VNF deployment actions, and a weighted Dijkstra algorithm completes the routing path mapping. During training, the algorithm uses a reward function to balance latency, cost, and stability, ultimately outputting a robust SFC deployment scheme with adaptive capabilities.
[0066] This embodiment considers a computing power controller and several distributed computing resource pools (denoted as...). Multiple routers (denoted as) ) and physical links (denoted as A computing network (CPN) composed of computing power networks. Figure 2 This demonstrates the business delivery process within the CPN.
[0067] Taking video streaming services as an example, data needs to be transmitted from router A to router F, traversing virtual network functions (VNFs) such as firewalls, proxy servers, and video transcoders in sequence. In a CPN environment, computing resource availability is highly dynamic—the shared resource pool must simultaneously meet the demands of newly arriving service requests as well as long-running backend services such as online transactions, payment systems, and email applications. Given that the failure of any single VNF instance can disrupt the entire service chain, the dynamic heterogeneous nature of CPN must be integrated into the SFC deployment process to ensure reliable service delivery even under resource fluctuations.
[0068] Figure 3 This paper demonstrates the overall workflow of the SDPC architecture, highlighting key components such as environment interaction, feature extraction, enforcer (policy) and critic (value) networks, action selection, path determination, reward evaluation, loss calculation, and parameter update. The system receives the environment state, including SFC requests and network topology①. Given the large number of nodes and their multidimensional resource attributes (such as CPU capacity, processing latency, and bandwidth), traditional state representations often struggle to capture inter-node dependencies and the global topology. Therefore, this system employs a convolutional neural network feature extractor to process the high-dimensional state tensor: after processing through two two-dimensional convolutional layers with ReLU activation functions, a compact low-dimensional feature vector is output via a fully connected layer②. This vector serves as input to both the policy network and the value network.
[0069] The policy network generates a probability distribution of possible deployment nodes in the VNF, and samples deployment actions from it after SoftMax normalization. Simultaneously, the value network estimates the value of the current state to calculate the advantage function ③ and ④ for policy optimization. After the action is executed, a weighted Dijkstra algorithm determines the routing path. The environment then provides immediate feedback through rewards and the next state ⑤. Using the PPO loss function, the parameters ⑥ of the convolutional neural network, policy network, and value network are updated through backpropagation. This iterative process progressively optimizes the deployment strategy.
[0070] In this embodiment, simulation experiments were conducted to verify the actual effect of the present invention. To better illustrate the effect of the present invention, comparisons were made using existing algorithms ILP_Greedy and NFVdeep. To ensure the repeatability and reliability of the results, simulations were performed on a desktop computer equipped with an 11th generation Intel® Core™ i5-11400 processor (running frequency 2.60GHz), 16GB of memory, and running the Windows 10 operating system. The experiments were conducted in a Python environment, targeting a small network topology with 6 nodes and 11 edges, and a USANet network with 24 nodes and 43 edges. The tools used included Gurobi 9.5, Stable-Baselines3 (version 1.5.0), and PyTorch (version 1.13).
[0071] Its effects mainly include the following points: (1) Convergence speed of algorithms in small networks Figure 4 illustrates the convergence behavior of the SDPC and NFVdeep algorithms in small-scale network environments. As training progresses, both training loss and value loss show a continuous decreasing trend, eventually reaching a stable convergence state. These findings confirm the convergence characteristics and training stability of the two algorithms in small-scale networks, thus verifying the effectiveness of their training process and the rationality of their underlying design.
[0072] (2) Average reception rate of different numbers of SFC requests in small networks
[0073] Figure 5 The average acceptance rate of different algorithms under varying traffic demands is shown, where acceptance rate is defined as the ratio of the number of successfully deployed SFCs to the total number of SFC requests. The results show that when the number of SFC requests reaches 40 and the traffic demand gradually increases from 4 to 8, all algorithms achieve a 100% acceptance rate. This indicates that under sufficient resource conditions, all algorithms can maximize the service request acceptance rate.
[0074] As traffic demand increases to 9, Figure 5This indicates that the Deterministic_ILP algorithm achieves the highest acceptance rate for SFC requests. However, this performance improvement comes at the cost of a significant increase in resource violation rate, such as... Figure 8 As shown, this indicates that the algorithm has limited robustness in managing resource constraints under high load conditions. In contrast, although SDPC and Robust_ILP have slightly lower receive rates, both maintain zero violation rates across all traffic levels, highlighting their superior stability and adaptability when dealing with stringent resource constraints. Notably, SDPC also outperforms NFVdeep in terms of receive rate, demonstrating a better balance between service efficiency and resource feasibility in resource-constrained network environments.
[0075] (3) Key optimization indicators
[0076] Figure 6 illustrates the performance of the evaluated algorithms on several key optimization metrics, including the overall objective value, end-to-end communication latency, and resource utilization cost. Figure 6(a) depicts the average objective value under different SFC traffic demands. It is evident that the objective value increases accordingly with increasing SFC traffic demand. This phenomenon stems from the positive correlation between traffic demand and communication latency and resource consumption; higher traffic leads to longer latency and higher costs. Comparative analysis further shows that Deterministic_ILP and Robust_ILP algorithms perform consistently well and outperform other methods. The SDPC algorithm follows closely behind, exhibiting near-optimal performance, while NFVdeep performs relatively poorly. Furthermore, Figures 6(b) and 6(c) further demonstrate that SDPC achieves near-optimal results in both communication latency and resource cost, highlighting its superior performance in overall optimization.
[0077] (4) Changes in SFC traffic demand in small networks
[0078] Figure 7 and Figure 8 The average resource utilization and resource constraint violation rate are shown separately for small networks as SFC traffic demand changes. Figure 7 As shown, resource utilization increases significantly with increasing SFC traffic demand, reflecting a corresponding increase in resource consumption and thus improving overall resource utilization efficiency. Furthermore, the results indicate that the Deterministic_ILP and Robust_ILP algorithms achieve comparable resource utilization, effectively utilizing available network resources. Simultaneously, the figure shows that SDPC and NFVdeep exhibit similar resource utilization levels.
[0079] Figure 8The study demonstrates the default rates of resource constraints under varying traffic demands. The deterministic integer programming algorithm maintains a low default rate under light traffic conditions, but the default rate increases significantly with increasing traffic. Notably, the proposed robust optimization algorithm consistently and effectively enforces resource constraints, highlighting its superior ability to maintain resource stability and ensure service quality.
[0080] (5) Algorithm convergence speed in large-scale network environments
[0081] To evaluate the scalability of the SDPC algorithm, this example compares it with the ILP_Greedy and NFVdeep algorithms in a large-scale network environment. Figure 9 shows the training loss and value loss convergence curves of the two algorithms, SDPC and NFVdeep. The results show that both convergence speeds are comparable, and the number of training iterations required is similar. Furthermore, both methods can stably converge to a steady state as training progresses. These findings confirm that SDPC and NFVdeep possess training stability and convergence characteristics, thus validating their effectiveness and applicability in large-scale network environments.
[0082] (6) Average reception rate of different numbers of SFC requests in large-scale networks
[0083] Figure 10 The average acceptance rate for different numbers of SFC requests in large-scale networks is shown. The results demonstrate that the SDPC algorithm consistently achieves the highest acceptance rate. Although the overall acceptance rate decreases with increasing SFC request volume, the absolute number of successfully accepted SFCs continues to rise, proving that the algorithm can effectively manage the increasing volume of service requests under high load conditions. Figure 11 shows key performance evaluation metrics under different SFC request volumes. By comparing Figures 11(a), 11(b), and 11(c), it can be found that the SDPC algorithm outperforms competing algorithms in terms of optimization objective, end-to-end communication latency, and resource utilization cost, highlighting its superior performance and robustness in large-scale dynamic network environments. The ILP_Greedy algorithm, due to its sequential non-global deployment strategy (which cannot complete the entire SFC allocation in one step), may converge to a local optimum. The NFVdeep algorithm lacks a comprehensive understanding of the underlying network topology and service request characteristics, resulting in limited decision-making efficiency.
[0084] In summary, the robust optimization scheme significantly reduces the resource constraint violation rate while achieving performance comparable to deterministic SFC deployment methods, thus ensuring enhanced robustness of the system in uncertain environments. Furthermore, the proposed SDPC algorithm achieves near-optimal performance in small-scale networks and surpasses state-of-the-art methods in large-scale network scenarios.
Claims
1. A robust SFC deployment method based on DRL in a dynamic computing power network, characterized in that, The method includes the following steps: (1) Construct a computing power network and its corresponding system model, which includes a computing power controller, a distributed computing resource pool and network links. The computing power controller senses the dynamic heterogeneous characteristics of the computing power network in real time, obtains the computing power, link bandwidth capacity and propagation delay physical parameters of the resource pool, and receives an SFC request set composed of source node, target node, VNF sequence and QoS constraints, and converts it into a deployment task to be processed. (2) In view of the random fluctuation characteristics of the distributed computing resource pool, a reference distribution is extracted from historical data and prediction information. The uncertainty set between the actual resource distribution and the reference distribution is quantified by KL divergence. A fault tolerance threshold is introduced to adjust the degree of conservatism. The hard resource capacity constraint is reconstructed into a probability constraint model to realize the mathematical measurement of resource fluctuation risk. (3) The probabilistic constraint optimization problem is transformed by the Lagrange multiplier method and KKT optimality conditions. A hybrid algorithm combining Newton's iteration method and the bisection method is used to search for and determine the robust SFC deployment decision threshold that satisfies the fault tolerance upper limit. The robust SFC deployment decision threshold serves as a safety boundary for resource allocation, mapping the underlying random and uncertain constraints to the upper-layer computable robust criteria. (4) Based on the SFC deployment decision threshold, construct a multi-objective optimization model with the comprehensive objectives of minimizing end-to-end communication latency, reducing resource usage costs, and maximizing total service throughput; The objective optimization model follows VNF placement constraints, traffic conservation constraints, and resource capacity robustness constraints to ensure that the generated deployment strategy can withstand resource fluctuations in the worst case while meeting stringent QoS requirements. (5) The SDPC algorithm based on the approximate policy optimization framework and convolutional neural network is called to solve the problem. The high-dimensional network topology and resource features are extracted by the convolutional neural network. The policy network generates VNF deployment actions and combines the weighted Dijkstra algorithm to complete the routing path mapping. Finally, a robust SFC deployment scheme with adaptive capabilities is output.
2. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 1, characterized in that, The system model includes a computing power network model and an SFC request model: In the aforementioned computing power network model, the set of computing resource pools is denoted as... A single computing resource pool is represented as , its in Time-based computing capability is represented as physical link The propagation delay is expressed as The starting point of the physical link is a distributed computing resource pool or router, denoted as... The destination is recorded as The cost per unit of computing resource in a distributed computing resource pool is defined as... The cost of a unit bandwidth resource on a physical link is defined as ; In the SFC request model described above, using Indicates SFC request The set of, where and Indicates the source node and the target node; SFC request The traffic demand is denoted as , Indicates SFC request Maximum tolerable delay, Indicates SFC request The topology, ,in Indicates SFC request VNF set, Indicates SFC request A virtual link set, where SFC is composed of instances of different VNFs; Define binary variables Indicates virtual network function Whether deployed in the computing resource pool Define binary variables above. Indicates virtual link Is it mapped to a physical link? Using binary variables Indicates SFC request Whether the deployment was successful; VNF The computing resources required to process one unit of flow are denoted as . This virtual network function The delay required to process one unit of traffic is denoted as . .
3. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 1 or 2, characterized in that, The constraints of the computing network mentioned in step (1) include: Computational resource constraints: Wherein, it is assumed that there exists In VNF type, its set is denoted as Define variables VNF type For the duration of operation; Bandwidth constraints: Service quality constraints: In the formula, Request for SFC At any moment The end-to-end delay.
4. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 1, characterized in that, The uncertainty set mentioned in step (2) is defined as follows: in, For computing resource pool The actual distribution of computing power It is a reference distribution. This represents the distance constraint given by the KL divergence calculation. Indicates distribution Expectations; The probability constraint model is as follows: in, This represents the fault tolerance threshold of the computing power network (CPN), quantifying the maximum permissible probability that available computing power cannot meet the required level.
5. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 1, characterized in that, The robust SFC deployment decision threshold mentioned in step (3) is expressed as: Its determination process includes: The probabilistic constraint optimization problem is transformed into a convex optimization problem, and the worst-case resource violation probability is derived using the Lagrange multiplier method and KKT optimality conditions. Based on the property that the probability of resource violation in the worst case monotonically increases with the robust SFC allocation decision, a bisection method is used in the interval... Search within to satisfy The solution, where To ensure The empirical coefficient, This is the resource violation probability mapping function in the worst-case scenario.
6. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 1, characterized in that, The objective function of the multi-objective optimization model in step (4) is: in, yes Matrix vectors, It means Matrix vectors, and It is the SFC deployment decision matrix. , and These are the weighting coefficients for latency, cost, and throughput, respectively. Request for SFC At any moment The cost of using resources.
7. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 6, characterized in that, The SFC request At any moment resource usage costs The calculation method is as follows: 。 8. The robust SFC deployment method based on DRL under dynamic computing power network according to claim 1, characterized in that, The loss function of the SDPC algorithm in step (5) is constructed as follows: in, For policy network parameters, Indicates the current strategy Compared with previous strategies In state Next action The probability ratio, Represents the dominance function. To truncate the threshold, , These are the weighting coefficients. It is the value of state estimation. The corresponding real-state value is obtained through time-difference learning. For policy entropy, It is the probability distribution of the policy network output.