A method for reliable deployment of SFC based on multi-agent proximal policy optimization
By employing a multi-agent proximal policy optimization method, combined with Markov decision processes and noise value functions, the problems of training stability and scalability elasticity in SFC deployment are solved, achieving reasonable resource allocation and optimization of reliability and cost, thereby improving the stability and efficiency of network services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2023-02-20
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to effectively improve the training stability and scalability of algorithms in multi-agent systems during SFC deployment, and fail to effectively combine reliability and deployment cost for comprehensive optimization, resulting in vulnerability of end-to-end network services.
We adopt a multi-agent proximal policy optimization approach. By designing availability schemes, establishing utility functions and Markov decision process models, and combining the KL divergence method and policy scaling, we achieve reasonable resource allocation and minimize deployment costs. We employ a framework of centralized training and step-by-step execution and use a noise value function to reduce the impact of overfitting.
It optimizes reliability and deployment cost under resource constraints, reduces end-to-end latency, improves the balance of resource allocation, and maintains good scalability as the number of agents increases.
Smart Images

Figure CN116156565B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology and relates to a reliable deployment method for SFC based on multi-agent near-end policy optimization. Background Technology
[0002] 5G network software-defined networking is considered a revolutionary cluster of technologies that encourages agility, programmability, and resilience by promoting software-oriented architecture. The most prominent candidate technologies for this software-defined paradigm are software-defined networking and network function virtualization, where physical network functions are replaced by virtual network functions. These functions are performed by industry-standard physical machines (e.g., commodity servers, switches / storage nodes, etc.). These virtual functions are linked in a strict processing order to form a service function chain to provide the diverse network services required by users and emerging applications.
[0003] NFV (Network Functions Virtualization) has greatly improved many aspects of future communication networks, such as automating network operation and providing resilient services. Nevertheless, end-to-end network service vulnerabilities still exist because many failures can occur. Therefore, meeting the reliability requirements of user services is crucial for any network service provider. Mobile users typically request not only specific VNF (Virtual Network Function) services but also a certain level of reliability. Network reliability is defined as the network's ability to provide stable service to ensure a reliable level of operation.
[0004] Heuristic methods rely on well-defined manual rules, thus machine learning-based approaches have garnered significant attention for addressing the reliable deployment of Service Chain Components (SFCs). Current research on SFC deployment typically focuses on single-objective optimization for reliability, rarely considering other factors comprehensively. Furthermore, while numerous studies have addressed SFC deployment through reinforcement learning, few have extended the training scenario to multi-agent systems, and few have examined improving the algorithm's training stability and scalability in response to increasing business demands. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a reliable SFC deployment method based on multi-agent near-end policy optimization, which optimizes reliability and deployment cost under the constraints of underlying resources, effectively reduces end-to-end latency, improves the balance of resource allocation, and has good scalability when the number of agents increases.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A reliable SFC deployment method based on multi-agent proximal policy optimization specifically includes the following steps:
[0008] S1: In the scenario of network function virtualization, design an availability scheme based on function distribution, establish a utility function based on availability probability, and propose reliability penalty schemes for load balancing and latency tolerance difference respectively.
[0009] S2: Under the condition of satisfying service latency constraints, establish the SFC reliable deployment optimization problem of maximizing joint availability and minimizing cost, and transform the problem into a Markov decision process model;
[0010] S3: The KL divergence method is used to ensure that optimization is completed in the confidence region, and the confidence region constraint is further implemented through policy proportional pruning;
[0011] S4: In a multi-agent system, the overall framework is based on centralized training and step-by-step execution. Each decision-maker adopts a proximal policy optimization algorithm. Random noise implicitly affects the advantage function by interfering with the centralized value network, so as to reduce the overfitting effect caused by the sampling advantage value bias.
[0012] Furthermore, in step S1, the network function virtualization (NFV) scenario includes a physical layer, a virtual layer, a control layer, and an application layer. The physical layer is a general-purpose underlying network holding basic resources (the physical layer consists of servers and links, acting as the bottom layer of the proposed architecture; once selected as an embedded baseboard node or link for a virtual network request, it will be responsible for processing and forwarding user data streams). The virtual layer categorizes and groups user needs into services, constructing virtual networks from these needs. The control layer performs comprehensive analysis and scheduling, completes decisions at each stage, and performs real-time monitoring. The application layer is primarily responsible for statistically analyzing current service types and needs, and transmitting stored information to the virtualization layer for analysis and operation.
[0013] Furthermore, in step S1, function distribution refers to adding VNF replicas after the deployment of VNF (Virtual Network Function), thereby reducing the risk of network service interruption. Each VNF replica will consume the same computing resources as the primary VNF. Considering the network reliability requirements and the backup habits of users in reality, only one replica needs to be set when it is necessary. At this time, VNF availability means that at least one instance of its primary VNF and replica VNF is available.
[0014] Furthermore, in step S1, the end-to-end delay of SFC includes processing delay and transmission delay, denoted by D. i Let represent the end-to-end delay of the i-th SFC, then its representation in time slot t is:
[0015]
[0016] The total processing delay P for the i-th SFC i This is related to the VNF mapping situation and is represented in time slot t as follows:
[0017]
[0018] in, This indicates that j VNFs in the i-th SFC are deployed to server v, and F represents the set of SFCs in the network. Let N represent the set of VNFs formed on the i-th SFC. s ={n1,n2,…n m} represents a set of m servers; Let m represent the single-node processing latency. i Let β represent the data packet size and β represent the processing rate coefficient. In time slot t, it is represented as:
[0019]
[0020] Where, ω i (t) represents the number of data packets actually arriving at the i-th SFC, following the parameter λ. i The Poisson distribution; This indicates the proportion of CPU resources allocated to server v. This represents the resource capacity held by the v-th server;
[0021] For the total communication delay T of the i-th SFC link i This is also related to the VNF mapping situation, and it is represented in time slot t as:
[0022]
[0023] Where jk represents the link connecting the j-th and k-th adjacent VNFs on the i-th SFC, E i Represents the set of links on the i-th SFC; Let jk represent a Boolean variable. When the link jk of the i-th SFC is mapped to the underlying link uv, then... L represents the set of links between nodes, and uv represents the connection n. u and n v The underlying link; This represents the corresponding communication delay, which is related to the amount of data to be transmitted, and can be expressed in time slot t as:
[0024]
[0025] in, This indicates the bandwidth resource requirement.
[0026] Furthermore, in step S1, the reliability penalty includes two parts: establishing an SLA protocol penalty based on node load rate; and assuming... Let be the CPU remaining rate of the v-th node, then its calculation formula is:
[0027]
[0028] Regarding load penalties, α c Indicates the resource overload warning value, ε c This represents the penalty imposed on the portion of CPU resource availability that falls below a warning threshold. The greater the difference from the warning threshold, the higher the penalty. This is the SLA penalty for server v violating load in the network. In time slot t, it is represented as:
[0029]
[0030] Regarding latency penalties, latency warning values τ are set for different types of SFCs. i , End-to-end delay exceeds τ i The portion will be subject to SLA penalties, denoted by a unit penalty coefficient of ε. d Penalty for SFC violating the delay protocol in Article i In time slot t, it is represented as:
[0031]
[0032] The usability score is measured based on the usability probability. but
[0033] The availability calculation formula for the j-th VNF on the i-th SFC is:
[0034]
[0035] in, This represents the set of primary replicas of the j-th VNF on the i-th SFC placed on server v.
[0036] Furthermore, in step S2, the total deployment cost Z in the network sum It can be expressed as the sum of three parts, that is
[0037] Z sum (t)=Z1(t)+Z2(t)+Z3(t)
[0038] The cost expressions for each part are as follows:
[0039]
[0040]
[0041]
[0042] For the i-th SFC, This represents the operating cost of the j-th main VNF on server v. This represents the cost of bandwidth used by jk on the physical link UV. A boolean variable indicating whether VNFj sets the VNF on server v, λ3 represents the unit cost of resources occupied by the replica, λ4 represents the unit cost of using the server's scheduling controller, and ω v This represents the unit cost of operating the scheduling controller;
[0043] Furthermore, in step S2, a joint optimization objective for reliable SFC deployment is established, and the utility function designed after considering various aspects is as follows:
[0044] U(t)=σ1S(t)-σ2E(t)-σ3Z sum (t)
[0045] Where S(t) represents the average network availability, E(t) represents the sum of load and latency penalties, and the coefficient σ q ,q=1,2,3 represent the corresponding weight coefficients of each item; the above formula needs to be performed under the constraints, firstly for the basic mapping related to VNF, link and replica, then for the capacity constraints including both computing resources and link resources, and then for the availability and latency requirements proposed for reliability.
[0046] Furthermore, in step S2, the optimization problem of reliable SFC deployment is transformed into an MDP model, represented by a quadruple M = <S, A, P, R>.
[0047] The state space S is defined as the mapping state information of SFC, the running state information of the node scheduling controller, and the node CPU resource remaining rate information. Therefore, for time slot t, s t ∈S represents the sum of three parts s t ={K(t),ω(t),η c (t)}, where K(t)=[K i (t)], K i (t) represents the mapping state information of the i-th SFC; ω(t) = [ω v ],
[0048] The action space A is defined as the mapping of each chain's main VNF, the placement of replica VNFs, and CPU allocation. Therefore, for time slot t, a t ∈A is represented as a t ={δ(t),Φ(t),X(t)}, where,
[0049] For the state transition probability p(s) t+1 |s t ,a t ) is defined as being in state s t Next, execute action a t After that, the state information will be transferred to the new time slot. t+1 The transition probability distribution is P: S×A×S→R.
[0050] Since the optimization objective is to maximize network availability and minimize deployment cost, and in order to satisfy the constraints, the reward function is defined as R(t) = kU(t), where k is a coefficient greater than 0.
[0051] Furthermore, step S3 specifically includes: introducing a KL constraint term to limit the KL divergence difference between the old and new policy functions. At this point, the objective function is maximized under the constraint of limiting the gradient update magnitude, expressed as:
[0052]
[0053]
[0054] in, This represents the average value calculated over the training trajectory, π. θ (a t |s t ) indicates a new strategy. Indicates the original strategy, δ θ Indicates the KL divergence limit value;
[0055] Furthermore, we transform it into an unconstrained optimization form, and combine it with a policy scaling method. At this point, the maximization objective is rewritten as a pruned objective function, i.e.
[0056]
[0057] Where, r t (θ) represents the ratio of the old and new strategies, and clip(·) represents the ratio of r to r. t A clipping function that limits the size of (θ). Used to control this defined range.
[0058] Furthermore, step S4 specifically includes the following steps:
[0059] S41: Classify and chain the network services requested by users;
[0060] S42: Reset the environment for SFC deployment and initialize the parameters of each actor and critic network;
[0061] S43: The agent selects actions in a local area, places VNFs and VNF replicas, allocates node computing resources, and obtains decision rewards and new state information for SFC deployment.
[0062] S44: Repeat the decision steps and store the trajectory until the maximum number of steps in the iteration is reached;
[0063] S45: Randomly select samples and add noise;
[0064] S46: Calculate the noise value function using a reduced version of the generalized dominance estimation method to obtain the dominance function;
[0065] S47: During training, the objective function and joint loss function are calculated, and then the critic network and actor network are updated using the Adam method;
[0066] S48: Repeat steps S42 to S47 until all decision-makers' models converge or the round deadline expires.
[0067] The beneficial effects of this invention are as follows: Under the resource constraints of physical servers and links, this invention rationally arranges resources during the deployment process, thereby jointly optimizing reliability and deployment cost. It adopts a local near-end strategy optimization and a multi-agent learning framework with centralized training and distributed execution at the upper layer. It combines the noise value function and the generalized advantage function estimation method of training trajectory to maximize the improvement of agent training effect.
[0068] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0069] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0070] Figure 1 This is a flowchart of the SFC reliable deployment method based on multi-agent proximal strategy optimization according to the present invention;
[0071] Figure 2 This is a system architecture diagram for enabling network function virtualization in this invention;
[0072] Figure 3 This invention provides a reliable SFC serial-parallel deployment scheme.
[0073] Figure 4 This invention provides a service function chain deployment framework based on multi-agent reinforcement learning.
[0074] Figure 5 This is a diagram of the network structure for multi-agent proximal strategy optimization in this invention. Detailed Implementation
[0075] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0076] Please see Figures 1-5 This invention provides a reliable SFC deployment method based on multi-agent proximal policy optimization, see [link to relevant documentation]. Figure 1 The method specifically includes the following steps:
[0077] S1: Design an availability scheme based on function distribution, establish a utility function based on availability probability, and propose reliability penalty schemes for load balancing and latency tolerance difference, respectively.
[0078] In step S1, function distribution refers to adding VNF replicas after VNF deployment, thereby reducing the risk of network service interruption. Each VNF replica will consume the same computing resources as the primary VNF. Considering network reliability requirements and users' backup habits in reality, only one replica needs to be set when it is necessary. At this time, VNF availability means that at least one instance of its primary VNF and replica VNF is available.
[0079] Uneven load can lead to network congestion and instability, while excessive business processing latency can cause network instability, resulting in decreased reliability.
[0080] S2: Under the constraint of service latency, establish a stochastic optimization problem that combines maximizing availability and minimizing cost, and transform this problem into a Markov decision process model.
[0081] In step S2, the stochastic optimization problem requires designing a utility function for comprehensive evaluation. The goal is to minimize this utility function under all constraints, representing how to deploy each primary VNF, set up replicas, and allocate computing resources, so as to minimize the deployment cost of SFC while maximizing the reliability of network services.
[0082] S3: The KL divergence method is used to ensure that the optimization is completed in the confidence region, and further confidence region constraints are achieved through policy proportional pruning.
[0083] In step S3, the trust region method is transformed into an unconstrained optimization form, and the trust region constraint is implemented through policy scaling. Using this method for policy optimization results in more stable training performance compared to the traditional stochastic gradient ascent method.
[0084] S4: In a multi-agent system, the overall framework is based on centralized training and step-by-step execution. Each decision-maker adopts a proximal policy optimization algorithm. Random noise implicitly affects the advantage function by interfering with the centralized value network, so as to reduce the overfitting effect caused by the sampling advantage value bias.
[0085] In step S4, for the centralized training and step-by-step execution framework, each agent has a local actor and critic network. The actor network solves the policy only through local observations, while the critic network receives the actions of each agent and then calculates the centralized value function.
[0086] In a multi-agent system, different agents are represented as users with different business needs. Each agent adopts a baseline method of proximal policy optimization and continuously interacts with the environment to learn individual policies. At this time, the decision-making process is extended to a distributed partially observable Markov decision process.
[0087] S4 specifically includes the following steps:
[0088] S41: Classify and chain the network services requested by users;
[0089] S42: Reset the environment for SFC deployment and initialize the parameters of each actor and critic network;
[0090] S43: The agent selects actions in a local area, places VNFs and replicas, allocates node computing resources, and obtains decision rewards and new state information for SFC deployment.
[0091] S44: Repeat the decision steps and store the trajectory until the maximum number of steps in the iteration is reached;
[0092] S45: Randomly select samples and add noise;
[0093] S46: Calculate the noise value function using a reduced version of the generalized dominance estimation method to obtain the dominance function;
[0094] S47: During training, the objective function and joint loss function are calculated, and then the critic network and actor network are updated using Adam.
[0095] S48: Repeat steps S42 to S47 until all decision-makers' models converge or the round deadline expires.
[0096] See Figure 2 Network Functions Virtualization (VNF) scenarios comprise four components: the physical layer, the control layer, the virtual layer, and the application layer. The physical layer includes the underlying server nodes and links, serving as the foundation of the proposed architecture and providing the basic resources for VNF instantiation (once selected as an embedded baseboard node or link for virtual network requests, it is responsible for processing and forwarding user data streams). The control layer primarily performs real-time monitoring of network information, load analysis for network decisions, and execution of resource allocation strategies. The virtual layer, relative to the physical layer, categorizes and groups user needs, constructing the needs into virtual networks. The application layer is responsible for statistics and storage of various tenant applications.
[0097] A physical network, consisting of a large number of nodes and links, is modeled as an undirected graph G. s =(N s ,L). N s ={n1,n2,…n m Let} be a collection of m servers that provide the computing resources needed to process network functions, and each underlying server can instantiate multiple network functions. This represents the resource capacity held by the v-th server. L = {l uv |n u ,n v ∈N s} represents the set of links between nodes, and uv represents the number of links n. u and n v The underlying link, whose maximum available bandwidth resources are represented as Set up a scheduling controller for each node to schedule availability replicas, and define a boolean variable ω. v ={0,1}, when the scheduling controller of the v-th node is running, ω v =1, which means that a VNF replica exists on the server it is on.
[0098] The virtual network is modeled as a directed graph G. v = (V, P). The set of SFCs in the network is denoted as F, and the i-th SFC is represented as a directed graph. V i Let P be the VNF set on the i-th SFC. i Let represent the set of virtual links on the i-th SFC. For the j-th VNF on the i-th SFC, This represents the amount of computing resources allocated to physical node v. jk represents the link connecting the j-th and k-th adjacent VNFs on the i-th SFC. This indicates the amount of bandwidth resources allocated to it by the underlying link uv.
[0099] The application layer provides a solution for building SFCs in the virtual layer, and various applications use SFCs as carriers to provide various services to users.
[0100] See Figure 3 , Figure 3 The SFC serial-parallel reliable deployment scheme of the present invention does not adopt the method of backup of adjacent nodes to improve reliability. Instead, it considers whether to add VNF replicas at the deployed node. When in use, the processing can be completed by using the primary VNF or any VNF in the replica pool. This parallel method increases the availability probability of VNF, thereby reducing the risk of network request failure.
[0101] If a node is configured with replicas, usually only one is needed. This is because the improvement in availability gradually decreases as the number of replicas increases, and this also more closely reflects the actual operation of real users. The part without replicas is a local serial connection, while the part with replicas is a local parallel connection. The overall serial-parallel system can effectively improve service reliability.
[0102] See Figure 4 , Figure 4 This is a service function chain deployment framework based on multi-agent reinforcement learning. Users with various business needs are treated as different agents, each assigned a number as needed. Each agent possesses local observation information, makes decisions to obtain rewards, and then jumps to the next new state value based on environmental state information. Through continuous interaction with the environment, each agent learns the optimal deployment strategy. The agents cooperate to serve arriving requests. Each agent has access to all resources in the environment and selects certain network resources to meet its deployment needs. Their common goal is to obtain the maximum cumulative shared reward.
[0103] By employing a multi-agent learning approach, the optimal placement and resource scheduling scheme is designed to meet various design requirements. This deployment framework features autonomy, coordination, and distribution, and the agents can also communicate and integrate with each other.
[0104] See Figure 5 , Figure 5The diagram shows the network structure of the multi-agent proximal policy optimization system in this invention. Traditional reinforcement learning is difficult to adapt to the scenario of multi-agent systems because a single agent will face the problem of environmental instability when performing independent distributed learning, making it difficult to train the best policy. However, if centralized reinforcement learning is used, in addition to the large action space, this centralized approach will lead to a large signaling overhead for interaction. The best way to solve the above problems is to adopt a method based on centralized training and distributed execution.
[0105] In a multi-agent scenario, the policy ratio of agent a is expressed as:
[0106]
[0107] The objective function to be maximized is expressed as:
[0108]
[0109] Where B represents the batch size, S represents the policy entropy, and σ represents the entropy parameter. Let τ represent the training trajectory. If we denote the discounted future rewards (rewards-to-go), then the loss function L(φ) to be minimized can be expressed as:
[0110]
[0111] To address the overfitting problem of the strategy, we consider adding noise, i.e.
[0112]
[0113] Among them, a noise As a weight for the noise values, another implicit method is to change the advantage function through the value function. Let the sampled Gaussian noise vector be represented as... The value function with noise can then be expressed as:
[0114]
[0115] The deployment method proposed in this invention is based on a framework of centralized training and step-by-step execution. Each agent has a local actor and commentator network. The actor network only needs local observations to solve the policy, while the commentator network needs to input the actions of all agents to obtain a centralized value function.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A reliable deployment method for SFC based on multi-agent proximal policy optimization, characterized in that, The method specifically includes the following steps: S1: In the scenario of network function virtualization, design an availability scheme based on function distribution, establish a utility function based on availability probability, and propose reliability penalty schemes for load balancing and latency tolerance difference respectively. In step S1, the network function virtualization scenario includes a physical layer, a virtual layer, a control layer, and an application layer. The physical layer is a general-purpose underlying network that holds basic resources. The virtual layer classifies and groups user needs into virtual networks. The control layer performs comprehensive analysis and scheduling, makes decisions at each stage, and monitors in real time. The application layer is responsible for statistically analyzing current service types and needs and transmitting the stored information to the virtualization layer for analysis and operation. Feature distribution refers to adding a VNF replica after VNF deployment; SFC's end-to-end latency includes processing latency and transmission latency, using Indicates the first The end-to-end delay of an SFC is then in The time slot is represented as: For the Total processing latency of each SFC This is related to the VNF mapping situation. The time slot is represented as: in, Indicates the first In the SFC A VNF is deployed to the server. superior, F This represents the set of SFCs in the network. j Indicates the first A VNF, Indicates the first The set of VNFs formed on the SFC, for A collection of servers; Indicates the single-node processing latency; let... Indicates the data packet size. Representing the processing rate coefficient, then exist The time slot is represented as: in, Indicates the first The number of data packets actually arriving at each SFC. Indicates server The proportion of CPU resources allocated to it. Indicates the first The resource capacity held by each server; For the Total communication delay of SFC link This is also related to the VNF mapping situation, which is in The time slot is represented as: in, jk Indicates the connection of the first Adjacent SFCs The and the first A link in a VNF, Indicates SFC number The set of links on an SFC; Represents a Boolean variable. L This represents the set of links between nodes. UV Indicates connection and The underlying link; This indicates the corresponding communication latency, which is related to the amount of data to be transmitted. The time slot is represented as: in, This indicates the bandwidth resource requirement; The reliability penalty consists of two parts: establishing a penalty based on the node load rate SLA protocol; and assuming... It is the first The CPU remaining rate of each node is calculated using the following formula: Regarding load penalties, This indicates the resource overload warning value. This represents the penalty imposed on the portion of CPU resource availability below a warning threshold. The greater the difference from the warning threshold, the higher the penalty. (This applies to servers in a network.) SLA penalty for violating load section exist The time slot is represented as: Regarding latency penalties, latency warning values are set for different types of SFCs. End-to-end latency exceeds The portion will be subject to SLA penalties, assuming a unit penalty coefficient of 1. , No. Penalties for SFC Violation of Delay Protocol exist The time slot is represented as: The usability score is measured based on the usability probability. ,but No. Article 1 of SFC A VNF is placed on the server. The availability calculation formula is as follows: in, Indicates the first Article 1 of SFC A VNF is placed on the server. The main copy collection on; S2: Under the condition of satisfying service latency constraints, establish the SFC reliable deployment optimization problem of maximizing joint availability and minimizing cost, and transform the problem into a Markov decision process model; In step S2, the total deployment cost in the network It can be expressed as the sum of three parts, that is The cost expressions for each part are as follows: in, Indicates the first The main VNF on the server Operating costs on express In physical link The cost of using bandwidth, VNF Is it on the server? Set a boolean variable in the VNF. This represents the unit cost of resources used by a copy. This represents the unit cost of using the server's scheduling controller. This represents the unit cost of operating the scheduling controller; The joint optimization objective for reliable SFC deployment is established, and the utility function designed after considering various factors is as follows: Among them, coefficient This represents the weight coefficients corresponding to each item; Indicates the average availability of the network. This represents the sum of load and latency penalties. The above formula must be performed under certain constraints. First, there are the basic mappings related to VNFs, links, and replicas. Then there are the capacity constraints, including both computing resources and link resources. Finally, there are the availability and latency requirements for reliability. The optimization problem of reliable SFC deployment is transformed into an MDP model, using a quadruple. To indicate; For the state space The information is defined as the mapping status information of SFC, the running status information of the node scheduling controller, and the node CPU resource remaining rate information. Therefore, it is for time slots. , Represented as the sum of three parts ,in, , Indicates the first The mapping status information of each SFC; ; ; For action space Defined as the mapping of each chain's main VNF, the placement of replica VNFs, and CPU allocation, therefore targeting time slots. , Represented as ,in, , , ; For state transition probability Defined in state Next, carry out the action. After that, the state information will be transferred to the new time slot. The transition probability distribution is ; Since the optimization objective is to maximize network availability and minimize deployment cost, and in order to satisfy the constraints, the reward function is defined as follows: , where k is a coefficient greater than 0; S3: The KL divergence method is used to ensure optimization is completed within the confidence region, and further confidence region constraints are achieved through policy scaling. Specifically, this includes introducing a KL constraint term to limit the KL divergence difference between the old and new policy functions. At this point, the objective function is maximized under the constraint of limiting the gradient update magnitude, expressed as: in, This indicates that the average value is calculated over the training trajectory. Indicating a new strategy, Indicates the original strategy, Indicates the KL divergence limit value; Furthermore, we transform it into an unconstrained optimization form, and combine it with a policy scaling method. At this point, the maximization objective is rewritten as a pruned objective function, i.e. in, This represents the ratio of the old to the new strategy. Indicates will A clipping function that limits the size of the clipping function. Used to control this defined range; S4: In a multi-agent system, the overall framework is based on centralized training and step-by-step execution. Each decision-maker employs a proximal policy optimization algorithm. Random noise implicitly influences the advantage function by interfering with the centralized value network, thereby reducing the overfitting effect caused by sampling advantage bias. Specifically, the following steps are included: S41: Classify and chain the network services requested by users; S42: Reset the environment for SFC deployment and initialize the parameters of each actor and critic network; S43: The agent selects actions in a local area, places VNFs and VNF replicas, allocates node computing resources, and obtains decision rewards and new state information for SFC deployment. S44: Repeat the decision steps and store the trajectory until the maximum number of steps in the iteration is reached; S45: Randomly select samples and add noise; S46: Calculate the noise value function using a reduced version of the generalized dominance estimation method to obtain the dominance function; S47: During training, the objective function and joint loss function are calculated, and then the critic network and actor network are updated using the Adam method; S48: Repeat steps S42 to S47 until all decision-makers' models converge or the round deadline expires.