Service function chaining route optimization method for industrial metaverse network

By acquiring network status and SFC request characteristics in real time within the industrial metaverse network, and utilizing the Actor-Critic network and multi-level attention mechanism, the resource contention and dependency issues among concurrent SFC requests are resolved, improving the collaborative optimization quality of SFC routing and ensuring the stability of long-term experience quality and efficient network operation.

CN122372490APending Publication Date: 2026-07-10NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2026-05-28
Publication Date
2026-07-10

Smart Images

  • Figure CN122372490A_ABST
    Figure CN122372490A_ABST
Patent Text Reader

Abstract

This invention provides a service function chain routing optimization method for industrial metaverse networks. It acquires real-time basic network state parameters of the underlying physical infrastructure and service attribute characteristics of concurrent collaborative service function chain requests (SFC requests). For each SFC request, it extracts a dimensionality-reduced candidate path set and routing attribute feature vector. It then extracts the relative quality features of the candidate paths, resource contention features among concurrent requests, and hidden state features containing time correlations. This yields the final conflict-free routing path. The method also observes the actual instantaneous QoS of the underlying layer after the SFC service traffic transmission is completed online. A continuous mapping model is constructed, enabling continuous online optimization of SFC routing decisions. This method ensures long-term QoE stability of dynamic networks, improves the collaborative routing quality of SFCs, optimizes instantaneous transmission quality while significantly suppressing long-term QoE decline, and enhances overall stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a service function link optimization method for industrial metaverse networks, belonging to the field of communication network and routing optimization technology. Background Technology

[0002] With the continuous evolution of next-generation information and communication technologies, the Industrial Internet has gradually developed towards an industrial metaverse that supports digital twins, immersive interaction, and intelligent collaboration. Industrial metaverse applications typically involve multi-stage service workflows, high-frequency interactions, and stringent continuity requirements. To support this complex business logic, Collaborative Service Function Chains (SFCs) are widely used. An SFC arranges multiple service functions with sequential dependencies according to business logic, forming an ordered chain structure composed of a set of virtual network functions. The implementation of an SFC not only needs to meet the computational, bandwidth, and latency resource requirements of each functional stage but also needs to ensure that service traffic is transmitted orderly along feasible paths between adjacent stages.

[0003] In the dynamic industrial metaverse network environment, the routing decisions of the Cooperative Service Function Chain (SFC) face extremely complex network state changes. Currently, existing SFC routing technologies are mainly divided into two categories: the first category is the traditional method based on mathematical programming and heuristic design, which aims to achieve constraint satisfaction and resource-efficient routing; the second category is the method based on deep reinforcement learning, which attempts to improve adaptability in dynamic and time-varying network environments.

[0004] However, existing SFC routing technologies still have the following significant problems and shortcomings in practical applications. Existing methods are mostly limited to independent path selection on a single chain, lacking consideration for the dependencies and resource contention between concurrent SFC requests. Existing routing strategies overemphasize instantaneous Quality of Service (QoS) optimization, failing to effectively characterize the cumulative effect of experience deviations over time, and struggling to suppress long-term Quality of Experience (QoE) deficiencies. Furthermore, their dynamic collaborative decision-making capabilities in time-varying environments are insufficient. Existing deep reinforcement learning models have architectural limitations in capturing path interactions, multi-request collaboration, and handling temporal dynamics, making it difficult to output stable and efficient joint optimization strategies in complex dynamic environments.

[0005] In summary, existing technologies still have many shortcomings in the field of SFC routing. There is an urgent need for a method that can guarantee the long-term QoE stability of dynamic networks and the resource allocation and scheduling in SFC concurrent scenarios, so as to improve the cooperative routing quality of SFC. Summary of the Invention

[0006] The purpose of this invention is to provide a service function link optimization method for industrial metaverse networks to solve the problem that existing technologies lack consideration for the dependencies between concurrent SFC requests and resource competition, making it difficult to guarantee long-term experience quality.

[0007] The technical solution of this invention is:

[0008] A method for optimizing service function links in an industrial metaverse network includes the following steps:

[0009] S1. Real-time acquisition of basic network status parameters of the underlying network physical infrastructure and service attribute characteristics of concurrent collaborative service function chain requests (SFC requests), and extraction of dimensionality-reduced candidate path sets and routing attribute feature vectors for each SFC request based on a multi-topology heuristic strategy.

[0010] S2. The basic network state parameters, service attribute features and routing attribute feature vectors obtained in step S1 are synchronously input into the shared feature extraction module. The shared feature extraction module extracts the relative quality features of candidate paths, resource competition features among concurrent requests and hidden state features containing time correlation through path attention and SFC attention mechanisms.

[0011] S3. Input the relative quality features of the candidate paths, resource contention features among concurrent requests, and hidden state features obtained in step S2 into the Actor-Critic network. The Actor network in the Actor-Critic network combines the feasibility mask and performs intra-batch conflict resolution to output the final conflict-free routing path.

[0012] S4. Based on the allocated final conflict-free routing path, send flow forwarding table entries to the network data plane to actually forward SFC service traffic, and observe the actual underlying instantaneous quality of service (QoS) after the SFC service traffic transmission is completed online.

[0013] S5. Construct a continuous mapping model to transform the actual underlying instantaneous quality of service (QoS) into a user-perceived experience quality index. By calculating the experience deviation between the actual underlying instantaneous QoS and the target experience quality (QoE), dynamically evolve and update the virtual QoE deficit queue corresponding to each SFC request.

[0014] S6. Calculate the joint reward function including the instantaneous QoE deficiency penalty term and the stability regularization term, and store the state, action, reward and related training information of the current decision period into the experience replay buffer. Based on the iterative execution of steps S1-S6 in each decision period, the recurrent proximal policy optimization algorithm R-PPO is used to iteratively train the recurrent Actor-Critic network based on multi-level attention, including the shared feature extraction module of step S2 and the Actor-Critic network of step S3, to achieve continuous online optimization of the SFC routing decision strategy.

[0015] Further, step S1 specifically includes:

[0016] S11. Model the underlying physical infrastructure of the industrial metaverse network as a weighted directed or undirected graph. ,in, Represents a set of physical nodes. Represents a set of physical links; for each physical link It acquires and records basic network status parameters in real time, including the limited bandwidth capacity of the links. Current available bandwidth and propagation delay;

[0017] S12. Perform concurrent SFC request batch processing and obtain service attribute characteristics;

[0018] S13. The end-to-end routing adopts a finite candidate path abstraction to obtain a candidate path set. ;

[0019] S14. During the current decision-making period For candidate path set Each path in Based on the real-time resource status of the current network, calculate and extract routing attribute feature vectors. :

[0020]

[0021] in, Representing a path The number of hops is a static network attribute; Indicating the current decision-making period Along the path The estimated end-to-end delay calculated cumulatively; Indicating the current decision-making period Path after considering the remaining capacity of the current link The minimum available bandwidth among all links.

[0022] Further, step S12 specifically involves,

[0023] S121. The Software-Defined Network Controller (SDN Controller) collects online SFC requests in a time-interval driven manner, and aggregates concurrent requests arriving during the current decision-making period into a batch for joint processing; for any SFC request within the batch... Controller Extraction Based on the service attribute characteristics, construct a 7-dimensional request attribute vector. :

[0024]

[0025] in, This indicates the minimum bandwidth requirement for SFC to operate normally. This represents the maximum tolerable end-to-end transmission delay. This indicates the number of virtual network function nodes included in the service function chain. This indicates the time it takes for the SFC request to reach the SDN controller. This indicates the lifespan of an SFC request in the network. This indicates that the SFC is requesting the current virtual QoE deficit queue value. A unique global identifier representing an SFC request;

[0026] S122. Under the Fixed Virtual Network Function (VNF) deployment rules, the SDN controller resolves and determines the physical source nodes of the first and last VNFs carrying the SFC request. With the target node .

[0027] Further, step S13 specifically involves,

[0028] S131, Regarding the SFC request The source node obtained by parsing With the target node In the weighted graph of the underlying physical network In this process, multiple topological heuristic strategies are invoked in parallel or serially to perform pathfinding calculations, and all pathfinding results are aggregated to generate an initial path pool.

[0029] S132. The SDN controller performs hash deduplication or set intersection operations on the generated initial path pool to remove completely duplicated redundant path data and obtain the deduplicated path set.

[0030] S133. After performing multi-level conditional sorting on the deduplicated path set, the SDN controller extracts the top-ranked path from the head of the queue. The paths together form an SFC request. Fixed-size candidate path set .

[0031] Furthermore, in step S2, the shared feature extraction module extracts the relative quality features of candidate paths, resource competition features among concurrent requests, and hidden state features containing temporal correlations through inter-path attention and SFC attention mechanisms. Specifically,

[0032] S21. During the current decision-making period Obtain a global network state vector summarizing the overall network resource utilization and congestion level. For any SFC request within the batch Request its attribute vector With the global network state vector The vectors are concatenated and fused, and then input into a multilayer perceptron module consisting of learnable parameters. In the process, request-level context features are generated. :

[0033]

[0034] in, Indicates a splicing operation;

[0035] S22, Request for SFC of Feature vectors of candidate paths Path encoder using shared parameters Perform independent mapping to obtain the embedded intrinsic routing features of the path. ;

[0036] S23. To capture the relative quality of different candidate paths under the same request, the SFC request... All path embeddings are fed into the multi-head self-attention module for processing, where attention weights are calculated. By combining residual connections, the relative quality features of the k-th candidate path for the i-th SFC request are obtained. :

[0037]

[0038]

[0039] in, Indicates the first The candidate path for the first Attention weights for each candidate path; This represents the value mapping matrix in the attention mechanism; Indicates the first The first SFC request The original path embedding of the candidate paths; The feature dimension of a single attention head; Represents the query vector mapping matrix; Represents the key vector mapping matrix;

[0040] S24. The relative quality features of candidate paths. The request-level context features obtained in step S21 splicing, through the integration of MLP modules Generate context-aware path utility descriptors :

[0041] ;;

[0042] S25. Context-aware path utility descriptor The Transformer encoder mechanism is used to model dependencies between requests and extract resource contention features among concurrent requests. ;

[0043] S26. Resource contention characteristics among concurrent requests After mean pooling, the input is given to the gated recurrent unit (GRU) and combined with the hidden state h from the previous decision period t-1. t-1 The update yields the current GRU hidden state h, which includes time-related information. t .

[0044] Further, step S25 specifically involves,

[0045] S251, Context-Aware Path Utility Descriptor Perform aggregation processing to calculate the representative SFC request. SFC-level initial embedding vector of overall routing requirements and features :

[0046] ;

[0047] S252, For concurrent arrivals within the current decision-making period For each SFC request, obtain its corresponding SFC-level initial embedding set. ;

[0048] S253. To calculate the resource contention and collaborative allocation relationships among concurrent requests, an independent and learnable parameter matrix is ​​used. , , For each SFC-level initial embedding set Each embedding vector in Perform a linear mapping to generate query vectors, key vectors, and value vectors for computing the attention mechanism;

[0049] S254, the first in the quantified batch The first SFC request and the first The degree of mutual influence among SFC requests when competing for shared network resources is used to calculate the query vector. With key vector Is it the dot product, and what scaling factor is introduced? To prevent gradient vanishing due to excessively large dot product values, the Softmax function is then applied for normalization to obtain the attention weights between SFCs. :

[0050]

[0051] in, The dimension of the key vector Indicates the number of batches within the current decision-making period SFC-level initial embedding vector for each SFC request. Index for SFC requests within the batch;

[0052] S255. Based on the calculated attention weights between SFCs The value vector of all SFC requests within the batch We perform weighted summation and introduce residual connections to calculate intermediate fusion features. :

[0053]

[0054] in, This represents the value vector mapping matrix in the SFC inter-attention mechanism, used to map the first... SFC-level initial embedding vector for each SFC request The linear mapping is converted into a value vector, thus participating in the weighted aggregation of resource competition features among SFCs;

[0055] S256, Intermediate Fusion Features After being fed sequentially into the layer normalization module and the position feedforward network for nonlinear processing, the resource contention characteristics among concurrent requests are obtained. .

[0056] Further, in step S3, the Actor network in the Actor-Critic network, combined with a feasibility mask and after intra-batch conflict resolution, outputs the final conflict-free routing path. Specifically,

[0057] S31. Perform candidate path feature concatenation and joint scoring calculation within the Actor network of the Actor-Critic network: For any SFC request within the batch. Each candidate path The relative quality features of candidate paths Resource contention characteristics among concurrent requests Global network state vector and the current GRU hidden state, which includes time-related information. The concatenated feature vectors are then input into a two-layer multilayer perceptron module with ReLU activation and layer normalization. ( After that, the original routing utility score of the candidate path is obtained. :

[0058]

[0059] in, Indicates a splicing operation;

[0060] S32. Introduce a hard resource constraint filtering mechanism and construct a feasibility mask matrix: for candidate paths Set bandwidth tolerance factor With delay tolerance factor If the candidate path Meet resource constraints: minimum available bandwidth And estimate end-to-end delay If the condition is met, the path is deemed preliminarily feasible; otherwise, the corresponding position of the candidate path that does not meet the resource constraints is assigned negative infinity or a masking identifier in the feasibility mask matrix.

[0061] S33. Apply the feasibility mask matrix constructed in step S32 to the original routing utility score of the candidate paths obtained in step S31. The above process yields the route utility score after masking; for candidate paths that meet the resource constraints, the corresponding original route utility score is retained; for candidate paths that do not meet the resource constraints, the corresponding original route utility score is set to negative infinity or a masking flag.

[0062] S34. Perform a normalized Softmax operation on the masked routing utility score to obtain the normalized routing policy probability distribution. Used to indicate observations in the current batch Next, the policy network selects the first... The first SFC request The probability of each candidate path;

[0063] S35. Based on the normalized routing policy probability distribution, sample actions and generate a set of proposed joint routing actions in parallel for all N concurrent SFC requests within the batch. ,in This is represented as an SFC request. The index of the selected candidate path;

[0064] S36. Perform batch conflict resolution: Simulate deducting link resources on the proposed path according to the set order. If the total allocated bandwidth b of the links is detected... i Exceeding physical capacity : If the conflict is triggered, the conflict resolution mechanism will be activated, and the SFC request affected by the conflict will be retried to select a suboptimal alternative candidate path. If the SFC request still cannot be successfully deployed after traversing all feasible paths under the current remaining resource status, a rejection action will be generated for the SFC request that cannot be successfully deployed, and the final conflict-free routing path will be obtained.

[0065] S51. Construct a continuous mapping model to transform the actual underlying instantaneous Quality of Service (QoS) into user-perceived experience quality indicators. :

[0066]

[0067] in, Request for SFC The set basic bandwidth requirements Request for SFC During the current decision-making period The actual available bottleneck bandwidth, min is the minimum value function;

[0068] S52, the SDN controller requests SFC to maintain an independent virtual QoE deficit queue. The SDN controller obtains the preset target experience quality level. Calculate the user's subjective perception of experience quality indicators. With the target experience quality level The discrepancies in experience between periods will be analyzed, and updates will be made for the next period. Deficit queue state :

[0069]

[0070] Where min is the maximum value function.

[0071] Further, step S6 specifically involves,

[0072] S61, for each SFC request Construct a joint reward function that includes an instantaneous QoE deficit penalty term and a stability regularization term. :

[0073]

[0074]

[0075] in, This refers to the user's subjective perception of experience quality. This is a clipping and normalization process performed to prevent gradient explosion; For the current decision-making period Virtual QoE deficit queue, Q max This represents the preset normalized upper limit of the virtual QoE deficit queue; The penalty intensity coefficient, The target experience quality level; when the user's subjective perception of experience quality indicators. When the target is not met, the accumulated deficit queue value is used to amplify the penalty, and the strategy is forced to prioritize compensating for the accumulated unsatisfactory SFC requests. For QoE stability regularization, This is a regularization weight used to penalize the variance between the actual experience and the target value;

[0076] S62. The SDN controller averages and aggregates the rewards of all SFC requests within batch t(t) in the current decision period to obtain the overall batch reward. :

[0077] ;

[0078] S63. Determine the resource contention characteristics of each SFC request in the current batch. Perform mean pooling to obtain batch-level aggregated features, and then combine the batch-level aggregated features with the global network state vector. and the current GRU hidden state, which includes time-related information. Perform cascading and splicing to form the current batch status. Then, the values ​​are input into the Critic network to obtain the state value estimate. After completing the environmental interaction, the SDN controller will display the batch status for the current decision-making period. Joint actions Batch overall reward Logarithmic probability of actions and state value estimate As an experience tuple, it is stored in the experience replay buffer;

[0079] S64. After collecting the trajectory with the preset unfolding length, calculate the time difference error using the generalized dominance estimation algorithm. And calculate the dominance function to reduce variance. With target return value :

[0080]

[0081]

[0082]

[0083] in, As a discount factor, To control the attenuation parameter of the bias-variance tradeoff, This indicates that the Critic network is based on the current batch state. Output state value estimate, This indicates that the Critic network is based on the state of the next batch. Output state value estimate, Indicates the parameters of the Critic network. This represents the old Critic network parameters before the update. Indicates time The target return value; This represents the old Critic network before the parameter update and its relationship to the current state. State value estimate; Indicates from the current moment From the beginning to the end The time difference error corresponding to each time step is used to measure the deviation between the immediate reward at that time step and the Critic network state value estimate.

[0084] S65. Utilize the dominance function calculated in step S64. With target return A joint training objective function is constructed to jointly update the parameters of the recurrent Actor-Critic network based on multi-level attention.

[0085] The beneficial effects of this invention are:

[0086] I. This invention provides a service function link optimization method for industrial metaverse networks. Addressing the challenges of intense competition for resources in multi-concurrent SFCs and the difficulty in guaranteeing long-term experience quality, this invention offers a collaborative SFC routing method driven by experience quality in industrial metaverse networks. By introducing a virtual QoE deficit queue to quantify the time accumulation of experience deviations and constructing a cyclic reinforcement learning routing strategy that integrates multi-level attention mechanisms and deficit-aware rewards, this method can ensure the long-term QoE stability of dynamic networks, improve the collaborative routing quality of SFCs, effectively solve the path coupling and conflict problems of concurrent requests in dynamic network environments, optimize instantaneous transmission quality while significantly suppressing the decline in long-term QoE, and improve overall stability.

[0087] II. Compared with existing technologies, this invention overcomes the problems of local decision-making bias and experience degradation caused by resource contention, which are prone to occur in traditional single-step greedy algorithms, by adopting a global perspective concurrent multi-SFC collaborative routing decision-making mechanism, thus significantly improving the end-to-end average experience quality level. The policy network achieves collaborative optimization of QoE level and service consistency, thereby maintaining high network stability while ensuring high-quality service.

[0088] Third, this service function link optimization method for industrial metaverse networks introduces a virtual QoE deficit queue combined with a deficit-aware reward mechanism to quantify and dynamically constrain the experience difference caused by routing decisions deviating from the target service level. This effectively avoids the rapid accumulation of experience liabilities over time and significantly reduces the system's QoE. It not only improves the satisfaction of request targets and reduces the risk of QoE violations, but also significantly increases the minimum QoE threshold in the worst case and narrows the fluctuation range of service quality by optimizing the global resource allocation strategy, ensuring high reliability under high concurrency and extreme network load scenarios.

[0089] Fourth, this service function link optimization method for industrial metaverse networks, while ensuring a high-quality end-to-end service experience, effectively avoids unnecessary long-distance detours through multi-level feature extraction and environmental conflict resolution mechanisms using intelligent strategies, significantly reduces the average path hop count, and thus forms a more compact data forwarding path, greatly saving the consumption of physical underlying link and node resources. Attached Figure Description

[0090] Figure 1 This is a flowchart illustrating the service function link optimization method for industrial metaverse networks according to an embodiment of the present invention.

[0091] Figure 2 This is a schematic diagram of the QoE-driven SFC routing decision and R-PPO closed-loop optimization process provided in the embodiment;

[0092] Figure 3 This is a schematic diagram illustrating the shared feature extraction module, Actor network, and Critic network in the embodiment. Detailed Implementation

[0093] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0094] The embodiment provides a service function link routing optimization method for industrial metaverse networks, such as... Figure 1 This includes the following steps:

[0095] S1. Real-time acquisition of basic network status parameters of the underlying network physical infrastructure and service attribute characteristics of concurrent collaborative service function chain requests (SFC requests), and extraction of dimensionality-reduced candidate path sets and routing attribute feature vectors for each SFC request based on a multi-topology heuristic strategy.

[0096] S11. Model the underlying physical infrastructure of the industrial metaverse network as a weighted directed or undirected graph. ,in, Represents a set of physical nodes. Represents a set of physical links; for each physical link It acquires and records basic network status parameters in real time, including the limited bandwidth capacity of the links. Current available bandwidth and propagation delay;

[0097] S12. Perform concurrent SFC request batch processing and obtain service attribute characteristics;

[0098] S121. The Software-Defined Network Controller (SDN Controller) collects online SFC requests in a time-interval driven manner, and aggregates concurrent requests arriving during the current decision-making period into a batch for joint processing; for any SFC request within the batch... Controller Extraction Service attribute characteristics Construct a 7-dimensional request attribute vector:

[0099]

[0100] in, This indicates the minimum bandwidth requirement for SFC to operate normally. This represents the maximum tolerable end-to-end transmission delay. This indicates the number of virtual network function nodes included in the service function chain. This indicates the time it takes for the SFC request to reach the SDN controller. This indicates the lifespan of an SFC request in the network. This indicates that the SFC is requesting the current virtual QoE deficit queue value. A unique global identifier representing an SFC request;

[0101] S122. Under the Fixed Virtual Network Function (VNF) deployment rules, the SDN controller resolves and determines the physical source nodes of the first and last VNFs carrying the SFC request. With the target node .

[0102] S13. The end-to-end routing adopts a finite candidate path abstraction to obtain a candidate path set. ;

[0103] S131, Regarding the SFC request The source node obtained by parsing With the target node In the weighted graph of the underlying physical network In this process, multiple topological heuristic strategies are invoked in parallel or serially to perform pathfinding calculations, and all pathfinding results are aggregated to generate an initial path pool.

[0104] In step S131, the heuristic algorithm may employ the shortest path algorithm with the fewest hops, the minimum delay path algorithm, the bandwidth priority path algorithm, and the path algorithm that prioritizes traversing the deployed intermediate VNF physical nodes.

[0105] S132. Since the different heuristic algorithms mentioned above may find the same physical path under a specific topology, the SDN controller performs hash deduplication or set intersection operation on the generated initial path pool to remove completely duplicated path redundant data and obtain the deduplicated path set.

[0106] S133. After performing multi-level conditional sorting on the deduplicated path set, the SDN controller extracts the top-ranked path from the head of the queue. The paths together form an SFC request. Fixed-size candidate path set .

[0107] In step S133, obtaining a fixed-size candidate path set effectively reduces the complexity of the action space for the reinforcement learning agent. If the total number of paths after deduplication is insufficient... If the path is not found to be a duplicate, then all paths are retained. A multi-level conditional sorting process is performed on the deduplicated path set from step S132 to ensure the highest quality of the remaining paths: First, each path is sorted in descending order based on whether it satisfies the preference condition of sequentially traversing intermediate VNF nodes, with paths that meet the condition given the highest priority; within paths with the same primary sorting condition (i.e., both satisfying or not satisfying VNF traversal), a secondary descending sort is further performed based on the estimated bottleneck bandwidth of the paths.

[0108] In step S13, the end-to-end routing adopts a finite candidate path abstraction, which can avoid the action space explosion caused by pathfinding in the entire network physical topology.

[0109] S14. During the current decision-making period For candidate path set Each path in Based on the real-time resource status of the current network, calculate and extract routing attribute feature vectors. :

[0110]

[0111] in, Representing a path The number of hops is a static network attribute; Indicating the current decision-making period Along the path The estimated end-to-end delay calculated cumulatively; Indicating the current decision-making period Path after considering the remaining capacity of the current link The minimum available bandwidth across all links. The extracted feature vector will serve as direct input data for the subsequent intelligent policy network.

[0112] S2. The basic network state parameters, service attribute features and routing attribute feature vectors obtained in step S1 are synchronously input into the shared feature extraction module. The shared feature extraction module extracts the relative quality features of candidate paths, resource competition features among concurrent requests and hidden state features containing time correlation through path attention and SFC attention mechanisms.

[0113] In step S2, the shared feature extraction module extracts the relative quality features of candidate paths, resource competition features among concurrent requests, and hidden state features containing temporal correlations through inter-path attention and SFC attention mechanisms. Specifically, such as... Figure 3 :

[0114] S21. During the current decision-making period Obtain a global network state vector summarizing the overall network resource utilization and congestion level. For any SFC request within the batch Request its attribute vector With the global network state vector The vectors are concatenated and fused, and then input into a multilayer perceptron module consisting of learnable parameters. In the process, request-level context features are generated. :

[0115]

[0116] in, Indicates a splicing operation;

[0117] In step S21, long-term QoE deficit information can be effectively integrated with the current macroscopic state of the network.

[0118] S22, Request for SFC of Feature vectors of candidate paths Path encoder using shared parameters Perform independent mapping to obtain the embedded intrinsic routing features of the path. .

[0119] S23. To capture the relative quality of different candidate paths under the same request, the SFC request... All path embeddings are fed into the multi-head self-attention module for processing, where attention weights are calculated. By combining residual connections, the relative quality features of the k-th candidate path for the i-th SFC request are obtained. :

[0120]

[0121]

[0122] in, Indicates the first The candidate path for the first Attention weights for each candidate path; This represents the value mapping matrix in the attention mechanism; Indicates the first The first SFC request The original path embedding of the candidate paths; The feature dimension of a single attention head; Represents the query vector mapping matrix; Represents the key vector mapping matrix;

[0123] S24. The relative quality features of candidate paths. The request-level context features obtained in step S21 splicing, through the integration of MLP modules Generate context-aware path utility descriptors :

[0124] ;;

[0125] S25. Context-aware path utility descriptor The Transformer encoder mechanism is used to model dependencies between requests and extract resource contention features among concurrent requests. .

[0126] S251, Context-Aware Path Utility Descriptor Perform aggregation processing to calculate the representative SFC request. SFC-level initial embedding vector of overall routing requirements and features :

[0127] .

[0128] S252, For concurrent arrivals within the current decision-making period For each SFC request, obtain its corresponding SFC-level initial embedding set. .

[0129] S253. To calculate the resource contention and collaborative allocation relationships among concurrent requests, an independent and learnable parameter matrix is ​​used. , , For each SFC-level initial embedding set Each embedding vector in Perform a linear mapping to generate query vectors, key vectors, and value vectors for computing the attention mechanism.

[0130] S254, the first in the quantified batch The first SFC request and the first The degree of mutual influence among SFC requests when competing for shared network resources is used to calculate the query vector. With key vector Is it the dot product, and what scaling factor is introduced? To prevent gradient vanishing due to excessively large dot product values, the Softmax function is then applied for normalization to obtain the attention weights between SFCs. :

[0131]

[0132] in, The dimension of the key vector Indicates the number of batches within the current decision-making period SFC-level initial embedding vector for each SFC request. This is the index for SFC requests within the batch.

[0133] In step S254, the attention weights between SFCs It accurately describes the first step in making joint routing decisions. The SFC request needs to address the first... The level of contextual attention invested in each SFC request.

[0134] S255. Based on the calculated attention weights between SFCs The value vector of all SFC requests within the batch We perform weighted summation and introduce residual connections to calculate intermediate fusion features. :

[0135]

[0136] in, This represents the value vector mapping matrix in the SFC inter-attention mechanism, used to map the first... SFC-level initial embedding vector for each SFC request The linear mapping is transformed into a value vector, thus participating in the weighted aggregation of resource competition features among SFCs.

[0137] S256, Intermediate Fusion Features After being fed sequentially into the layer normalization module and the position feedforward network for nonlinear processing, the resource contention characteristics among concurrent requests are obtained. .

[0138] In step S25, to address the link capacity contention problem that easily arises when multiple SFCs concurrently share the underlying network, the embodiment utilizes the Transformer encoder mechanism to model the dependency between requests, obtaining resource contention characteristics among concurrent requests. This enables the recurrent Actor-Critic network based on multi-level attention in subsequent steps to detect and avoid potential resource conflicts with other concurrent SFC requests in advance when selecting a routing path for a single SFC request.

[0139] S26. Resource contention characteristics among concurrent requests After mean pooling, the input is given to the gated recurrent unit (GRU) and combined with the hidden state h from the previous decision period t-1. t-1 The update yields the current GRU hidden state h, which includes time-related information. t .

[0140] S3. Input the relative quality features of the candidate paths, the resource contention features among concurrent requests, and the hidden state features obtained in step S2 into the Actor-Critic network. The Actor network in the Actor-Critic network combines the feasibility mask and performs intra-batch conflict resolution to output the final conflict-free routing path.

[0141] In step S3, the Actor network in the Actor-Critic network, combined with a feasibility mask and after intra-batch conflict resolution, outputs the final conflict-free route path. Specifically, as follows: Figure 3 :

[0142] S31. Perform candidate path feature concatenation and joint scoring calculation: For any SFC request within the batch... Each candidate path The relative quality features of candidate paths Resource contention characteristics among concurrent requests Global network state vector and the current GRU hidden state, which includes time-related information. The concatenated feature vectors are then input into a two-layer multilayer perceptron module with ReLU activation and layer normalization. ( After that, the original routing utility score of the candidate path is obtained. :

[0143]

[0144] in, This indicates a splicing operation.

[0145] S32. Introduce a hard resource constraint filtering mechanism and construct a feasibility mask matrix: for candidate paths Set bandwidth tolerance factor With delay tolerance factor If the candidate path Meet resource constraints: minimum available bandwidth And estimate end-to-end delay If the conditions are met, the path is deemed preliminarily feasible; otherwise, the corresponding position of the candidate path that does not meet the resource constraints is assigned a negative infinity or a masking mark in the feasibility mask matrix, so that it can be physically removed in subsequent calculations.

[0146] In step S32, the embodiment introduces a hard resource constraint filtering mechanism, which can prevent the model from making ineffective explorations on obviously infeasible paths and ensure the network security of the underlying network.

[0147] S33. Apply the feasibility mask matrix constructed in step S32 to the original routing utility score of the candidate paths obtained in step S31. The above process yields the route utility score after masking. For candidate paths that meet the resource constraints, the original route utility score is retained. For candidate paths that do not meet the resource constraints, the original route utility score is set to negative infinity or a masking flag.

[0148] S34. Perform a normalized Softmax operation on the masked routing utility score to obtain the normalized routing policy probability distribution. Used to indicate observations in the current batch Next, the policy network selects the first... The first SFC request The probability of each candidate path.

[0149] S35. Based on the normalized routing policy probability distribution, sample actions and generate a set of proposed joint routing actions in parallel for all N concurrent SFC requests within the batch. ,in This is represented as an SFC request. The index of the selected candidate path.

[0150] S36. Perform batch conflict resolution: Simulate deducting link resources on the proposed path according to the set order. If the total allocated bandwidth b of the links is detected... i Exceeding physical capacity : If the conflict is triggered, the conflict resolution mechanism will be activated, and the SFC request affected by the conflict will be retried to select a suboptimal alternative candidate path. If the SFC request still cannot be successfully deployed after traversing all feasible paths under the current remaining resource status, a rejection action will be generated for the SFC request that cannot be successfully deployed, and the final conflict-free routing path will be obtained.

[0151] In step S36, since the joint action in step S35 is generated in parallel for multiple SFCs within a batch, the masking mechanism can only guarantee that a single request meets the constraints, and cannot guarantee the feasibility of joint resources when requests are concurrent within a batch. Therefore, the proposed joint action... Before the batch is executed, conflicts within the batch are resolved through the network environment interaction module.

[0152] S4. Based on the allocated final conflict-free routing path, issue flow table entries to the network data plane to actually forward SFC service traffic, and observe the actual underlying instantaneous quality of service (QoS) after the SFC service traffic transmission is completed online.

[0153] S41. After environmental conflict resolution in step S3, the device is successfully assigned to the final physical path. SFC request The SDN controller generates corresponding flow tables or routing forwarding table entries based on the path information and distributes them to the underlying physical network. The relevant forwarding nodes in the process. Simultaneously, in the resource management module, this path is officially locked and deducted. Each physical link Corresponding request bandwidth This is to prevent subsequent requests from causing resource conflicts.

[0154] S42. After the network node receives the forwarding table entries, the data plane begins actual operation. The data flow generated by the service side will follow the defined physical path. The data is sequentially and systematically transmitted through virtual network function instances deployed on different physical nodes, completing end-to-end data transmission and service chaining across multiple processing stages. During this process, the underlying physical network synchronously evolves and updates the global network state based on the real-time resource usage of all concurrent flows.

[0155] S43. During the actual forwarding of SFC service traffic or after a single service is completed, the network performance monitoring module collects data on the real network environment experienced by the SFC request. The network will observe and record the actual QoS indicators, specifically including the actual end-to-end latency experienced by the traffic. and the actual available bottleneck bandwidth provided on this path Due to the fluctuations in queuing delays and the coupling effects of concurrent resource allocation under dynamic network conditions, the actual measured values ​​collected here are... and Typically, it deviates from the candidate estimated features generated based on static topology in step S1. and .

[0156] S44. In step S3, during the batch conflict resolution, due to the severe shortage of remaining network resources and the inability to successfully deploy after traversing all feasible paths, the environment ultimately determines that the deployment is rejected (i.e., an empty path is allocated). The SFC request will be intercepted by the SDN controller at the access layer and its traffic forwarding process will be terminated. This will release any temporary memory or signaling resources that the SDN controller may have occupied during the calculation phase, generate a drop / reject log, and set the actual QoS observation value for the current decision period to zero or a maximum penalty value.

[0157] S5. Construct a continuous mapping model to transform the actual underlying instantaneous quality of service (QoS) into a user-perceived experience quality index. By calculating the experience deviation between the actual underlying instantaneous QoS and the target experience quality (QoE), dynamically evolve and update the virtual QoE deficit queue corresponding to each SFC request.

[0158] S51. Construct a continuous mapping model to transform the actual underlying instantaneous Quality of Service (QoS) into user-perceived experience quality indicators. :

[0159]

[0160] in, Request for SFC The set basic bandwidth requirements Request for SFC During the current decision-making period The actual available bottleneck bandwidth, min is the minimum value function;

[0161] In step S51, since the underlying objective transmission indicators cannot be directly equated to the user's subjective perception of the service experience, and bandwidth allocation exceeding the upper limit of service demand will not lead to a linear and infinite increase in experience, the embodiment constructs a continuous truncation mapping model from QoS to QoE, compares the actual bottleneck bandwidth with the required bandwidth, and finally... Values ​​strictly normalized to Within the range, it accurately represents the degree to which current physical network resources can immediately meet the user's service needs.

[0162] S52, the SDN controller requests SFC to maintain an independent virtual QoE deficit queue. The SDN controller obtains the preset target experience quality level. Calculate the user's subjective perception of experience quality indicators. With the target experience quality level The discrepancies in experience between periods will be analyzed, and updates will be made for the next period. Deficit queue state :

[0163]

[0164] Where max is the maximum value function.

[0165] In step S52, because the Industrial Metaverse service has extremely high requirements for interaction continuity, in order to quantify the cumulative effect of a single service quality degradation on the long-term business experience, the SDN controller maintains an independent virtual QoE deficit queue for each active SFC request on the network. And update the next period according to the rules of evolution. Deficit queue state: whenever the SDN controller actually provides Below target level When positive experience differences are generated, they will be accumulated and piled up in the deficit queue, representing the accumulation of user experience dissatisfaction; conversely, when actual... If the amount exceeds the target, the excess will be used to offset historical backlogs; The operation ensures that the value of the deficit queue is always non-negative.

[0166] In step S5, based on stochastic network optimization and Lyapunov stability theory, the aforementioned virtual queue evolution mechanism rigorously transforms the long-term QoE satisfaction objective in multi-SFC concurrent scenarios into maintaining the time-averaged stability of the virtual QoE deficit queues for all active requests. The updated request queue states are shown below. This will serve as a key environmental feedback feature, flowing to the next step to participate in the calculation of the joint reward function and the evolution of the strategy model.

[0167] S6. Calculate the joint reward function including the instantaneous QoE deficiency penalty term and the stability regularization term, and store the state, action, reward, and related training information of the current decision period into the experience replay buffer. Based on the iterative execution of steps S1-S6 in each decision period, the Recurrent Proximal Policy Optimization (R-PPO) algorithm is used to iteratively train the recurrent Actor-Critic network based on multi-level attention, including the shared feature extraction module of step S2 and the Actor-Critic network of step S3, to achieve continuous online optimization of the SFC routing decision strategy. Figure 2 :

[0168] S61, for each SFC request Construct a joint reward function that includes an instantaneous QoE deficit penalty term and a stability regularization term. :

[0169]

[0170]

[0171] in, This refers to the user's subjective perception of experience quality. This is a clipping and normalization process performed to prevent gradient explosion; For the current decision-making period Virtual QoE deficit queue, Q max This represents the preset normalized upper limit of the virtual QoE deficit queue; The penalty intensity coefficient, The target experience quality level; when the user's subjective perception of experience quality indicators. When the target is not met, the accumulated deficit queue value is used to amplify the penalty, and the strategy is forced to prioritize compensating for the accumulated unsatisfactory SFC requests. For QoE stability regularization, This is a regularization weight used to penalize the variance between the actual experience and the target value.

[0172] In step S61, in order to maintain numerical stability in policy gradient optimization and to approximately implement the Lyapunov drift plus penalty principle, this invention abandons the traditional single QoE reward and innovatively requests each SFC. A single-step joint reward function with three constraints was constructed. Part 1 For instantaneous network experience benefits, Part Two The third part is the QoE stability regularization term, which can reduce drastic fluctuations over time and improve long-term service stability.

[0173] S62. The SDN controller averages and aggregates the rewards of all SFC requests within batch t(t) in the current decision period to obtain the overall batch reward. :

[0174] ;

[0175] S63. Determine the resource contention characteristics of each SFC request in the current batch. Perform mean pooling to obtain batch-level aggregated features, and then combine the batch-level aggregated features with the global network state vector. and the current GRU hidden state, which includes time-related information. Perform cascading and splicing to form the current batch status. Then, the values ​​are input into the Critic network to obtain the state value estimate. After completing the environmental interaction, the SDN controller will display the batch status for the current decision-making period. Joint actions Batch overall reward Logarithmic probability of actions and state value estimate As an experience tuple, it is stored in the experience replay buffer;

[0176] S64. After collecting the trajectory with the preset unfolding length, calculate the time difference error using the generalized dominance estimation algorithm. And calculate the dominance function to reduce variance. With target return value :

[0177]

[0178]

[0179]

[0180] in, As a discount factor, To control the attenuation parameter of the bias-variance tradeoff, This indicates that the Critic network is based on the current batch state. Output state value estimate, This indicates that the Critic network is based on the state of the next batch. Output state value estimate, Indicates the parameters of the Critic network. This represents the old Critic network parameters before the update. Indicates time The target return value; This represents the old Critic network before the parameter update and its relationship to the current state. State value estimate; Indicates from the current moment From the beginning to the end The time difference error corresponding to each time step is used to measure the deviation between the immediate reward at that time step and the Critic network state value estimate.

[0181] S65. Utilize the dominance function calculated in step S64. With target return A joint training objective function is constructed to jointly update the parameters of the recurrent Actor-Critic network based on multi-level attention.

[0182] S651. Constructing the agent objective function of the policy network. :

[0183]

[0184]

[0185] in, Indicates the current policy network parameters. This represents the expectation of samples at each time step in the sampling trajectory, where min represents the minimum value function. This represents the probability ratio between the current strategy and the old strategy. Indicates time The estimate of the dominance function, This represents the clipping function. This represents the clipping hyperparameter. This indicates that the current policy network is in batch state. Select the joint routing action The probability, This indicates that the old policy network is in batch state. Select the joint routing action The probability of.

[0186] In step S651, a policy network agent objective function with a pruning mechanism is constructed, introducing a pruning mechanism for near-end policy optimization. Preset pruning hyperparameters are then used. By limiting the step size of policy updates, the agent objective function of the policy network is constructed. When the advantage function If positive, meaning that the current routing action performs well, the probability increase of the new policy for that action is limited to a maximum of [value missing]. Conversely, the lower limit should not be lower than This effectively prevents a single update from causing excessive disruption to the routing policy distribution.

[0187] S652. Constructing the value loss function :

[0188]

[0189] In step S652, a mean squared error evaluation function for the value network is constructed. The value network is responsible for evaluating the expected long-term returns of the current network state, providing a benchmark reference for updating the policy network. This is achieved by minimizing the estimated value of the value network output. Returns to true goals The mean square error between them guides the value network parameters. The optimization of the value network allows it to more accurately predict resource congestion trends and queue backlog risks in dynamic network environments by continuously minimizing this error.

[0190] S653, Introducing Entropy Rewards for Policy Distribution :

[0191]

[0192] in, The information entropy represents the current routing policy distribution; Indicates the current batch status Below, the probability distribution of the policy network's output actions for all candidate routes is given; where the symbol " "" indicates all possible candidate actions.

[0193] In step S653, an entropy reward for policy distribution is introduced. To prevent the agent from getting trapped in local optima too early in multi-SFC route selection decisions, such as repeatedly choosing a path that seems to have few hops but is highly prone to micro-burst congestion, an entropy reward for policy distribution is added to the objective function. The higher the entropy value, the more dispersed the probability of the policy network selecting different candidate paths. This mechanism actively encourages the model to explore the unknown network path space extensively in the early stages of training.

[0194] S654. Constructing the joint training objective function :

[0195]

[0196] Where c1 and c2 are weighting coefficients;

[0197] The S655 and SDN controllers compute the joint training objective function through the automatic differentiation engine of the deep learning framework. Regarding strategy parameters and value parameters The gradient is calculated; at the same time, the maximum absolute value of the gradient is limited by gradient pruning technique. Finally, the optimizer is used to perform gradient ascent or to perform gradient descent after negating the objective function, thus completing the parameter update iteration of the Actor-Critic network in this round of iteration.

[0198] This invention addresses the challenges of intense competition for resources in concurrent Service Function Chaining (SFC) systems and the difficulty in guaranteeing long-term experience quality in industrial metaverse networks. It provides a collaborative SFC routing method driven by experience quality in these networks. By introducing a virtual QoE deficit queue to quantify the time accumulation of experience deviations and constructing a cyclic reinforcement learning routing strategy that integrates multi-level attention mechanisms and deficit-aware rewards, this method ensures long-term QoE stability in dynamic networks, improves the collaborative routing quality of SFCs, effectively solves the path coupling and conflict problems of concurrent requests in dynamic network environments, optimizes instantaneous transmission quality while significantly suppressing long-term QoE decline, and enhances overall stability.

[0199] Compared with existing technologies, this invention overcomes the problems of local decision-making bias and experience degradation caused by resource contention that are prone to occur in traditional single-step greedy algorithms by adopting a global perspective concurrent multi-SFC collaborative routing decision-making mechanism, significantly improving the end-to-end average experience quality level. The policy network achieves collaborative optimization of QoE level and service consistency, thereby maintaining high network stability while ensuring high-quality service.

[0200] This service function link optimization method for industrial metaverse networks introduces a virtual QoE deficit queue combined with a deficit-aware reward mechanism to quantify and dynamically constrain the experience difference caused by routing decisions deviating from the target service level. This effectively avoids the rapid accumulation of experience liabilities over time and significantly reduces the system's QoE. It not only improves the satisfaction of request targets and reduces the risk of QoE violations, but also significantly increases the minimum QoE threshold in the worst case and narrows the fluctuation range of service quality by optimizing the global resource allocation strategy, ensuring high reliability under high concurrency and extreme network load scenarios.

[0201] This service function link optimization method for industrial metaverse networks, while ensuring a high-quality end-to-end service experience, effectively avoids unnecessary long-distance detours through multi-level feature extraction and environmental conflict resolution mechanisms using intelligent strategies. It significantly reduces the average path hop count, thereby forming a more compact data forwarding path and greatly saving the consumption of physical underlying link and node resources.

[0202] The foregoing has shown and described the main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention, all of which fall within the scope of the claims. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A service function link optimization method for industrial metaverse networks, characterized in that: Includes the following steps, S1. Real-time acquisition of basic network status parameters of the underlying network physical infrastructure and service attribute characteristics of concurrent collaborative service function chain requests (SFC requests), and extraction of candidate path sets and routing attribute feature vectors for each SFC request based on a multi-topology heuristic strategy. S2. The basic network state parameters, service attribute features and routing attribute feature vectors obtained in step S1 are synchronously input into the shared feature extraction module. The shared feature extraction module extracts the relative quality features of candidate paths, resource competition features among concurrent requests and hidden state features containing time correlation through path attention and SFC attention mechanisms. S3. Input the relative quality features of the candidate paths, resource contention features among concurrent requests, and hidden state features obtained in step S2 into the Actor-Critic network. The Actor network in the Actor-Critic network combines the feasibility mask and performs intra-batch conflict resolution to output the final conflict-free routing path. S4. Based on the allocated final conflict-free routing path, send flow forwarding table entries to the network data plane to actually forward SFC service traffic, and observe the actual underlying instantaneous quality of service (QoS) after the SFC service traffic transmission is completed online. S5. Construct a continuous mapping model to transform the actual underlying instantaneous quality of service (QoS) into a user-perceived experience quality index. By calculating the experience deviation between the actual underlying instantaneous QoS and the target experience quality (QoE), dynamically evolve and update the virtual QoE deficit queue corresponding to each SFC request. S6. Calculate the joint reward function including the instantaneous QoE deficiency penalty term and the stability regularization term, and store the state, action, reward and related training information of the current decision period into the experience replay buffer. Based on the iterative execution of steps S1-S6 in each decision period, the recurrent proximal policy optimization algorithm R-PPO is used to iteratively train the recurrent Actor-Critic network based on multi-level attention, including the shared feature extraction module of step S2 and the Actor-Critic network of step S3, to achieve continuous online optimization of the SFC routing decision strategy.

2. The service function link routing optimization method for industrial metaverse networks as described in claim 1, characterized in that: Step S1 is as follows: S11. Model the underlying physical infrastructure of the industrial metaverse network as a weighted directed or undirected graph. ,in, Represents a set of physical nodes. Represents a set of physical links; for each physical link It acquires and records basic network status parameters in real time, including the limited bandwidth capacity of the links. Current available bandwidth and propagation delay; S12. Perform concurrent SFC request batch processing and obtain service attribute characteristics; S13. The end-to-end routing adopts a finite candidate path abstraction to obtain a candidate path set. ; S14. During the current decision-making period For candidate path set Each path in Based on the real-time resource status of the current network, calculate and extract routing attribute feature vectors. : , in, Representing a path The number of hops is a static network attribute; Indicating the current decision-making period Along the path The estimated end-to-end delay calculated cumulatively; Indicating the current decision-making period Path after considering the remaining capacity of the current link The minimum available bandwidth among all links.

3. The service function link routing optimization method for industrial metaverse networks as described in claim 2, characterized in that: Step S12, specifically, S121. The Software-Defined Network Controller (SDN Controller) collects online SFC requests in a time-interval driven manner, and aggregates concurrent requests arriving during the current decision-making period into a batch for joint processing; for any SFC request within the batch... The SDN controller extracts the SFC request. Based on the service attribute characteristics, construct the corresponding 7-dimensional request attribute vector. : , in, This indicates the minimum bandwidth requirement for SFC to operate normally. This represents the maximum tolerable end-to-end transmission delay. This indicates the number of virtual network function nodes included in the service function chain. This indicates the time it takes for the SFC request to reach the SDN controller. This indicates the lifespan of an SFC request in the network. This indicates that the SFC is requesting the current virtual QoE deficit queue value. A unique global identifier representing an SFC request; S122. Under the Fixed Virtual Network Function (VNF) deployment rules, the SDN controller resolves and determines the physical source nodes of the first and last VNFs carrying the SFC request. With the target node .

4. The service function link routing optimization method for industrial metaverse networks as described in claim 2, characterized in that: Step S13, specifically, S131, Regarding the SFC request The source node obtained by parsing With the target node In the weighted graph of the underlying physical network In this process, multiple topological heuristic strategies are invoked in parallel or serially to perform pathfinding calculations, and all pathfinding results are aggregated to generate an initial path pool. S132. The SDN controller performs hash deduplication or set intersection operations on the generated initial path pool to remove completely duplicated redundant path data and obtain the deduplicated path set. S133. After performing multi-level conditional sorting on the deduplicated path set, the SDN controller extracts the top-ranked path from the head of the queue. The paths together form an SFC request. Fixed-size candidate path set .

5. The service function link routing optimization method for industrial metaverse networks as described in any one of claims 1-4, characterized in that: In step S2, the shared feature extraction module extracts the relative quality features of candidate paths, resource competition features among concurrent requests, and hidden state features containing temporal correlations through inter-path attention and SFC attention mechanisms. Specifically, S21. During the current decision-making period Obtain a global network state vector summarizing the overall network resource utilization and congestion level. ; For any SFC request within the batch The requested attribute vector is compared with the global network state vector. The vectors are concatenated and fused, and then input into a multilayer perceptron module consisting of learnable parameters. In the process, request-level context features are generated. : , in, Indicates a splicing operation; S22, Request for SFC of Feature vectors of candidate paths Path encoder using shared parameters Perform independent mapping to obtain the embedded intrinsic routing features of the path. ; S23. To capture the relative quality of different candidate paths under the same request, the SFC request... All path embeddings are fed into the multi-head self-attention module for processing, where attention weights are calculated. By combining residual connections, the relative quality features of the k-th candidate path for the i-th SFC request are obtained. : , , in, Indicates the first The candidate path for the first Attention weights for each candidate path; This represents the value mapping matrix in the attention mechanism; Indicates the first The first SFC request The original path embedding of the candidate paths; The feature dimension of a single attention head; Represents the query vector mapping matrix; Represents the key vector mapping matrix; S24. The relative quality features of candidate paths. The request-level context features obtained in step S21 splicing, through the integration of MLP modules Generate context-aware path utility descriptors : ; S25. Context-aware path utility descriptor The Transformer encoder mechanism is used to model dependencies between requests and extract resource contention features among concurrent requests. ; S26. Resource contention characteristics among concurrent requests After mean pooling, the input is given to the gated recurrent unit (GRU) and combined with the hidden state h from the previous decision period t-1. t-1 The update yields the current GRU hidden state h, which includes time-related information. t .

6. The service function link routing optimization method for industrial metaverse networks as described in claim 5, characterized in that: Step S25, specifically, S251, Context-Aware Path Utility Descriptor Perform aggregation processing to calculate the representative SFC request. SFC-level initial embedding vector of overall routing requirements and features : ; S252, For concurrent arrivals within the current decision-making period For each SFC request, obtain its corresponding SFC-level initial embedding set. ; S253. To calculate the resource contention and collaborative allocation relationships among concurrent requests, an independent and learnable parameter matrix is ​​used. , , For each SFC-level initial embedding set Each embedding vector in Perform a linear mapping to generate query vectors, key vectors, and value vectors for computing the attention mechanism; S254, the first in the quantified batch The first SFC request and the first The degree of mutual influence among SFC requests when competing for shared network resources is used to calculate the query vector. With key vector Is it the dot product, and what scaling factor is introduced? To prevent gradient vanishing due to excessively large dot product values, the Softmax function is then applied for normalization to obtain the attention weights between SFCs. : , in, The dimension of the key vector Indicates the number of batches within the current decision-making period SFC-level initial embedding vector for each SFC request. Index for SFC requests within the batch; S255. Based on the calculated attention weights between SFCs The value vector of all SFC requests within the batch We perform weighted summation and introduce residual connections to calculate intermediate fusion features. : , in, This represents the value vector mapping matrix in the SFC inter-attention mechanism, used to map the first... SFC-level initial embedding vector for each SFC request The linear mapping is converted into a value vector, thus participating in the weighted aggregation of resource competition features among SFCs; S256, Intermediate Fusion Features After being fed sequentially into the layer normalization module and the position feedforward network for nonlinear processing, the resource contention characteristics among concurrent requests are obtained. .

7. The service function link routing optimization method for industrial metaverse networks as described in any one of claims 2-4, characterized in that: In step S3, the Actor network in the Actor-Critic network, combined with a feasibility mask and after intra-batch conflict resolution, outputs the final conflict-free route path. Specifically, S31. Perform candidate path feature concatenation and joint scoring calculation within the Actor network of the Actor-Critic network: For any SFC request within the batch. Each candidate path The relative quality features of candidate paths Resource contention characteristics among concurrent requests Global network state vector and the current GRU hidden state, which includes time-related information. The concatenated feature vectors are then input into a two-layer multilayer perceptron module with ReLU activation and layer normalization. ( After that, the original routing utility score of the candidate path is obtained. : , in, Indicates a splicing operation; S32. Introduce a hard resource constraint filtering mechanism and construct a feasibility mask matrix: for candidate paths Set bandwidth tolerance factor With delay tolerance factor If the candidate path Meet resource constraints: minimum available bandwidth And estimate end-to-end delay If the condition is met, the path is deemed preliminarily feasible; otherwise, the corresponding position of the candidate path that does not meet the resource constraints is assigned negative infinity or a masking identifier in the feasibility mask matrix. S33. Apply the feasibility mask matrix constructed in step S32 to the original routing utility score of the candidate paths obtained in step S31. The above process yields the route utility score after masking; for candidate paths that meet the resource constraints, the corresponding original route utility score is retained; for candidate paths that do not meet the resource constraints, the corresponding original route utility score is set to negative infinity or a masking flag. S34. Perform a normalized Softmax operation on the masked routing utility score to obtain the normalized routing policy probability distribution. Used to indicate observations in the current batch Next, the policy network selects the first... The first SFC request The probability of each candidate path; S35. Based on the normalized routing policy probability distribution, sample actions and generate a set of proposed joint routing actions in parallel for all N concurrent SFC requests within the batch. ,in This is represented as an SFC request. The index of the selected candidate path; S36. Perform batch conflict resolution: Simulate deducting link resources on the proposed path according to the set order. If the total allocated bandwidth b of the links is detected... i Exceeding physical capacity : If the conflict is triggered, the conflict resolution mechanism will be activated, and the SFC request affected by the conflict will be retried to select a suboptimal alternative candidate path. If the SFC request still cannot be deployed successfully after traversing all feasible paths under the current remaining resource status, a rejection action will be generated for the SFC request that cannot be deployed successfully, and the final conflict-free routing path will be obtained.

8. The service function link routing optimization method for industrial metaverse networks as described in any one of claims 1-4, characterized in that: Step S5, specifically, S51. Construct a continuous mapping model to transform the actual underlying instantaneous Quality of Service (QoS) into user-perceived experience quality indicators. : , in, Request for SFC The set basic bandwidth requirements Request for SFC During the current decision-making period The actual available bottleneck bandwidth, min is the minimum value function; S52, the SDN controller requests SFC to maintain an independent virtual QoE deficit queue. The SDN controller obtains the preset target experience quality level. Calculate the user's subjective perception of experience quality indicators. With the target experience quality level The discrepancies in experience between periods will be analyzed, and updates will be made for the next period. Deficit queue state : , Where max is the maximum value function.

9. The service function link routing optimization method for industrial metaverse networks as described in any one of claims 1-4, characterized in that: Step S6, specifically, S61, for each SFC request Construct a joint reward function that includes an instantaneous QoE deficit penalty term and a stability regularization term. : , , in, This refers to the user's subjective perception of experience quality. This is a clipping and normalization process performed to prevent gradient explosion; For the current decision-making period Virtual QoE deficit queue, Q max This represents the preset normalized upper limit of the virtual QoE deficit queue; The penalty intensity coefficient, The target experience quality level; when the user's subjective perception of experience quality indicators. When the target is not met, the accumulated deficit queue value is used to amplify the penalty, and the strategy is forced to prioritize compensating for the accumulated unsatisfactory SFC requests. For QoE stability regularization, This is a regularization weight used to penalize the variance between the actual experience and the target value; S62. The SDN controller averages and aggregates the rewards of all SFC requests within batch 𝓑(t) in the current decision period to obtain the overall batch reward. : ; S63. Determine the resource contention characteristics of each SFC request in the current batch. Perform mean pooling to obtain batch-level aggregated features, and then combine the batch-level aggregated features with the global network state vector. and the current GRU hidden state, which includes time-related information. Perform cascading and splicing to form the current batch status. Then, the values ​​are input into the Critic network to obtain the state value estimate. After completing the environmental interaction, the SDN controller will display the batch status for the current decision-making period. Joint actions Batch overall reward Logarithmic probability of actions and state value estimate As an experience tuple, it is stored in the experience replay buffer; S64. After collecting the trajectory with the preset unfolding length, calculate the time difference error using the generalized dominance estimation algorithm. And calculate the dominance function to reduce variance. With target return value : , , , in, As a discount factor, To control the attenuation parameter of the bias-variance tradeoff, This indicates that the Critic network is based on the current batch state. Output state value estimate, This indicates that the Critic network is based on the state of the next batch. Output state value estimate, Indicates the parameters of the Critic network. This represents the old Critic network parameters before the update. Indicates time The target return value; This represents the old Critic network before parameter updates and its relationship to the current state. State value estimate; Indicates from the current moment From the beginning to the end The time difference error corresponding to each time step is used to measure the deviation between the immediate reward at that time step and the Critic network state value estimate. S65. Utilize the dominance function calculated in step S64. With target return A joint training objective function is constructed to jointly update the parameters of the recurrent Actor-Critic network based on multi-level attention.