A method for security isolation and threat detection of virtualized network slices

By constructing a ternary fusion framework of zero trust and deep reinforcement learning, combined with multi-agent collaboration and adaptive linkage mechanisms, the dynamic adaptability and linkage response of security isolation and threat detection in virtualized network slicing environments are solved, achieving efficient security protection and rapid response.

CN122640183APending Publication Date: 2026-08-25HUNAN UNIV OF SCI & ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610739425.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In existing virtualized network slicing environments, security isolation and threat detection solutions face challenges such as insufficient adaptability to dynamic resource changes, lack of coordinated response from independent systems, insufficient adaptability of traditional detection methods, and insufficient cross-slice security status awareness.

Method used

A three-element fusion framework integrating zero trust, deep reinforcement learning, and intelligent linkage is constructed. It adopts a hierarchical threat model, multi-level identity verification and authorization mechanism, and combines a multi-agent collaborative architecture and adaptive linkage mechanism to achieve quantitative trust assessment, multi-agent threat detection and adaptive policy linkage, forming a closed-loop feedback control system.

Benefits of technology

It enables proactive and intelligent defense against virtualized network slice environments, improves the matching between security policies and business needs, enhances the accuracy and response speed of complex threat detection, and reduces false alarm rate and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640183A_ABST
    Figure CN122640183A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of network security, and specifically relates to a security isolation and threat detection method for a virtualized network slice, specific steps of the security isolation and threat detection method being as follows: virtualized network slice threat model construction, dynamic security isolation algorithm based on zero trust, deep reinforcement learning threat detection model design, and isolation strategy and detection result self-adaptive linkage mechanism; existing schemes mostly adopt static weights or binary judgments, a multi-time scale dynamic trust decay model is constructed in this paper, and three evaluations of historical behavior, real-time behavior and environmental context are fused; existing schemes only realize two-layer linkage of detection and response, three-layer intelligent linkage of detection, evaluation and execution is designed in this paper, and a closed-loop optimization from threat awareness to protection strategy is established; existing schemes rely on traditional machine learning or single-agent reinforcement learning, a multi-agent cooperation architecture is adopted in this paper, and complex threat detection accuracy is significantly improved through information sharing and collaborative decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, specifically to a method for secure isolation and threat detection of virtualized network slices. Background Technology

[0002] Network slicing security research originated from the 5G standardization process, and its technological evolution can be divided into three stages: the boundary protection stage (2018-2019), mainly based on traditional firewalls and VPNs; the dynamic isolation stage (2020-2021), introducing SDN / NFV technologies to achieve flexible control; and the intelligent protection stage (2022-present), integrating AI technologies to improve adaptive capabilities. International research focuses on standards development and theoretical modeling, while domestic research emphasizes engineering implementation and performance optimization.

[0003] Virtualized network slicing relies on software-defined networking and network function virtualization technologies to achieve multi-tenant network services through resource abstraction and virtualization. This architecture employs a design pattern that separates the control plane from the data plane, enabling dynamic creation, configuration, and management of slices. However, multi-tenant sharing of physical infrastructure introduces significant security risks. Malicious tenants may use side-channel attacks to obtain sensitive information from other slices, denial-of-service attacks caused by resource contention can threaten the quality of slice services, and vulnerabilities in the virtualization layer make cross-slice attacks possible.

[0004] Traditional network security boundaries become very blurred in virtualized environments, which allows attackers to leverage the characteristics of virtualization technology to carry out more covert and complex attack behaviors;

[0005] Security isolation technology has evolved from static to dynamic, and from coarse-grained to fine-grained. Early static isolation methods based on VLANs / VPNs were complex to configure and lacked flexibility. Subsequent developments, such as dynamic policy configuration of SDN controllers and NFV virtual network function chain technology, have significantly improved the accuracy and manageability of isolation.

[0006] Current network slicing security isolation technologies are mainly based on Virtual Private Networks (VPNs), traffic engineering, and access control mechanisms. VLAN and VPN-based isolation schemes achieve traffic separation between slices through logical identifiers, but suffer from configuration complexity issues in virtualized environments. Software-defined boundary (SDB) technologies provide fine-grained access control using encrypted tunnels and dynamic policies, but face performance overhead and scalability limitations. Micro-segmentation technologies achieve more precise isolation granularity through network function virtualization, but lack adaptive response capabilities to dynamic threats. Zero-trust architecture offers a new approach to slice isolation, emphasizing continuous authentication and the principle of least privilege, but its specific implementation mechanisms in virtualized network slicing environments still require further research and improvement.

[0007] Threat detection technology has gradually evolved from traditional rule - based and statistical analysis methods to intelligent detection methods based on machine learning. Breakthroughs in deep learning for automatic feature extraction and the advantages of reinforcement learning in adapting to dynamic environments have provided new technical paths for threat detection in virtualized environments.

[0008] Threat detection technologies in virtualized environments mainly include signature - based detection, abnormal behavior analysis, and machine learning methods. Signature matching technology identifies known threats through predefined attack patterns, with high accuracy but difficulty in dealing with zero - day attacks. Statistical analysis methods detect abnormal activities by establishing baselines of normal behavior, capable of discovering unknown threats but with a high false - positive rate. Deep learning technology improves detection accuracy using the feature - learning ability of neural networks and performs well in processing large - scale network traffic. Reinforcement learning optimizes detection strategies by interacting with the environment and can adapt to dynamically changing threat environments. Heterogeneous graph neural networks can process graph - structured data containing different types of nodes and edges, better completing the representation learning of complex structures and having very wide applications in the communication field. However, existing methods lack targeted design in specific scenarios of virtualized network slices and are difficult to effectively handle the detection challenges brought by complex interactions between slices and dynamic resource allocation.

[0009] However, existing applications of reinforcement learning mainly focus on resource scheduling optimization, and research on the coordination between threat detection and security protection is relatively lagging.

[0010] Existing security isolation and threat detection solutions face multiple technical challenges in the virtualized network slice environment. Static isolation policies cannot adapt to the dynamic changes of slice resources and real - time adjustments of service requirements, resulting in a mismatch between security policies and business needs. Independent threat detection systems lack effective coordination with isolation mechanisms and cannot achieve the linkage response between detection results and protection strategies. The feature engineering of traditional detection methods relies on expert knowledge and is insufficient in adapting to complex attacks in virtualized environments. The balance problem between performance overhead and security intensity restricts the practical deployment of the solution. Especially in resource - constrained edge - computing scenarios, the lack of cross - slice security state awareness and global threat situation analysis capabilities affects the overall security protection effect and response speed.

[0011] In response to the above - mentioned technical challenges, existing research lacks a comprehensive solution that deeply integrates the zero - trust architecture, deep reinforcement learning, and linkage mechanism. This application innovatively constructs a ternary fusion framework of zero - trust, deep reinforcement learning, and intelligent linkage, achieving an organic combination of quantitative trust assessment, multi - agent threat detection, and adaptive policy linkage, breaking through the limitations of traditional security technologies being independent of each other, and providing a new path for proactive intelligent defense in virtualized network slices. Summary of the Invention

[0012] The purpose of this invention is to provide a method for secure isolation and threat detection of virtualized network slices, so as to solve the problems mentioned in the background art.

[0013] To achieve the above objectives, the present invention provides the following technical solution: a method for secure isolation and threat detection of virtualized network slices, the specific steps of which are as follows:

[0014] Step 1: Construction of Virtualized Network Slice Threat Model: The virtualized network slice threat model adopts a layered modeling method, dividing threats into four dimensions: physical layer, virtualization layer, slice layer and application layer. The threat model adopts a layered architecture design, from bottom to top: physical layer, virtualization layer, slice layer and application layer.

[0015] Step 2: Zero-trust-based dynamic security isolation algorithm: Adopting the core principle of never trusting and always verifying, a multi-layered authentication and authorization mechanism is constructed. The algorithm achieves secure isolation between slices through quantitative trust assessment and dynamic policy adjustment; it calculates the trust level of entities by comprehensively considering multiple dimensions; the micro-segment isolation strength is dynamically adjusted according to the threat level and resource importance.

[0016] The algorithm uses a strategy-as-code approach to automate the deployment and updating of isolation rules, supports elastic scaling of slice resources, ensures the confidentiality and integrity of data transmission between slices through encrypted communication and data protection mechanisms, adjusts the isolation strategy in real time based on threat intelligence and security event triggers, and realizes the transformation from passive protection to active defense based on a quantitative model.

[0017] Step 3: Design of Deep Reinforcement Learning Threat Detection Model: The deep reinforcement learning threat detection model adopts a multi-agent architecture, with each agent responsible for threat monitoring of a specific slice or network region. The deep reinforcement learning threat detection framework adopts a multi-agent collaborative architecture, in which multiple agents achieve distributed threat detection and joint optimization through information sharing and collaborative decision-making. It covers multi-agent collaboration mechanisms, state space construction, reward function design, and policy optimization processes. This framework leverages the information interaction and collaborative decision-making between agents to achieve intelligent detection and response to complex threats in virtualized network slice environments.

[0018] Step 4: Adaptive Linkage Mechanism between Isolation Strategy and Detection Results: The adaptive linkage mechanism establishes a closed-loop feedback control system between threat detection results and isolation strategies. The threat level assessment module calculates the threat severity based on the detection results, threat type, and scope of impact, triggering the corresponding level of isolation response. The strategy decision engine uses a combination of fuzzy logic and expert systems to automatically select the optimal combination of isolation strategies. The linkage mechanism supports progressive isolation upgrades, from traffic restrictions and privilege downgrades to multi-level responses of complete isolation. The real-time strategy optimization module continuously adjusts linkage parameters and decision thresholds by analyzing historical data and current network status. The mechanism also includes fault recovery and rollback strategies to ensure rapid restoration of normal service in the event of false alarms. The linkage effect evaluation module monitors the effectiveness of isolation measures and provides feedback information for strategy optimization.

[0019] Preferably, in step one, physical layer threats originate from hardware devices and infrastructure, virtualization layer threats exploit vulnerabilities in virtualization technology, slicing layer threats target the isolation mechanism between slices, and application layer threats attack upper-layer business services. The threat propagation paths include vertical attacks (cross-layer penetration) and lateral diffusion (same-layer spread), forming a complete threat situation analysis framework.

[0020] Preferably, in step two, the algorithm is formally defined as follows:

[0021] Definition 1 (Zero Trust State Space): S = {s_auth, s_authz, s_monitor, s_isolate}, representing authentication, authorization, monitoring, and isolation states respectively;

[0022] Definition 2 (Action Space): A = {allow, deny, quarantine(t), monitor(level)}, where t is the quarantine time and level is the monitoring level;

[0023] Definition 3 (State transition function): δ: S×A×T→S, satisfies the Markov property.

[0024] Preferably, in step two, the formula for calculating entity trust level is:

[0025] (1)

[0026] Where T(e,t) represents the trust level of entity e at time t, T hist (e) represents the basic trust value based on historical behavior, T behav (e,t) represents the trust score for the current behavior, where T is the trust score for the current behavior. context (e,t) is the environmental context trust value, where α, β, and γ are weight coefficients and satisfy α+β+γ=1;

[0027] The probability of access permission is calculated as follows:

[0028] (2)

[0029] Among them, P access For the probability of access permission, A i w is the standardized value of the i-th attribute. i Let θ be the corresponding weight, θ be the decision threshold, and n be the total number of attributes. When P access Access requests are permitted when the threshold is exceeded.

[0030] The formula for calculating the isolation level is:

[0031] (3)

[0032] Among them, L isolation For isolation level, R threat For threat risk value, C resource L is the resource importance coefficient. max This is the highest level of isolation.

[0033] Preferably, in step three, the model design includes the following core components:

[0034] (1) State space construction encodes network traffic characteristics, system behavior indicators, resource usage and security event information into high-dimensional state vectors, and uses attention mechanism and graph neural network to extract spatiotemporal correlation features and construct dependency graphs between slices;

[0035] The multi-head attention mechanism uses 8 attention heads, each with a dimension of 64, and calculates temporal feature weights through a Query-Key-Value mechanism.

[0036] (4)

[0037] Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k With the key vector dimension, this mechanism effectively captures the temporal dependence of threat behavior. The graph neural network adopts a 3-layer GCN architecture with hidden layer dimensions of [256, 128, 64]. The adjacency matrix is ​​constructed based on the communication frequency between slices.

[0038] (5)

[0039] Among them, A ij f represents the connection weight between nodes i and j. i and f j Let i and j be the feature vectors of nodes i and j, respectively, σ be the bandwidth parameter, and the message passing mechanism be:

[0040] (6);

[0041] Among them, h v (l+1) Let W be the hidden state of node v at level l+1. l Here, is the weight matrix of the l-th layer, AGG is the aggregation function, N(v) is the set of neighbors of node v, σ is the activation function, and spatiotemporal feature fusion is achieved by combining the temporal attention output and the spatial graph convolution result through an adaptive weight mechanism:

[0042] (7)

[0043] Among them, h final For the final fusion feature, h temporal For temporal attention output, h spatial The result is the spatial graph convolution, where λ is the adaptive weight parameter.

[0044] State representations, which integrate statistical features, sequence patterns, and graph structure information, can capture complex attack behavior patterns.

[0045] (2) In terms of reward function design, multiple objectives such as detection accuracy, false alarm rate, detection latency and resource consumption are comprehensively considered. The positive reward mechanism is used to encourage correct threat identification and timely response, while the negative reward mechanism is used to punish false alarms and missed alarms. The dynamic reward adjustment strategy allocates weights based on the severity of the threat and the impact on business, so as to balance detection performance and system efficiency.

[0046] (3) The state space adopts a 288-dimensional feature vector: 96-dimensional network traffic features (packet size, protocol distribution, connection mode), 96-dimensional system behavior features (CPU, memory, I / O utilization), 48-dimensional resource usage features (bandwidth, storage usage), and 48-dimensional security event features (alarm type, severity).

[0047] (4) Multi-agent cooperation mechanism: The number of agents is based on It is determined that each active slice deploys a local detection agent, with an additional global coordinator. The communication topology adopts a hybrid structure of star and local mesh. The local agent is connected to the coordinator in a star topology, and mesh sidechains are established between adjacent slices for lateral threat perception. The weighted voting weights are dynamically updated based on the historical accuracy of each agent over the past 10 windows.

[0048] (8)

[0049] Coordinator's weighted score If a threat is identified and isolation linkage is triggered, when S∈[0.4,0.6), the coordinator retrieves the complete features and makes a second determination through the PPO policy network, outputs the final confidence score, and broadcasts the result to each local agent to complete the closed loop.

[0050] (5) The reward function weights are allocated based on business priorities: detection accuracy w1 = 0.4 (core objective), false alarm penalty w2 = -0.3 (reduce operation and maintenance costs), response delay penalty w3 = -0.2 (ensure real-time performance), and resource consumption penalty w4 = -0.1 (system efficiency).

[0051] (6) Training hyperparameter settings: PPO algorithm, learning rate 3e-4, batch size 256, experience replay buffer 10000, ε-greedy policy decayed from 0.9 to 0.1, gradient clipping threshold 0.5.

[0052] Preferably, in step three, the framework includes four core layers: a state perception layer, a feature extraction layer, a decision layer, and an execution layer. The state perception layer is responsible for the real-time collection of network traffic and system behavior data. The feature extraction layer extracts spatiotemporal correlation features through attention mechanisms and graph neural networks. The decision layer performs threat judgment based on a multi-agent reinforcement learning algorithm. The execution layer outputs detection results and response strategies.

[0053] Compared with the prior art, the beneficial effects of the present invention are:

[0054] 1) Trust assessment model: Existing solutions mostly use static weights or binary judgments. This paper constructs a multi-timescale dynamic trust decay model, which integrates historical behavior, real-time behavior and environmental context for triple assessment.

[0055] 2) Linkage Mechanism Architecture: Existing solutions only achieve two-layer linkage of detection and response. This paper designs a three-layer intelligent linkage of detection, evaluation, and execution to establish a closed-loop optimization from threat perception to protection strategy.

[0056] 3) Collaborative detection algorithm: Existing solutions rely on traditional machine learning or single-agent reinforcement learning. This paper adopts a multi-agent collaborative architecture, which significantly improves the detection accuracy of complex threats through information sharing and collaborative decision-making.

[0057] The core breakthrough of this application's fusion solution lies in the bidirectional driving coupling of zero trust and DRL. The trust assessment result is directly used as the state input of DRL, and the DRL detection conclusion updates the trust score in reverse. The two form a closed loop rather than operating independently, which is a design that is generally lacking in existing solutions. Attached Figure Description

[0058] Figure 1 A layered diagram of the threat model for virtualized network slicing;

[0059] Figure 2 Architecture diagram of a deep reinforcement learning threat detection framework;

[0060] Figure 3 The graph shows the performance simulation curves for threat detection. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Example:

[0063] Please see Figure 1-3 The present invention provides a technical solution:

[0064] A method for secure isolation and threat detection of virtualized network slices, the specific steps of which are as follows:

[0065] Step 1: Construction of Virtualized Network Slice Threat Model: The virtualized network slice threat model adopts a layered modeling method, dividing threats into four dimensions: physical layer, virtualization layer, slice layer and application layer. The threat model adopts a layered architecture design, from bottom to top: physical layer, virtualization layer, slice layer and application layer.

[0066] Table 1 Threat Classification Table for Virtualized Network Slices

[0067] Threat Level Main threat types Typical attack methods Risk level probability of occurrence Physical layer Hardware tampering Malicious chip implantation, firmware modification high Low (5%) Equipment damage Physical equipment damage, line interruption middle Low (3%) Virtualization layer Virtual machine escape Hypervisor vulnerability exploitation high China (12%) Container Breakout Container runtime vulnerabilities Medium and high Middle (15%) Resource pool pollution Malicious Mirror Injection middle China (22%) slice layer Cross-slice access Bypassing the isolation mechanism high China (23%) Slice hijacking Stealing administrative privileges high Low (10%) Resource contention attack Malicious resource predation middle High (35%) Application layer Malicious code injection SQL injection, XSS attack middle High (35%) Data breach Theft of sensitive information high Middle (25%) Denial-of-service attack Application-layer DoS middle High (40%)

[0068] Table 1 shows the risk level and probability distribution of threats at each layer. As can be seen from the table, application layer threats have the highest probability of occurrence (35%-40%), but their risk level is relatively low; while high-risk threats at the physical layer and slicing layer, although having a lower probability of occurrence, will cause serious impact once they occur.

[0069] Threat propagation analysis indicates that attackers typically employ either a "bottom-up" or "top-down" penetration strategy.

[0070] Typical attack paths include:

[0071] (1) Physical layer attack → Virtualization layer control → Slice layer penetration → Application layer destruction;

[0072] (2) Application layer vulnerability → slice layer privilege acquisition → virtualization layer escape → complete system control. The impact of the threat has a cascading amplification effect; a single point of breach may lead to damage to multiple slices.

[0073] like Figure 1 As shown, Figure 1 The model details the distribution of specific threat types, vertical attack propagation paths, and horizontal threat diffusion patterns in the four-layer threat model. The model introduces a threat propagation graph to describe attack paths and the scope of impact, and uses probabilistic risk assessment to quantify the probability of occurrence and potential losses of different threats.

[0074] Taking cross-slice data leakage as an example: Attackers first gain initial access through web application vulnerabilities, then exploit container escape techniques to breach virtualization boundaries, and finally bypass slice isolation mechanisms to access data in neighboring slices. This attack chain involves three layers: application layer → virtualization layer → slice layer, demonstrating the hierarchical propagation characteristics of the threat model. The threat model also considers the dynamic characteristics of the slice lifecycle, establishing a threat evolution pattern and time correlation analysis framework, providing a theoretical basis for security policy formulation and threat prediction.

[0075] Step 2: Zero-trust-based dynamic security isolation algorithm: Adopting the core principle of never trusting and always verifying, a multi-layered authentication and authorization mechanism is constructed. The algorithm achieves secure isolation between slices through quantitative trust assessment and dynamic policy adjustment; it calculates the trust level of entities by comprehensively considering multiple dimensions; the micro-segment isolation strength is dynamically adjusted according to the threat level and resource importance.

[0076] The algorithm uses a strategy-as-code approach to automate the deployment and updating of isolation rules, supports elastic scaling of slice resources, ensures the confidentiality and integrity of data transmission between slices through encrypted communication and data protection mechanisms, adjusts the isolation strategy in real time based on threat intelligence and security event triggers, and realizes the transformation from passive protection to active defense based on a quantitative model.

[0077] Algorithm complexity analysis: The zero-trust isolation algorithm has a time complexity of O(n×m+k²), where n is the number of entities, m is the number of attributes, and k is the number of policy rules; the space complexity is O(n×m).

[0078] The inference complexity of a deep reinforcement learning model is O(|S||A|+L²d), where |S| is the size of the state space, |A| is the size of the action space, L is the sequence length, and d is the feature dimension.

[0079] Compared to traditional methods, while the complexity of this algorithm is increased (ACL algorithm has O(n)), in actual tests, the average response time for processing 1000 entities is 9.4ms, approximately 18 times faster than the 0.5ms of traditional firewalls, but with a significant improvement in security. In resource-constrained edge computing scenarios, the complexity can be reduced by 60% by lowering feature dimensions and simplifying the collaboration mechanism, while still maintaining over 90% detection accuracy.

[0080] Step 3: Design of Deep Reinforcement Learning Threat Detection Model: The deep reinforcement learning threat detection model adopts a multi-agent architecture, with each agent responsible for threat monitoring of a specific slice or network region. The deep reinforcement learning threat detection framework adopts a multi-agent collaborative architecture, in which multiple agents achieve distributed threat detection and joint optimization through information sharing and collaborative decision-making. It covers multi-agent collaboration mechanisms, state space construction, reward function design, and policy optimization processes. This framework leverages the information interaction and collaborative decision-making between agents to achieve intelligent detection and response to complex threats in virtualized network slice environments.

[0081] Step 4: Adaptive Linkage Mechanism between Isolation Strategy and Detection Results: The adaptive linkage mechanism establishes a closed-loop feedback control system between threat detection results and isolation strategies. The threat level assessment module calculates the threat severity based on the detection results, threat type, and scope of impact, triggering the corresponding level of isolation response. The strategy decision engine uses a combination of fuzzy logic and expert systems to automatically select the optimal combination of isolation strategies. The linkage mechanism supports progressive isolation upgrades, from traffic restrictions and privilege downgrades to multi-level responses of complete isolation. The real-time strategy optimization module continuously adjusts linkage parameters and decision thresholds by analyzing historical data and current network status. The mechanism also includes fault recovery and rollback strategies to ensure rapid restoration of normal service in the event of false alarms. The linkage effect evaluation module monitors the effectiveness of isolation measures and provides feedback information for strategy optimization.

[0082] In step one, physical layer threats originate from hardware devices and infrastructure, virtualization layer threats exploit vulnerabilities in virtualization technology, slicing layer threats target the isolation mechanism between slices, and application layer threats attack upper-layer business services. Threat propagation paths include vertical attacks (cross-layer penetration) and lateral diffusion (same-layer spread), forming a complete threat situation analysis framework.

[0083] In step two, the algorithm is formally defined as follows:

[0084] Definition 1 (Zero Trust State Space): S = {s_auth, s_authz, s_monitor, s_isolate}, representing authentication, authorization, monitoring, and isolation states respectively;

[0085] Definition 2 (Action Space): A = {allow, deny, quarantine(t), monitor(level)}, where t is the quarantine time and level is the monitoring level;

[0086] Definition 3 (State transition function): δ: S×A×T→S, satisfies the Markov property.

[0087] Algorithm 1: Zero-trust dynamic isolation algorithm

[0088] Input: Entity set E = {e1, ..., e} n}, attribute matrix M, policy rule set P, trust threshold τ, decision threshold θ

[0089] Output: Isolation decision set D

[0090] Algorithm flow:

[0091] 1: for each eᵢ ∈ E do T hist (eᵢ) ← 0.5; D ←θ end for;

[0092] 2: for each eᵢ ∈ E do;

[0093] 3: T hist (eᵢ)← 0.7·T hist (eᵢ)+ 0.3·Score(f hist );

[0094] 4: T behav (eᵢ, t) ← AnomalyDetect(eᵢ, t);

[0095] 5: T(eᵢ, t) ← α·T hist (eᵢ) + β·T behav (eᵢ, t) + γ·T ctx (eᵢ, t) / / Formula (1);

[0096] 6: end for;

[0097] 7: for each request(eᵢ, rⱼ) do;

[0098] 8: if T(eᵢ, t) < τ then aᵢ ← deny;

[0099] 9: else calculate P access ← σ(∑wᵢAᵢ - θ) / / Formula (2);

[0100] 10: if P access ≥ θ then aᵢ ← allow;

[0101] 11: else if P access ≥ θ - 0.1 then aᵢ ← monitor;

[0102] 12: else calculate Lisolation And perform differential piecewise isolation aᵢ ← quarantine / / formula (3);

[0103] 13: D←D∪ {(eᵢ, aᵢ)};

[0104] 14: end for;

[0105] 15: Update T hist (eᵢ) refreshes P when a security event is triggered;

[0106] 16: return D;

[0107] In step two, the formula for calculating entity trust level is:

[0108] (1)

[0109] Where T(e,t) represents the trust level of entity e at time t, T hist (e) represents the basic trust value based on historical behavior, T behav (e,t) represents the trust score for the current behavior, where T is the trust score for the current behavior. context (e,t) is the environmental context trust value, where α, β, and γ are weight coefficients and satisfy α+β+γ=1;

[0110] The physical meaning of this formula lies in calculating the comprehensive trust level of entities through multi-dimensional fusion, avoiding the limitations of single-dimensional evaluation. In terms of parameter settings, the weight coefficients are dynamically adjusted according to the network environment and security requirements. In high-security scenarios, α is set to 0.3, β is set to 0.5, and γ is set to 0.2, highlighting the importance of the current behavior. In stable network environments, α is set to 0.5, β is set to 0.3, and γ is set to 0.2, relying more on historical trust foundations.

[0111] Parameter optimization employed a grid search method, conducting 49 combinations of tests for α, β, and γ ∈ [0.1, 0.7] under the constraint α + β + γ = 1. The optimal parameters were determined using 5-fold cross-validation with the objective function of overall safety performance F = 0.4 × accuracy + 0.3 × (1 - false alarm rate) + 0.3 × response speed. Controlled-variable experiments showed that the F-value reached its highest value of 0.892 when α = 0.3, β = 0.5, and γ = 0.2, representing an 8.3% improvement compared to the second-best combination. Sensitivity analysis indicated that system performance fluctuations were less than 5% when parameters varied within ±20%, demonstrating the robustness of the parameter settings. The decision threshold θ was determined through ROC curve analysis. For the high-safety slice, θ = 0.8 resulted in a false positive rate below 2%, while for the general slice, θ = 0.6 achieved the optimal balance between detection rate and efficiency.

[0112] The probability of access permission is calculated as follows:

[0113] (2)

[0114] Among them, P access For the probability of access permission, A i w is the standardized value of the i-th attribute. i Let θ be the corresponding weight, θ be the decision threshold, and n be the total number of attributes. When P access Access requests are permitted when the threshold is exceeded.

[0115] This formula maps multi-attribute scores to the [0,1] interval using the Sigmoid function, achieving probabilistic decision-making for access control. Physically, this means that the greater the attribute weight and the higher the attribute value, the greater the probability of access permission. Parameter settings include: the decision threshold θ is set according to the security level, with 0.8 for high-security slices and 0.6 for general slices; attribute weights wi are assigned based on attribute importance, with 0.4 for identity authentication, 0.3 for behavioral analysis, and 0.3 for environmental context.

[0116] The formula for calculating the isolation level is:

[0117] (3)

[0118] Among them, L isolation For isolation level, R threat For threat risk value, C resource L is the resource importance coefficient. max This is the highest level of isolation.

[0119] This formula ensures that the isolation strength has a logarithmic relationship with the risk level, avoiding excessive isolation that could negatively impact system performance. The physical significance of the logarithmic function design lies in achieving non-linear growth in isolation strength, avoiding over-isolation at low risk and under-isolation at high risk. The parameter settings are based on: the threat risk value R_threat is quantified through threat level, ranging from [0,10], where 0-3 represents low risk, 4-6 medium risk, and 7-10 high risk; and the resource importance coefficient C... resource The isolation level is set according to the business criticality: core business is set to 1.5, general business is set to 1.0, and test business is set to 0.5. The maximum isolation level is L. max Set to 5, which corresponds to the fully isolated state.

[0120] In step three, the model design includes the following core components:

[0121] (1) State space construction encodes network traffic characteristics, system behavior indicators, resource usage and security event information into high-dimensional state vectors, and uses attention mechanism and graph neural network to extract spatiotemporal correlation features and construct dependency graphs between slices;

[0122] The multi-head attention mechanism uses 8 attention heads, each with a dimension of 64, and calculates temporal feature weights through a Query-Key-Value mechanism.

[0123] (4)

[0124] Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k With the key vector dimension, this mechanism effectively captures the temporal dependence of threat behavior. The graph neural network adopts a 3-layer GCN architecture with hidden layer dimensions of [256, 128, 64]. The adjacency matrix is ​​constructed based on the communication frequency between slices.

[0125] (5)

[0126] Among them, A ij f represents the connection weight between nodes i and j. i and f j Let i and j be the feature vectors of nodes i and j, respectively, σ be the bandwidth parameter, and the message passing mechanism be:

[0127] (6);

[0128] Among them, h v (l+1) Let W be the hidden state of node v at level l+1. l Here, is the weight matrix of the l-th layer, AGG is the aggregation function, N(v) is the set of neighbors of node v, σ is the activation function, and spatiotemporal feature fusion is achieved by combining the temporal attention output and the spatial graph convolution result through an adaptive weight mechanism:

[0129] (7)

[0130] Among them, h final For the final fusion feature, h temporal For temporal attention output, h spatial The result is the spatial graph convolution, where λ is the adaptive weight parameter.

[0131] State representations, which integrate statistical features, sequence patterns, and graph structure information, can capture complex attack behavior patterns.

[0132] (2) In terms of reward function design, multiple objectives such as detection accuracy, false alarm rate, detection latency and resource consumption are comprehensively considered. The positive reward mechanism is used to encourage correct threat identification and timely response, while the negative reward mechanism is used to punish false alarms and missed alarms. The dynamic reward adjustment strategy allocates weights based on the severity of the threat and the impact on business, so as to balance detection performance and system efficiency.

[0133] (3) The state space adopts a 288-dimensional feature vector: 96-dimensional network traffic features (packet size, protocol distribution, connection mode), 96-dimensional system behavior features (CPU, memory, I / O utilization), 48-dimensional resource usage features (bandwidth, storage usage), and 48-dimensional security event features (alarm type, severity).

[0134] (4) Multi-agent cooperation mechanism: The number of agents is based on It is determined that each active slice deploys a local detection agent, with an additional global coordinator. The communication topology adopts a hybrid structure of star and local mesh. The local agent is connected to the coordinator in a star topology, and mesh sidechains are established between adjacent slices for lateral threat perception. The weighted voting weights are dynamically updated based on the historical accuracy of each agent over the past 10 windows.

[0135] (8)

[0136] Coordinator's weighted score If a threat is identified and isolation linkage is triggered, when S∈[0.4,0.6), the coordinator retrieves the complete features and makes a second determination through the PPO policy network, outputs the final confidence score, and broadcasts the result to each local agent to complete the closed loop.

[0137] (5) The reward function weights are allocated based on business priorities: detection accuracy w1 = 0.4 (core objective), false alarm penalty w2 = -0.3 (reduce operation and maintenance costs), response delay penalty w3 = -0.2 (ensure real-time performance), and resource consumption penalty w4 = -0.1 (system efficiency).

[0138] (6) Training hyperparameter settings: PPO algorithm, learning rate 3e-4, batch size 256, experience replay buffer 10000, ε-greedy policy decayed from 0.9 to 0.1, gradient clipping threshold 0.5.

[0139] In step three, the framework comprises four core layers: a state perception layer, a feature extraction layer, a decision layer, and an execution layer. The state perception layer is responsible for the real-time collection of network traffic and system behavior data. The feature extraction layer extracts spatiotemporal correlation features through attention mechanisms and graph neural networks. The decision layer performs threat judgment based on a multi-agent reinforcement learning algorithm. The execution layer outputs detection results and response strategies.

[0140] System Implementation and Experimental Verification

[0141] Prototype system design and experimental environment construction:

[0142] The prototype system is built based on the OpenStack cloud platform and adopts a distributed microservices architecture to implement security isolation and threat detection functional modules. The core components of the system include a slice management controller, a zero-trust security gateway, a threat detection engine, and a联动 response module. The slice management controller is responsible for the creation, configuration, and lifecycle management of virtual network slices, integrating an SDN controller and an NFV orchestrator. The zero-trust security gateway is deployed at the slice boundary to implement dynamic access control and traffic monitoring. The threat detection engine is developed based on the TensorFlow deep learning framework and supports multi-agent reinforcement learning algorithms. The experimental environment consists of 20 physical servers configured with Intel Xeon processors and 64GB of memory, interconnected via 10Gbps Ethernet. The virtualization platform creates 100 virtual slice instances to simulate different business scenarios and user behaviors. The network topology adopts a three-layer architecture, including a core layer, an aggregation layer, and an access layer, with corresponding security protection devices deployed at each layer. The experimental dataset combines real network traffic and synthetic attack samples, covering normal business traffic and 15 typical network attack types.

[0143] Experiment on validating the effectiveness of the security isolation mechanism:

[0144] The experiment on validating the effectiveness of security isolation evaluates the protection ability of the isolation mechanism by designing various attack scenarios. The experiment sets up typical attack modes such as unauthorized access between slices, resource competition attacks, and side-channel information leakage, and measures the blocking effect and response time of the isolation mechanism. The testing method combines black-box testing and white-box testing to evaluate the system security from the perspectives of both attackers and defenders. The experimental results show that the zero-trust dynamic isolation mechanism can maintain effective isolation in the face of different-intensity attacks, and the attack success rate is reduced to less than 2.3%. The average response time for adjusting the isolation policy is 150 milliseconds, meeting the real-time protection requirements. The traffic isolation rate between slices reaches 99.7%, and the accuracy rate of resource access control is 98.5%. The experiment also verifies the stability of the isolation mechanism in high-concurrency scenarios. When processing 1000 access requests simultaneously, the system performance degradation is controlled within 8%. The micro-segmentation technology effectively reduces the scope of attack impact, controls potential losses within a single security domain, and avoids the occurrence of cascading security events;

[0145] Testing the accuracy and real-time performance of threat detection:

[0146] Threat detection performance testing evaluates the effectiveness of detection algorithms by constructing a comprehensive test set containing both known attacks and zero-day attacks. The test dataset is based on CICIDS2017 and NSL-KDD, combined with synthetic attack samples generated by GANs. These synthetic attack samples cover scenarios such as zero-day exploitation and multi-stage penetration, and are mixed with real samples at a 3:1 ratio. The test dataset contains a total of 5 million records, with malicious traffic accounting for 15%, covering various threat types including DDoS attacks, malware propagation, and data theft.

[0147] Table 2 Performance Comparison of Threat Detection Methods

[0148] method accuracy False alarm rate Detection delay F1-Score Snort IDS 76.3% 8.7% 15.2ms 0.709 SVM + Feature Engineering 82.7% 6.4% 8.9ms 0.786 CNN-LSTM 89.4% 4.2% 12.4ms 0.859 GraphSAINT 92.1% 3.8% 18.6ms 0.900 This article's method 96.8% 2.1% 3.2ms 0.948

[0149] As shown in the comparative analysis in Table 2, our proposed method significantly outperforms the baseline method in all three key metrics: accuracy, false positive rate, and detection latency. It achieves a 20.5 percentage point improvement over traditional IDS, a 4.7 percentage point improvement over the best deep learning baseline GraphSAINT, and a 83% reduction in detection latency. The multi-agent collaborative mechanism improves accuracy by 12.3% compared to the single-agent DQN method, demonstrating the effectiveness of the collaborative strategy.

[0150] In addition, Table 3 further compares with recent integrated solutions, including the ZT-DRL solution combining zero trust and deep learning (2024) and the heterogeneous graph neural network detection solution (2024). The detection accuracy of our method is improved by 3.2 and 2.5 percentage points respectively compared with these two solutions, and the detection latency is reduced by 78% and 81%, which verifies the effectiveness of the multi-agent collaboration and intelligent linkage mechanism compared with the latest similar solutions.

[0151] Table 3. Comparison with advanced solutions

[0152] Method type Specific Plan accuracy False alarm rate Detection delay Deep learning threat detection Heterogeneous Graph Neural Networks 94.3% 3.2% 16.8ms CNN-BiLSTM 91.7% 4.6% 14.2ms Advanced Zero Trust Solution Google BeyondCorp 88.5% 5.1% 25.6ms Microsoft Zero Trust 86.2% 6.3% 22.4ms This article's solution Zero Trust + DRL + Collaboration 96.8% 2.1% 3.2ms

[0153] As shown in Table 3, the extended comparative experiments show that the proposed scheme improves the detection accuracy by 2.5 percentage points and reduces the detection latency by 81% compared with the heterogeneous graph neural network. Compared with the mainstream zero-trust scheme, it has significant advantages in accuracy and real-time performance, verifying the effectiveness of the multi-agent collaboration and intelligent linkage mechanism.

[0154] The deep reinforcement learning detection model reached convergence after 10,000 iterations of training, achieving a detection accuracy of 96.8% on the test set, a false positive rate of 2.1%, and a false negative rate of 1.1%. Real-time performance test results show that the average detection latency of a single data packet is 3.2 milliseconds. This result was achieved through the following engineering optimizations: GCN inference is accelerated using sparse matrices, the multi-head attention mechanism is deployed on GPU parallel computing (NVIDIA A100, FP16 precision), multi-agent communication only transmits L1-level binary results (latency < 0.1ms), a single forward inference of the three-layer GCN and 8-head attention takes approximately 2.1ms, and the coordinator decision and communication overhead is approximately 1.1ms, totaling 3.2ms, which meets the real-time detection requirements of high-speed network environments. The model's generalization ability to novel attacks was verified through zero-day attack detection experiments, maintaining a detection rate of over 85% even for unseen attack patterns. To comprehensively evaluate detection performance, multi-dimensional simulation experiments were designed, and the results are as follows: Figure 3 As shown.

[0155] Figure 3 The study includes four key performance dimensions: (a) the trend of detection accuracy with network load, verifying the stability of the algorithm under high load; (b) a comparison of false alarm rate control effects, demonstrating the algorithm's accuracy advantage; (c) a temporal stability analysis of detection latency, proving real-time detection capability; and (d) the variation of system resource consumption with load, evaluating the algorithm's efficiency and scalability. Experimental results verify the comprehensive advantages of the proposed algorithm in terms of accuracy, real-time performance, and resource efficiency. The multi-agent cooperation mechanism significantly improves the detection performance, increasing accuracy by 12% compared to the single-agent method. The experiment also tested detection performance under different network load conditions. When network traffic increased to 10Gbps, the detection accuracy remained above 94%, demonstrating the algorithm's robustness and scalability.

[0156] Experimental analysis of the impact of synergistic mechanism performance:

[0157] Ablation Experiment Analysis: To verify the independent contribution of each module, four ablation control groups were designed:

[0158] ① Zero Trust Isolation Module Only (ZT-only) ② DRL Detection Module Only (DRL-only) ③ Zero Trust + DRL but No Linkage Mechanism (ZT+DRL) ④ Complete Solution (ZT+DRL+Linkage);

[0159] Experimental results show that ZT-only improves detection accuracy by 8.1% compared to the baseline, DRL-only by 12.3%, and ZT+DRL by 21.4%. The complete solution achieves the best accuracy of 96.8%, which is 4.2 percentage points higher than ZT+DRL. This demonstrates that the linkage mechanism brings additional performance gains on the basis of zero trust and DRL collaboration. All three modules make irreplaceable independent contributions to the final performance.

[0160] System resource consumption analysis: The introduction of the collaborative mechanism has a certain impact on the overall system performance, mainly reflected in three aspects: CPU utilization, memory usage, and network bandwidth consumption. Experimental monitoring shows that when the collaborative mechanism is running, the average CPU utilization increases by 15%, and the peak utilization increases by 22%. Memory usage increases by an additional 320MB, mainly used for storing threat intelligence, policy rules, and historical status information. Network control traffic overhead accounts for 3.2% of the total bandwidth, which is within an acceptable range.

[0161] Table 4. Performance Cost Comparison of Different Linkage Strategies

[0162] Linkage Strategy CPU growth rate Memory growth rate Average response delay System throughput decreased Basic threshold linkage 8% 12% 25ms 5% Rule Engine Linkage 15% 18% 45ms 12% This article features intelligent linkage. 15% 20% 52ms 8% Deep Collaboration 25% 32% 89ms 18%

[0163] As shown in Table 4, the intelligent linkage strategy presented in this paper is comparable to the rule engine linkage in terms of CPU and memory utilization, but the system throughput decreases by only 8%, which is 10 percentage points lower than the deep linkage strategy. Although the average response latency is slightly higher than the basic solution, it has a significant advantage in the accuracy of complex threat detection.

[0164] Performance balancing control mechanism: When the CPU utilization exceeds 85%, the system automatically reduces the detection frequency to 70%, and the accuracy drops by only 2.3%; when the network load exceeds 8Gbps, a lightweight model is used for replacement, which increases the processing latency by 15% but maintains 94% detection accuracy.

[0165] The performance degradation under high load is mainly due to: (1) the communication overhead of multi-agents increases exponentially with the load; (2) the secondary complexity of the attention mechanism becomes a bottleneck when dealing with large-scale data; and (3) the conflict between the real-time computing requirements of collaborative decision-making and resource competition. By using dynamic dimensionality reduction and hierarchical processing strategies, the performance loss under high load scenarios can be controlled within 15%.

[0166] Response Latency and Throughput Testing: The impact of the collaborative mechanism on system response latency and business throughput was evaluated through stress testing. Under normal load conditions, the average system response latency increased by 8 milliseconds, mainly due to threat detection and policy decision-making processes. Under high load scenarios, the latency increase was controlled within 15 milliseconds, and the business throughput decreased by 6.5%. The adaptive adjustment capability of the collaborative mechanism effectively mitigated the performance impact, automatically increasing detection accuracy when high-risk threats are detected and reducing monitoring intensity to optimize performance when the security situation is good.

[0167] Comparison of protection effectiveness under different attack scenarios:

[0168] To comprehensively evaluate the protective capabilities of the proposed collaborative mechanism, comparative experiments covering a variety of typical network attacks were designed. Three representative baseline methods were selected for the comparative experiments: traditional firewalls representing rule-based boundary protection schemes, Snort IDS representing signature-based intrusion detection schemes, and the network slicing security deployment method based on an improved Bayesian network model proposed by the BN slicing scheme. The experiments were conducted under the same network environment and attack intensity to ensure the objectivity and comparability of the results.

[0169] Table 5 evaluates the performance of different protection schemes under various attack scenarios from three dimensions: detection rate, false alarm rate, and response time. The experimental results show that the proposed collaborative mechanism has significant advantages in the face of complex multi-stage attacks and zero-day attacks.

[0170] Table 5 Comparison of protection effectiveness under different attack scenarios

[0171] Attack type Traditional firewalls Snort IDS BN slicing scheme This article's solution Attack type Traditional firewalls Snort IDS BN slicing scheme This article's solution Detection rate (%) Detection rate (%) Detection rate (%) Detection rate (%) DDoS attack 89.2 / 5.8 / 250 85.4 / 8.2 / 200 92.1 / 4.3 / 180 96.8 / 2.1 / 150 Malicious code injection 75.6 / 12.3 / 300 82.3 / 9.5 / 280 88.9 / 6.1 / 220 94.2 / 3.2 / 170 Side-channel attack 45.2 / 15.6 / - 52.8 / 13.4 / - 68.7 / 8.9 / 350 85.3 / 4.1 / 200 Resource contention attack 52.1 / 9.8 / 400 58.7 / 11.2 / 380 71.5 / 7.3 / 290 89.6 / 3.8 / 180 Zero-day attack 38.9 / 18.4 / - 43.2 / 16.7 / - 65.2 / 10.5 / - 82.7 / 5.4 / 250

[0172] Note: The data format is "detection rate (%) / false alarm rate (%) / response time (ms)", "-" indicates that it cannot be detected effectively.

[0173] The response time in Table 5 is the end-to-end protection response time, covering the entire process of threat detection, strategy decision-making and isolation execution. The detection latency in Tables 2 and 3 only refers to the time consumed in the model inference stage (at the level of 3.2ms). The two have different statistical standards and cannot be directly compared.

[0174] Experimental data show that the proposed collaborative mechanism achieved the best overall performance in all test scenarios. Compared with the optimal baseline solution, the detection rate was improved by an average of 15.3%, the false alarm rate was reduced by 52.8%, and the response time was reduced by 28.6%. In side-channel attacks and zero-day attacks that traditional protection solutions struggle to handle, the collaborative mechanism demonstrated significant technical advantages, with detection rates improved by 16.6% and 17.5%, respectively, providing more reliable security for virtualized network slicing environments.

[0175] Results analysis and discussion:

[0176] Quantitative analysis of experimental results:

[0177] Through statistical analysis of large-scale experimental data, this solution has achieved significant performance improvement in the security protection of virtualized network slicing. The experiment adopted multiple sets of control tests, collected more than 5 million network traffic data and 10,000 security event records to ensure the statistical significance of the results. The data analysis covered key dimensions such as system convergence, stability, scalability and cost-effectiveness.

[0178] Table 6 Statistical Analysis of System Performance Indicators

[0179] Performance dimension Indicator Name Test conditions Measured values Standard deviation Confidence interval (95%) Algorithm convergence Number of training convergence rounds 10 independent training sessions 8,750 wheels 245 [8,505, 8,995] Stability detection Accuracy Variance 24-hour continuous monitoring 0.0032 0.0008 [0.0024, 0.0040] System scalability Slice Count Threshold Incremental load test 500 slices - - Response consistency Delay standard deviation 1000 response tests 12.5ms 3.2 [9.3, 15.7] resource utilization rate Peak CPU usage High load scenarios 78.5% 5.1 [73.4, 83.6] Memory stability Memory leak rate 72-hour stress test 0.08MB / h 0.02 [0.06, 0.10]

[0180] As shown in Table 6, the statistical analysis results demonstrate that the deep reinforcement learning model can converge stably under different initial conditions, with an average convergence epoch of 8,750 epochs, indicating good convergence. The system maintains stable detection performance during long-term operation, with an accuracy variance of only 0.0032, proving the robustness of the algorithm. Scalability tests show that the system can support up to 500 concurrent network slices, meeting the needs of large-scale deployment. The standard deviation of response latency is controlled within 12.5 milliseconds, reflecting good service quality consistency. Resource consumption monitoring shows that the system's peak CPU usage is 78.5% under high load, with an extremely low memory leak rate, demonstrating the ability to operate stably for a long time.

[0181] Technical advantages and performance improvements of the solution:

[0182] The proposed collaborative security framework offers multiple technical advantages over existing solutions, primarily in terms of adaptability, accuracy, and real-time performance. The zero-trust dynamic isolation algorithm, through quantitative trust assessment and a multi-attribute decision model, achieves fine-grained access control and adaptive security policy adjustments. The deep reinforcement learning threat detection model possesses powerful feature learning capabilities and environmental adaptability, effectively addressing unknown attacks and dynamic threat environments. The collaborative linkage mechanism establishes a closed-loop feedback control between detection and protection, realizing intelligent coordination between threat perception and security response. Key technical innovations include a multi-dimensional trust assessment mechanism, a logarithmic function-based isolation level calculation method, a multi-agent collaborative detection architecture, and an adaptive policy optimization algorithm. Performance improvements are mainly reflected in significantly increased detection accuracy, a significantly reduced false alarm rate, a significantly faster response speed, and effective assurance of system scalability. The solution also boasts excellent backward compatibility, enabling seamless integration with existing network security devices and management systems, reducing deployment costs and technical risks.

[0183] Discussion on the feasibility of practical application:

[0184] From the perspectives of technological maturity and deployment complexity, this solution demonstrates strong feasibility for practical application. The core technologies employed, such as SDN, NFV, and deep learning, are widely used in industry, providing a solid technical foundation for implementation. Experimental verification shows that the solution meets the deployment requirements of production environments in terms of resource consumption, performance overhead, and system stability.

[0185] Table 7. Cost and Benefit Analysis of Solution Deployment

[0186] Deployment elements Initial investment Operating costs Expected returns hardware devices Infrastructure upgrade 500,000 yuan Annual maintenance cost: 80,000 yuan Reduce security incident losses by 2 million yuan per year. Software System Development and deployment costs: 800,000 yuan License fee: 120,000 yuan / year Improved operational efficiency saves 600,000 yuan per year Personnel training Training fee: 150,000 yuan Professional salary 250,000 yuan / year Reduce labor costs by 400,000 yuan / year System Integration Integration testing cost: 350,000 yuan Technical support fee: 100,000 yuan / year To avoid business interruption losses of 1.5 million yuan per year total 1.8 million yuan 550,000 yuan / year 4.5 million yuan / year

[0187] As shown in Table 7's economic benefit analysis, although the initial deployment of the solution requires an investment of 1.8 million yuan and annual operating costs of 550,000 yuan, the expected annual return can reach 4.5 million yuan by reducing security incident losses, improving operational efficiency, and lowering labor costs. The investment payback period is approximately 6 months, demonstrating good economic feasibility. Key factors to consider during deployment include compatibility modifications to existing systems, training needs for technical personnel, phased implementation strategies, and risk control measures. It is recommended to adopt a pilot-first, gradually expanding deployment approach to ensure smooth implementation and continuous optimization of the solution. The above cost and benefit data are calculated proportionally based on resource consumption monitoring results from a 20-node experimental environment (Table 6) and estimated with reference to the industry average for similar security system deployments. Specific figures may vary depending on the actual deployment scale and operating environment and are for reference only.

[0188] This application addresses the key challenges of security isolation and threat detection in virtualized network slicing environments by proposing a collaborative security framework that integrates zero-trust architecture, deep reinforcement learning, and intelligent linkage mechanisms.

[0189] The core contributions are summarized as follows:

[0190] (1) In terms of technological innovation, a zero-trust dynamic isolation algorithm based on quantitative trust assessment was constructed, which broke through the static limitations of traditional boundary protection; a multi-agent deep reinforcement learning threat detection model was designed to improve the adaptive identification capability of unknown threats; and a closed-loop linkage mechanism between detection results and isolation strategies was established to realize intelligent collaboration of security protection.

[0191] (2) In terms of performance improvement, experimental verification shows that the threat detection accuracy reaches 96.8%, the false alarm rate is controlled at 2.1%, and the attack success rate is reduced to 2.3%. Compared with the existing solutions, there are significant improvements in detection accuracy, response speed and protection effect.

[0192] (3) In terms of theoretical contributions, a four-layer threat model for virtualized network slicing was established, and a quantitative mapping relationship between trust metric and isolation strength was proposed, providing a theoretical basis for research in related fields.

[0193] Future research will focus on the following specific directions: (1) Research on cross-domain slicing security collaboration mechanism, specifically including unified coordination of slicing security strategies in multi-operator environments, threat intelligence sharing mechanism for cross-regional network slices, and trust transfer and verification methods in heterogeneous network environments; (2) Performance optimization strategies in large-scale deployment scenarios, involving lightweight detection algorithms under massive slicing concurrency, load balancing mechanisms for distributed threat detection, and resource-constrained optimization methods in edge computing environments; (3) Security architecture evolution for 6G networks, exploring security protection mechanisms for integrated air-space-ground network slices, intelligent reflector-assisted secure communication methods, and quantum security-enhanced slicing isolation technology.

[0194] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or basic characteristics. Therefore, the embodiments should be considered exemplary and non-limiting in all respects. The scope of the invention is defined by the appended claims rather than the foregoing description. Therefore, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0195] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for secure isolation and threat detection of virtualized network slices, characterized in that, The specific steps of this security isolation and threat detection method are as follows: Step 1: Construction of Virtualized Network Slice Threat Model: The virtualized network slice threat model adopts a layered modeling method, dividing threats into four dimensions: physical layer, virtualization layer, slice layer and application layer. The threat model adopts a layered architecture design, from bottom to top: physical layer, virtualization layer, slice layer and application layer. Step 2: Zero-trust-based dynamic security isolation algorithm: Adopting the core principle of never trusting and always verifying, a multi-layered authentication and authorization mechanism is constructed. The algorithm achieves secure isolation between slices through quantitative trust assessment and dynamic policy adjustment; it calculates the trust level of entities by comprehensively considering multiple dimensions; the micro-segment isolation strength is dynamically adjusted according to the threat level and resource importance. The algorithm uses a strategy-as-code approach to automate the deployment and updating of isolation rules, supports elastic scaling of slice resources, ensures the confidentiality and integrity of data transmission between slices through encrypted communication and data protection mechanisms, adjusts the isolation strategy in real time based on threat intelligence and security event triggers, and realizes the transformation from passive protection to active defense based on a quantitative model. Step 3: Design of Deep Reinforcement Learning Threat Detection Model: The deep reinforcement learning threat detection model adopts a multi-agent architecture, with each agent responsible for threat monitoring of a specific slice or network region. The deep reinforcement learning threat detection framework adopts a multi-agent collaborative architecture, in which multiple agents achieve distributed threat detection and joint optimization through information sharing and collaborative decision-making. It covers multi-agent collaboration mechanisms, state space construction, reward function design, and policy optimization processes. This framework leverages the information interaction and collaborative decision-making between agents to achieve intelligent detection and response to complex threats in virtualized network slice environments. Step 4: Adaptive Linkage Mechanism between Isolation Strategy and Detection Results: The adaptive linkage mechanism establishes a closed-loop feedback control system between threat detection results and isolation strategies. The threat level assessment module calculates the threat severity based on the detection results, threat type, and scope of impact, and triggers the corresponding level of isolation response. The strategy decision engine uses a combination of fuzzy logic and expert systems to automatically select the optimal combination of isolation strategies. The linkage mechanism supports progressive isolation upgrades, with multi-level responses ranging from traffic restrictions and permission downgrades to complete isolation. The real-time strategy optimization module continuously adjusts linkage parameters and decision thresholds by analyzing historical data and current network status. The mechanism also includes fault recovery and rollback strategies to ensure rapid restoration of normal service in the event of false alarms. The linkage effect evaluation module monitors the effectiveness of isolation measures and provides feedback information for strategy optimization.

2. The method for secure isolation and threat detection of virtualized network slices according to claim 1, characterized in that: In step one, physical layer threats originate from hardware devices and infrastructure, virtualization layer threats exploit vulnerabilities in virtualization technology, slicing layer threats target the isolation mechanism between slices, and application layer threats attack upper-layer business services. Threat propagation paths include vertical attacks (cross-layer penetration) and lateral diffusion (same-layer spread), forming a complete threat situation analysis framework.

3. The method for secure isolation and threat detection of virtualized network slices according to claim 1, characterized in that: In step two, the algorithm is formally defined as follows: Definition 1 (Zero Trust State Space): S = {s_auth, s_authz, s_monitor, s_isolate}, representing authentication, authorization, monitoring, and isolation states respectively; Definition 2 (Action Space): A = {allow, deny, quarantine(t), monitor(level)}, where t is the quarantine time and level is the monitoring level; Definition 3 (State transition function): δ: S×A×T→S, satisfies the Markov property.

4. The method for secure isolation and threat detection of virtualized network slices according to claim 1, characterized in that: In step two, the formula for calculating entity trust level is: (1) Where T(e,t) represents the trust level of entity e at time t, T hist (e) represents the basic trust value based on historical behavior, T behav (e,t) represents the trust score for the current behavior, where T is the trust score for the current behavior. context (e,t) is the environmental context trust value, where α, β, and γ are weight coefficients and satisfy α+β+γ=1; The probability of access permission is calculated as follows: (2) Among them, P access For the probability of access permission, A i w is the standardized value of the i-th attribute. i Let θ be the corresponding weight, θ be the decision threshold, and n be the total number of attributes. When P access Access requests are permitted when the threshold is exceeded. The formula for calculating the isolation level is: (3) Among them, L isolation For isolation level, R threat For threat risk value, C resource L is the resource importance coefficient. max This is the highest level of isolation.

5. The method for secure isolation and threat detection of virtualized network slices according to claim 1, characterized in that: In step three, the model design includes the following core components: (1) State space construction encodes network traffic characteristics, system behavior indicators, resource usage and security event information into high-dimensional state vectors, and uses attention mechanism and graph neural network to extract spatiotemporal correlation features and construct dependency graphs between slices; The multi-head attention mechanism uses 8 attention heads, each with a dimension of 64, and calculates temporal feature weights through a Query-Key-Value mechanism. (4) Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k With the key vector dimension as the key, this mechanism effectively captures the temporal dependence of threat behavior. The graph neural network adopts a 3-layer GCN architecture with hidden layer dimensions of [256, 128, 64]. The adjacency matrix is ​​constructed based on the communication frequency between slices. (5) Among them, A ij f is the connection weight between nodes i and j. i and f j Let i and j be the feature vectors of nodes i and j, respectively, σ be the bandwidth parameter, and the message passing mechanism be: (6); Among them, h v (l+1) Let W be the hidden state of node v at level l+1. l Let be the weight matrix of the l-th layer, AGG be the aggregation function, N(v) be the set of neighbors of node v, and σ be the activation function. Spatiotemporal feature fusion is achieved by combining the temporal attention output and the spatial graph convolution result through an adaptive weight mechanism. (7) Among them, h final For the final fusion feature, h temporal For temporal attention output, h spatial The result is the spatial graph convolution, where λ is the adaptive weight parameter. State representations, which integrate statistical features, sequence patterns, and graph structure information, can capture complex attack behavior patterns. (2) In terms of reward function design, multiple objectives such as detection accuracy, false alarm rate, detection latency and resource consumption are comprehensively considered. The positive reward mechanism is used to encourage correct threat identification and timely response, while the negative reward mechanism is used to punish false alarms and missed alarms. The dynamic reward adjustment strategy allocates weights based on the severity of the threat and the impact on business, so as to balance detection performance and system efficiency. (3) The state space adopts a 288-dimensional feature vector: 96-dimensional network traffic features (packet size, protocol distribution, connection mode), 96-dimensional system behavior features (CPU, memory, I / O utilization), 48-dimensional resource usage features (bandwidth, storage usage), and 48-dimensional security event features (alarm type, severity). (4) Multi-agent cooperation mechanism: The number of agents is based on It is determined that each active slice deploys a local detection agent, with an additional global coordinator. The communication topology adopts a hybrid structure of star and local mesh. The local agent is connected to the coordinator in a star topology, and mesh sidechains are established between adjacent slices for lateral threat perception. The weighted voting weights are dynamically updated based on the historical accuracy of each agent over the past 10 windows. (8) Coordinator's weighted score If a threat is identified and isolation linkage is triggered, when S∈[0.4,0.6), the coordinator retrieves the complete features and makes a second determination through the PPO policy network, outputs the final confidence score, and broadcasts the result to each local agent to complete the closed loop. (5) The reward function weights are allocated based on business priorities: detection accuracy w1 = 0.4 (core objective), false alarm penalty w2 = -0.3 (reduce operation and maintenance costs), response delay penalty w3 = -0.2 (ensure real-time performance), and resource consumption penalty w4 = -0.1 (system efficiency). (6) Training hyperparameter settings: PPO algorithm, learning rate 3e-4, batch size 256, experience replay buffer 10000, ε-greedy policy decayed from 0.9 to 0.1, gradient clipping threshold 0.

5.

6. The method for secure isolation and threat detection of virtualized network slices according to claim 1, characterized in that: In step three, the framework comprises four core layers: a state perception layer, a feature extraction layer, a decision layer, and an execution layer. The state perception layer is responsible for the real-time collection of network traffic and system behavior data. The feature extraction layer extracts spatiotemporal correlation features through attention mechanisms and graph neural networks. The decision layer performs threat judgment based on a multi-agent reinforcement learning algorithm. The execution layer outputs detection results and response strategies.