A network evaluation method, apparatus and storage medium

By monitoring and fault injection to evaluate key indicators of NFV networks and network elements, this technology addresses the problem that existing technologies have failed to assess the fault tolerance capabilities of NFV networks and network elements, and enables effective assessment and verification of network resilience and fault tolerance.

CN115708376BActive Publication Date: 2025-12-02CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110959933.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-20
Publication Date
2025-12-02
Estimated Expiration
2041-08-20

AI Technical Summary

Technical Problem

Existing technical solutions fail to effectively assess the fault tolerance of NFV networks and network elements, as well as their resilience to damage within acceptable business metrics.

Method used

By monitoring and collecting key indicators of the NFV network before and after fault injection, the fault tolerance capability of the NFV network and network elements is evaluated, including injecting faults layer by layer and controlling the fluctuation of predetermined indicators within a preset range, and using specific indicator combinations to evaluate the resilience of the network.

Benefits of technology

It enables the assessment of the service quality resilience of NFV networks and network elements under fault conditions, verifies the network's cloud-based scalability and fault tolerance, and reduces the impact of uncertainty on steady state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115708376B_ABST
    Figure CN115708376B_ABST
Patent Text Reader

Abstract

This invention discloses a network evaluation method, apparatus, and storage medium, comprising: determining the metrics to be evaluated in an NFV network; monitoring the metrics and collecting a first metric of the metrics before fault injection; injecting a fault; monitoring the metrics and collecting a second metric of the metrics after fault injection; restoring the NFV network to the state where the first metric is true; and evaluating the NFV network based on the second metric. Using this invention, a technical solution can be provided for evaluating the fault tolerance capability of an NFV network and the resilience of the NFV network to damage within acceptable ranges of service metrics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a network evaluation method, apparatus, and storage medium. Background Technology

[0002] For NFV (Network Function Virtualization) networks based on cloud resource pools, high availability is typically achieved by leveraging cloud computing's redundant resource load balancing technology and stateless system architecture, thereby improving overall network stability. Existing technical solutions mainly employ service stress testing, using system availability metrics or data integrity metrics to evaluate the overall network reliability.

[0003] The shortcomings of existing technology are:

[0004] Existing technical solutions do not address how to assess the fault tolerance of NFV networks and network elements, or the resilience of NFV networks and network elements to damage within acceptable service metrics. Summary of the Invention

[0005] This invention provides a network and network element evaluation method, apparatus, and storage medium to address the lack of a solution for evaluating the fault tolerance of NFV networks and network elements, as well as the resilience of NFV networks to damage within acceptable service metrics.

[0006] This invention provides the following technical solutions:

[0007] A network evaluation method, comprising:

[0008] Identify the metrics that need to be evaluated in the NFV network;

[0009] Monitor the aforementioned indicators and collect the first indicator of the aforementioned indicators before fault injection;

[0010] Injection failure;

[0011] Monitor the aforementioned indicator and collect a second indicator of the aforementioned indicator after fault injection;

[0012] Restoring the NFV network to the aforementioned metric is the first metric;

[0013] NFV networks are evaluated based on the second metric.

[0014] During implementation, it further includes:

[0015] When a fault is injected, the predetermined indicators of the NFV network are controlled to fluctuate within a preset range.

[0016] In practice, the predetermined indicator is one or a combination of the following indicators:

[0017] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0018] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0019] One or a combination of the following indicators on 4G core network elements:

[0020] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0021] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0022] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0023] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0024] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0025] One or a combination of the following metrics on 5G core network elements:

[0026] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0027] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0028] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0029] Success rate metrics on UDM;

[0030] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0031] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0032] Success rate type metrics and error type metrics on NSSF.

[0033] In practice, the predetermined indicator is one or a combination of the following indicators:

[0034] One or a combination of the following metrics on 4G base stations:

[0035] Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of uplink and / or downlink PDCP SDU bytes on the cell user plane and / or control plane;

[0036] Success rate metrics include: E-RAB establishment success rate (related to services), wireless connection success rate (related to services), RRC connection establishment success rate, and handover success rate.

[0037] Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time;

[0038] Error type metrics: wireless call drop rate, call drop rate, TCP-based wireless call drop rate, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops;

[0039] One or a combination of the following metrics on 5G base stations:

[0040] Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP PDU bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transmission TB, downlink PDSCH PRB available number, paging record received number, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of ERABs for NSASCG Split Bear type;

[0041] Success rate metrics include: RRC connection establishment success rate, Flow establishment success rate, PDUSESSION establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate.

[0042] Latency type indicators: RRC connection average and / or maximum setup time, RLC downlink data packet average processing latency, gNB inter-NG handover average time, gNB inter-Xn handover average time, Epsfallback service handover latency from 5G to 4G, and RLC downlink data packet average processing latency per slice cell;

[0043] Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses;

[0044] One or a combination of the following indicators on 4G core network elements:

[0045] On MME:

[0046] Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs;

[0047] Success rate metrics include: EPS attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate.

[0048] Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time;

[0049] Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures;

[0050] On ServingGW:

[0051] Traffic volume and capacity type metrics: SGW average and / or peak capacity utilization, user plane uplink and / or downlink traffic, average GTP uplink and / or downlink traffic per attached user of ServingGW, average and / or maximum number of attached users of ServingGW, average and / or maximum number of bearers of ServingGW, SGW data throughput capacity utilization;

[0052] Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate;

[0053] On PGW:

[0054] Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic.

[0055] Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate;

[0056] Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW;

[0057] Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface;

[0058] On HSS:

[0059] Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate;

[0060] Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate;

[0061] Error type indicator: Number of times the position update failed;

[0062] On PCRF:

[0063] Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity;

[0064] Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate;

[0065] Error type metric: Application session call drop rate;

[0066] One or a combination of the following metrics on 5G core network elements:

[0067] On AMF:

[0068] Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses;

[0069] Success rate metrics include: initial registration success rate, registration update success rate, switchover success rate, UECM registration failure success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate.

[0070] Latency-related metrics: Average initial registration time;

[0071] Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections;

[0072] On SMF:

[0073] Traffic volume type metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows;

[0074] Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate;

[0075] Latency type metric: Average duration of PDU session establishment process;

[0076] Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures;

[0077] On UPF:

[0078] Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface;

[0079] Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate;

[0080] Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface;

[0081] On UDM:

[0082] Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate;

[0083] On PCF:

[0084] Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy;

[0085] Success rate type indicators: AM strategy association establishment success rate, AM strategy association update success rate, AM strategy association deletion success rate, SM strategy association establishment success rate, SM strategy association update success rate, SM strategy association deletion success rate;

[0086] Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures;

[0087] On NRF:

[0088] Business volume type metric: Number of storage instances;

[0089] Success rate metrics: NF update initiation success rate, NF discovery success rate;

[0090] Error type indicators: Number of NF update failures and number of NF discovery failures;

[0091] On NSSF:

[0092] Success rate metrics: Network slice selection success rate;

[0093] Error type metric: Number of network slice selection failures.

[0094] In practice, when injecting faults, faults are injected layer by layer from the infrastructure layer to the VNF layer, according to the NFV network functional architecture; and / or,

[0095] Based on the NFV network functional architecture, faults are injected layer by layer from the VNF layer to the infrastructure layer.

[0096] In practice, when a fault is injected, it is in one or a combination of the following aspects:

[0097] The main NFV network service plane, backup NFV network service plane, management plane, network element component level, and infrastructure.

[0098] During implementation, one or a combination of the following faults may be injected into the NFV network service plane:

[0099] The main NFV network service injection failure was caused by a CloudOS database IP address conflict.

[0100] The main NFV network service injection fault resulted in the complete blocking of VIM and damage to some service virtual machines of network elements.

[0101] The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services.

[0102] A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region.

[0103] The main regional service injection failure was caused by high read / write latency in the CloudOS storage component.

[0104] The service injection failure in the main region was a CloudOS storage component switching failure.

[0105] The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

[0106] During implementation, the OMU fault injected into the management plane was VNF, which prevented access.

[0107] In practice, the metric is one or a combination of the following NFV network performance metrics:

[0108] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0109] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0110] One or a combination of the following indicators on 4G core network elements:

[0111] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0112] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0113] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0114] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0115] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0116] One or a combination of the following metrics on 5G core network elements:

[0117] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0118] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0119] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0120] Success rate metrics on UDM;

[0121] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0122] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0123] Success rate type metrics and error type metrics on NSSF.

[0124] In practice, evaluating an NFV network or network element based on the second indicator assesses the overall network service quality's resilience to damage when components in the NFV network or network element fail.

[0125] During implementation, NFV networks are evaluated according to the second indicator in one or a combination of the following ways:

[0126] NFV element resilience = α * (a1 * network component failure rate + a2 * compute / storage component failure rate + α) + β * OMU failure rate; where NFV element resilience is an evaluation index for assessing the ability of an NFV element to withstand service quality disruptions when components fail, α represents the user plane weight, β represents the management plane weight, and a1 * network component failure rate + β * OMU failure rate. i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0127] NFV network resilience = α * NFV network element failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability compute and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability compute and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability compute and storage component failure rate) * 100%; where NFV network resilience is an evaluation index for assessing the overall network service quality resilience against damage when NFV network elements, EMS, VNFM, and NFVO functional components fail; α represents the user plane weight, β represents the management plane weight, and b... i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0128] 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0129] 5G wireless network element resilience = (α * (a1 * ∑5G base station high availability network component failure rate + a2 * ∑5G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0130] 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0131] 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0132] 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0133] 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0134] 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch (Top of Rack), EOR switch (End of Row), network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0135] 5G core network resilience = (α * ∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability computing and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability computing and storage component failure rate +) * 100%, where high-availability components include the following components One or a combination thereof: TOR switch, EOR switch, network interface card, and computing and storage including one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 range from [0, 1]. NFV network element resilience = α * (a1 * network component failure rate + a2 * computing / storage component failure rate + ) + β * OMU failure rate; where, NFV network element resilience is an evaluation index for assessing the ability of network element services to withstand damage when components in an NFV network element fail, α represents the user plane weight, β represents the management plane weight, and a i{i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane.

[0136] In practice, evaluating an NFV network according to the second metric involves evaluating one or a combination of the following performance characteristics of the NFV network:

[0137] Wireless network element resilience, core network element resilience,

[0138] The wireless network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the base station, and the core network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the core network element.

[0139] In implementation, the aforementioned wireless network element resilience is used to verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of the network element; and / or,

[0140] The core network element resilience is used to verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements.

[0141] During implementation, the resilience of the NFV network, the resilience of the wireless network elements, the resilience of the core network elements, or a combination thereof, shall be assessed according to the second indicator in the following manner:

[0142] Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0143] Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0144] A network evaluation device, comprising:

[0145] The processor is used to read programs from memory and execute the following procedures:

[0146] Identify the metrics that need to be evaluated in the NFV network;

[0147] Monitor the aforementioned indicators and collect the first indicator of the aforementioned indicators before fault injection;

[0148] Injection failure;

[0149] Monitor the aforementioned indicator and collect a second indicator of the aforementioned indicator after fault injection;

[0150] Restoring the NFV network to the aforementioned metric is the first metric;

[0151] Evaluate NFV networks based on the second metric;

[0152] A transceiver is used to receive and send data under the control of a processor.

[0153] During implementation, it further includes:

[0154] When a fault is injected, the predetermined indicators of the NFV network are controlled to fluctuate within a preset range.

[0155] In practice, the predetermined indicator is one or a combination of the following indicators:

[0156] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0157] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0158] One or a combination of the following indicators on 4G core network elements:

[0159] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0160] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0161] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0162] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0163] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0164] One or a combination of the following metrics on 5G core network elements:

[0165] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0166] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0167] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0168] Success rate metrics on UDM;

[0169] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0170] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0171] Success rate type metrics and error type metrics on NSSF.

[0172] In practice, the predetermined indicator is one or a combination of the following indicators:

[0173] One or a combination of the following metrics on 4G base stations:

[0174] Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of uplink and / or downlink PDCP SDU bytes on the cell user plane and / or control plane;

[0175] Success rate metrics include: E-RAB establishment success rate (related to services), wireless connection success rate (related to services), RRC connection establishment success rate, and handover success rate.

[0176] Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time;

[0177] Error type metrics: wireless call drop rate, call drop rate, TCP-based wireless call drop rate, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops;

[0178] One or a combination of the following metrics on 5G base stations:

[0179] Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP PDU bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transmission TB, downlink PDSCH PRB available number, paging record received number, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of ERABs for NSA SCG Split Bear type;

[0180] Success rate metrics include: RRC connection establishment success rate, Flow establishment success rate, PDUSESSION establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate.

[0181] Latency type indicators: RRC connection average and / or maximum setup time, RLC downlink data packet average processing latency, gNB inter-NG handover average time, gNB inter-Xn handover average time, Epsfallback service handover latency from 5G to 4G, and RLC downlink data packet average processing latency per slice cell;

[0182] Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses;

[0183] One or a combination of the following indicators on 4G core network elements:

[0184] On MME:

[0185] Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs;

[0186] Success rate metrics include: EPS attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate.

[0187] Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time;

[0188] Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures;

[0189] On ServingGW:

[0190] Traffic volume and capacity type metrics: SGW average and / or peak capacity utilization, user plane uplink and / or downlink traffic, average GTP uplink and / or downlink traffic per attached user of ServingGW, average and / or maximum number of attached users of ServingGW, average and / or maximum number of bearers of ServingGW, SGW data throughput capacity utilization;

[0191] Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate;

[0192] On PGW:

[0193] Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic.

[0194] Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate;

[0195] Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW;

[0196] Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface;

[0197] On HSS:

[0198] Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate;

[0199] Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate;

[0200] Error type indicator: Number of times the position update failed;

[0201] On PCRF:

[0202] Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity;

[0203] Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate;

[0204] Error type metric: Application session call drop rate;

[0205] One or a combination of the following metrics on 5G core network elements:

[0206] On AMF:

[0207] Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses;

[0208] Success rate metrics include: initial registration success rate, registration update success rate, switchover success rate, UECM registration failure success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate.

[0209] Latency-related metrics: Average initial registration time;

[0210] Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections;

[0211] On SMF:

[0212] Traffic volume type metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows;

[0213] Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate;

[0214] Latency type metric: Average duration of PDU session establishment process;

[0215] Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures;

[0216] On UPF:

[0217] Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface;

[0218] Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate;

[0219] Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface;

[0220] On UDM:

[0221] Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate;

[0222] On PCF:

[0223] Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy;

[0224] Success rate type indicators: AM strategy association establishment success rate, AM strategy association update success rate, AM strategy association deletion success rate, SM strategy association establishment success rate, SM strategy association update success rate, SM strategy association deletion success rate;

[0225] Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures;

[0226] On NRF:

[0227] Business volume type metric: Number of storage instances;

[0228] Success rate metrics: NF update initiation success rate, NF discovery success rate;

[0229] Error type indicators: Number of NF update failures and number of NF discovery failures;

[0230] On NSSF:

[0231] Success rate metrics: Network slice selection success rate;

[0232] Error type metric: Number of network slice selection failures.

[0233] In practice, when injecting faults, faults are injected layer by layer from the infrastructure layer to the VNF layer, according to the NFV network functional architecture; and / or,

[0234] Based on the NFV network functional architecture, faults are injected layer by layer from the VNF layer to the infrastructure layer.

[0235] In practice, when a fault is injected, it is in one or a combination of the following aspects:

[0236] The main NFV network service plane, backup NFV network service plane, management plane, network element component level, and infrastructure.

[0237] During implementation, one or a combination of the following faults may be injected into the NFV network service plane:

[0238] The main NFV network service injection failure was caused by a CloudOS database IP address conflict.

[0239] The main NFV network service injection fault resulted in the complete blocking of VIM and damage to some service virtual machines of network elements.

[0240] The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services.

[0241] A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region.

[0242] The main regional service injection failure was caused by high read / write latency in the CloudOS storage component.

[0243] The service injection failure in the main region was a CloudOS storage component switching failure.

[0244] The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

[0245] During implementation, the OMU fault injected into the management plane was VNF, which prevented access.

[0246] In practice, the metric is one or a combination of the following NFV network performance metrics:

[0247] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0248] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0249] One or a combination of the following indicators on 4G core network elements:

[0250] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0251] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0252] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0253] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0254] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0255] One or a combination of the following metrics on 5G core network elements:

[0256] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0257] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0258] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0259] Success rate metrics on UDM;

[0260] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0261] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0262] Success rate type metrics and error type metrics on NSSF.

[0263] In practice, evaluating the NFV network based on the second indicator assesses the overall network service quality's resilience to damage when components in the NFV network fail.

[0264] During implementation, NFV networks are evaluated according to the second indicator in one or a combination of the following ways:

[0265] NFV element resilience = α * (a1 * network component failure rate + a2 * compute / storage component failure rate + α) + β * OMU failure rate; where NFV element resilience is an evaluation index for assessing the ability of an NFV element to withstand service quality disruptions when components fail, α represents the user plane weight, β represents the management plane weight, and a1 * network component failure rate + β * OMU failure rate. i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0266] NFV network resilience = α * NFV network element failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability compute and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability compute and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability compute and storage component failure rate) * 100%; where NFV network resilience is an evaluation index for assessing the overall network service quality resilience against damage when NFV network elements, EMS, VNFM, and NFVO functional components fail; α represents the user plane weight, β represents the management plane weight, and b... i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0267] 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0268] 5G wireless network element resilience = (α * (a1 * ∑5G base station high availability network component failure rate + a2 * ∑5G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0269] 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0270] 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0271] 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0272] 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0273] 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0274] 5G core network resilience = (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0275] In practice, evaluating an NFV network according to the second metric involves evaluating one or a combination of the following performance characteristics of the NFV network:

[0276] Wireless network element resilience, core network element resilience,

[0277] The wireless network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the base station, and the core network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the core network element.

[0278] In implementation, the aforementioned wireless network element resilience is used to verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of the network element; and / or,

[0279] The core network element resilience is used to verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements.

[0280] During implementation, the resilience of the NFV network, the resilience of the wireless network elements, the resilience of the core network elements, or a combination thereof, shall be assessed according to the second indicator in the following manner:

[0281] Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0282] Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0283] A network evaluation device, comprising:

[0284] The determination module is used to determine the metrics that need to be evaluated in the NFV network;

[0285] The monitoring module is used to monitor the indicators and collect the first indicators of the indicators before fault injection.

[0286] The fault module is used to inject faults.

[0287] The monitoring module is also used to monitor the indicator and collect a second indicator of the indicator after fault injection;

[0288] The fault module is also used to restore the NFV network to the first metric.

[0289] The evaluation module is used to evaluate NFV networks based on a second metric.

[0290] During implementation, it further includes:

[0291] The control module is used to control the fluctuation of predetermined indicators of the NFV network within a preset range when a fault is injected.

[0292] In implementation, the control module is further used for one or a combination of the following predetermined indicators:

[0293] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0294] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0295] One or a combination of the following indicators on 4G core network elements:

[0296] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0297] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0298] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0299] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0300] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0301] One or a combination of the following metrics on 5G core network elements:

[0302] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0303] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0304] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0305] Success rate metrics on UDM;

[0306] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0307] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0308] Success rate type metrics and error type metrics on NSSF.

[0309] In practice, the predetermined indicator is one or a combination of the following indicators:

[0310] One or a combination of the following metrics on 4G base stations:

[0311] Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of uplink and / or downlink PDCP SDU bytes on the cell user plane and / or control plane;

[0312] Success rate metrics include: E-RAB establishment success rate (related to services), wireless connection success rate (related to services), RRC connection establishment success rate, and handover success rate.

[0313] Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time;

[0314] Error type metrics: wireless call drop rate, call drop rate, TCP-based wireless call drop rate, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops;

[0315] One or a combination of the following metrics on 5G base stations:

[0316] Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP PDU bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transmission TB, downlink PDSCH PRB available number, paging record received number, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of ERABs for NSA SCG Split Bear type;

[0317] Success rate metrics include: RRC connection establishment success rate, Flow establishment success rate, PDUSESSION establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate.

[0318] Latency type indicators: RRC connection average and / or maximum setup time, RLC downlink data packet average processing latency, gNB inter-NG handover average time, gNB inter-Xn handover average time, Epsfallback service handover latency from 5G to 4G, and RLC downlink data packet average processing latency per slice cell;

[0319] Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses;

[0320] One or a combination of the following indicators on 4G core network elements:

[0321] On MME:

[0322] Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs;

[0323] Success rate metrics include: EPS attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate.

[0324] Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time;

[0325] Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures;

[0326] On ServingGW:

[0327] Traffic volume and capacity type metrics: SGW average and / or peak capacity utilization, user plane uplink and / or downlink traffic, average GTP uplink and / or downlink traffic per attached user of ServingGW, average and / or maximum number of attached users of ServingGW, average and / or maximum number of bearers of ServingGW, SGW data throughput capacity utilization;

[0328] Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate;

[0329] On PGW:

[0330] Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic.

[0331] Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate;

[0332] Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW;

[0333] Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface;

[0334] On HSS:

[0335] Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate;

[0336] Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate;

[0337] Error type indicator: Number of times the position update failed;

[0338] On PCRF:

[0339] Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity;

[0340] Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate;

[0341] Error type metric: Application session call drop rate;

[0342] One or a combination of the following metrics on 5G core network elements:

[0343] On AMF:

[0344] Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses;

[0345] Success rate metrics include: initial registration success rate, registration update success rate, switchover success rate, UECM registration failure success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate.

[0346] Latency-related metrics: Average initial registration time;

[0347] Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections;

[0348] On SMF:

[0349] Traffic volume type metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows;

[0350] Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate;

[0351] Latency type metric: Average duration of PDU session establishment process;

[0352] Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures;

[0353] On UPF:

[0354] Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface;

[0355] Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate;

[0356] Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface;

[0357] On UDM:

[0358] Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate;

[0359] On PCF:

[0360] Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy;

[0361] Success rate type indicators: AM strategy association establishment success rate, AM strategy association update success rate, AM strategy association deletion success rate, SM strategy association establishment success rate, SM strategy association update success rate, SM strategy association deletion success rate;

[0362] Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures;

[0363] On NRF:

[0364] Business volume type metric: Number of storage instances;

[0365] Success rate metrics: NF update initiation success rate, NF discovery success rate;

[0366] Error type indicators: Number of NF update failures and number of NF discovery failures;

[0367] On NSSF:

[0368] Success rate metrics: Network slice selection success rate;

[0369] Error type metric: Number of network slice selection failures.

[0370] In implementation, the fault module is further used to inject faults layer by layer from the infrastructure layer to the VNF layer, according to the NFV network functional architecture; and / or,

[0371] Based on the NFV network functional architecture, faults are injected layer by layer from the VNF layer to the infrastructure layer.

[0372] In implementation, the fault module is further used to inject faults in one or a combination of the following ways:

[0373] The main NFV network service plane, backup NFV network service plane, management plane, network element component level, and infrastructure.

[0374] In implementation, the fault module is further used to inject one or a combination of the following faults into the NFV network service plane:

[0375] The main NFV network service injection failure was caused by a CloudOS database IP address conflict.

[0376] The main NFV network service injection fault resulted in the complete blocking of VIM and damage to some service virtual machines of network elements.

[0377] The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services.

[0378] A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region.

[0379] The main regional service injection failure was caused by high read / write latency in the CloudOS storage component.

[0380] The service injection failure in the main region was a CloudOS storage component switching failure.

[0381] The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

[0382] In practice, the fault module is further used to inject an OMU fault that is inaccessible in the VNF into the management plane.

[0383] In implementation, the determination module is further used to determine that the indicator is one or a combination of the following NFV network performance indicators:

[0384] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0385] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0386] One or a combination of the following indicators on 4G core network elements:

[0387] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0388] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0389] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0390] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0391] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0392] One or a combination of the following metrics on 5G core network elements:

[0393] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0394] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0395] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0396] Success rate metrics on UDM;

[0397] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0398] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0399] Success rate type metrics and error type metrics on NSSF.

[0400] In practice, the predetermined indicator is one or a combination of the following indicators:

[0401] One or a combination of the following metrics on 4G base stations:

[0402] Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of uplink and / or downlink PDCP SDU bytes on the cell user plane and / or control plane;

[0403] Success rate metrics include: E-RAB establishment success rate (related to services), wireless connection success rate (related to services), RRC connection establishment success rate, and handover success rate.

[0404] Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time;

[0405] Error type metrics: wireless call drop rate, call drop rate, TCP-based wireless call drop rate, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops;

[0406] One or a combination of the following metrics on 5G base stations:

[0407] Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP PDU bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transmission TB, downlink PDSCH PRB available number, paging record received number, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of ERABs for NSA SCG Split Bear type;

[0408] Success rate metrics include: RRC connection establishment success rate, Flow establishment success rate, PDUSESSION establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate.

[0409] Latency type indicators: RRC connection average and / or maximum setup time, RLC downlink data packet average processing latency, gNB inter-NG handover average time, gNB inter-Xn handover average time, Epsfallback service handover latency from 5G to 4G, and RLC downlink data packet average processing latency per slice cell;

[0410] Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses;

[0411] One or a combination of the following indicators on 4G core network elements:

[0412] On MME:

[0413] Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs;

[0414] Success rate metrics include: EPS attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate.

[0415] Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time;

[0416] Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures;

[0417] On ServingGW:

[0418] Traffic volume and capacity type metrics: SGW average and / or peak capacity utilization, user plane uplink and / or downlink traffic, average GTP uplink and / or downlink traffic per attached user of ServingGW, average and / or maximum number of attached users of ServingGW, average and / or maximum number of bearers of ServingGW, SGW data throughput capacity utilization;

[0419] Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate;

[0420] On PGW:

[0421] Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic.

[0422] Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate;

[0423] Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW;

[0424] Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface;

[0425] On HSS:

[0426] Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate;

[0427] Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate;

[0428] Error type indicator: Number of times the position update failed;

[0429] On PCRF:

[0430] Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity;

[0431] Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate;

[0432] Error type metric: Application session call drop rate;

[0433] One or a combination of the following metrics on 5G core network elements:

[0434] On AMF:

[0435] Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses;

[0436] Success rate metrics include: initial registration success rate, registration update success rate, switchover success rate, UECM registration failure success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate.

[0437] Latency-related metrics: Average initial registration time;

[0438] Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections;

[0439] On SMF:

[0440] Traffic volume type metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows;

[0441] Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate;

[0442] Latency type metric: Average duration of PDU session establishment process;

[0443] Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures;

[0444] On UPF:

[0445] Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface;

[0446] Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate;

[0447] Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface;

[0448] On UDM:

[0449] Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate;

[0450] On PCF:

[0451] Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy;

[0452] Success rate type indicators: AM strategy association establishment success rate, AM strategy association update success rate, AM strategy association deletion success rate, SM strategy association establishment success rate, SM strategy association update success rate, SM strategy association deletion success rate;

[0453] Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures;

[0454] On NRF:

[0455] Business volume type metric: Number of storage instances;

[0456] Success rate metrics: NF update initiation success rate, NF discovery success rate;

[0457] Error type indicators: Number of NF update failures and number of NF discovery failures;

[0458] On NSSF:

[0459] Success rate metrics: Network slice selection success rate;

[0460] Error type metric: Number of network slice selection failures.

[0461] In practice, the evaluation module is further used to evaluate the NFV network based on the second indicator, which assesses the overall network service quality resilience to damage when components in the NFV network fail.

[0462] In implementation, the evaluation module is further used to evaluate the NFV network based on the second metric in one or a combination of the following ways:

[0463] NFV element resilience = α * (a1 * network component failure rate + a2 * compute / storage component failure rate + α) + β * OMU failure rate; where NFV element resilience is an evaluation index for assessing the ability of an NFV element to withstand service quality disruptions when components fail, α represents the user plane weight, β represents the management plane weight, and a1 * network component failure rate + β * OMU failure rate. i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0464] NFV network resilience = α * NFV network element failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability compute and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability compute and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability compute and storage component failure rate) * 100%; where NFV network resilience is an evaluation index for assessing the overall network service quality resilience against damage when NFV network elements, EMS, VNFM, and NFVO functional components fail; α represents the user plane weight, β represents the management plane weight, and b... i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0465] 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0466] 5G wireless network element resilience = (α * (a1 * ∑5G base station high availability network component failure rate + a2 * ∑5G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0467] 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0468] 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0469] 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0470] 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0471] 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0472] 5G core network resilience = (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0473] In implementation, the evaluation module is further used to evaluate the NFV network based on the second metric, which is to evaluate one or a combination of the following performance characteristics of the NFV network:

[0474] Wireless network element resilience, core network element resilience,

[0475] The wireless network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the base station, and the core network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the core network element.

[0476] In implementation, the aforementioned wireless network element resilience is used to verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of the network element; and / or,

[0477] The core network element resilience is used to verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements.

[0478] In implementation, the assessment module is further used to assess one or a combination of the following: the resilience of the NFV network's wireless network elements, the resilience of the core network elements, and so on, based on the second metric:

[0479] Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; α and β are set to 0.5, and a1 and a2 are recommended to be set to 1;

[0480] Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0481] A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program that performs the above-described network evaluation method.

[0482] The beneficial effects of this invention are as follows:

[0483] In the technical solution provided by the embodiments of the present invention, since the indicators to be evaluated in the NFV network are determined first, and then a fault is injected, the indicators after the fault injection are collected, and the NFV network is evaluated based on the indicators, a technical solution is provided to evaluate the fault tolerance capability of the NFV network and the degree of damage resistance of the NFV network within the acceptable range of service indicators. Attached Figure Description

[0484] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0485] Figure 1 This is a schematic diagram illustrating the implementation process of the network evaluation method in this embodiment of the invention;

[0486] Figure 2 This is a schematic diagram of the components of an NFV network in an embodiment of the present invention;

[0487] Figure 3 This is a schematic diagram of the relationships within the chaos engineering experimental system in an embodiment of the present invention;

[0488] Figure 4 This is a schematic diagram of the closed-loop system of the chaos engineering experiment in this embodiment of the invention;

[0489] Figure 5 This is a schematic diagram of the NFV network functional architecture in an embodiment of the present invention;

[0490] Figure 6 This is a schematic diagram of the network evaluation device in an embodiment of the present invention. Detailed Implementation

[0491] Existing technical solutions do not address how to assess the fault tolerance capability of NFV networks, or the resilience of NFV networks within acceptable service metrics. Therefore, this solution primarily addresses how to more accurately assess the fault tolerance capability of NFV networks, i.e., their resilience, also referred to as toughness in this invention embodiment, which is specifically divided into two parts: wireless network element toughness and wireless network toughness.

[0492] The specific embodiments of the present invention will now be described with reference to the accompanying drawings.

[0493] Figure 1 The diagram illustrates the implementation process of the network evaluation method, which may include:

[0494] Step 101: Determine the metrics to be evaluated in the NFV network;

[0495] Step 102: Monitor the indicator and collect the first indicator of the indicator before fault injection;

[0496] Step 103: Inject fault;

[0497] Step 104: Monitor the indicator and collect a second indicator of the indicator after fault injection;

[0498] Step 105: Restore the NFV network to the first indicator;

[0499] Step 106: Evaluate the NFV network based on the second metric.

[0500] The embodiment employs chaos engineering methods. By injecting faults into each layer / module of the NFV network, the impact on network service indicators is monitored. Then, the proportion of faulty components in each layer / module is analyzed to calculate NFV network resilience indicators, which are used to measure the reliability of the NFV network against damage. The implementation will be described in three main parts below:

[0501] 1. Establish a system framework for testing the reliability of NFV networks using chaos engineering;

[0502] Second: Inject faults into each layer / module and monitor the extent to which business metrics are affected;

[0503] Third: Calculate the system resilience index, which is used to measure the system's reliability in resisting damage.

[0504] In the cloud era, with the evolution of large-scale distributed systems and microservice architecture technologies, the behavior of each individual microservice is reasonable. However, the dependencies between services are complex, and in certain scenarios, the combination of dependency behaviors between microservices may lead to unexpected system problems. At the same time, due to the rapid iteration of business and technology, continuously ensuring the stability and health of the system will face significant challenges.

[0505] By employing chaos engineering methods, through the development and implementation of plans, clarifying the steady-state indicators of NFV networks, performing fault injection, checking NFV network steady-state indicators, repairing discovered problems, and automating continuous verification, the impact of uncertainties on NFV network steady-state can be minimized. This includes verifying the fault tolerance limits of NFV networks, the scalability of cloud services, the timeliness of monitoring and alarms, and the emergency response capabilities for locating and resolving problems. As a result, the fault tolerance of NFV network architecture can be improved, the efficiency of emergency fault handling can be increased, the fault recurrence rate can be reduced, and the overall reliability of NFV networks can be improved.

[0506] Figure 2 The diagram below illustrates the components of an NFV network. Typically, an NFV network consists of the following components, primarily a management plane and a service plane. The management plane includes NFVO (NFV Orchestrator) / NFVO+, VNFM (VNF Manager; VNF: Virtualization Network Function), and EMS (Element Management System). The service plane includes the infrastructure layer (within the data center), compute / storage layer, CloudOS (Cloud Operating System) layer, and VNF layer. Each layer contains a varying number of key components.

[0507] 1. Establish a system framework for testing the reliability of NFV networks using chaos engineering.

[0508] Figure 3The diagram illustrates the relationships within the chaos engineering experimental system. First, a chaos engineering experimental system is established to conduct large-scale fault drills and possesses orchestration capabilities to adapt to new services. The main functional architecture includes a fault injection capability layer, an orchestration management layer, and an application layer. The chaos engineering experimental system interfaces with the NFV network and network management system, enabling fault injection into various functional modules of the NFV network, comprehensive monitoring of the NFV network, and real-time closed-loop management for evaluating experimental results.

[0509] Figure 4 The diagram shows a closed-loop system for a chaos engineering experiment. A complete chaos engineering experiment is a continuously iterating closed-loop system, mainly including:

[0510] Experimental requirements (iteration), feasibility evaluation, monitoring indicator design, scenario design, fault injection capability orchestration, experimental plan development, experimental execution, environment restoration, result analysis, issue tracking, pipeline integration, experimental requirements (iteration)...

[0511] Starting with experimental requirements, a feasibility assessment is conducted to determine the experimental scope, design appropriate monitoring metrics, experimental scenarios and environments, and select suitable fault injection capabilities for orchestration. An experimental plan is established, and full communication with stakeholders of the experimental subjects is conducted to jointly execute the experimental process and collect pre-designed monitoring metrics. After the experiment is completed, the experimental environment is cleaned and restored, the experimental results are analyzed, root causes are traced, and problems are resolved. The above experimental scenarios are automated and integrated into the pipeline for regular execution. Afterward, new experimental scopes can be added, with continuous iteration and orderly improvement. The feasibility assessment needs to consider the architecture's fault resilience assessment, experimental environment selection (development environment, testing environment, pre-production environment, production environment), the minimum blast radius in the fault scenarios (the minimum impact of a fault on business during the exercise), and the experimental environment's recovery capability assessment.

[0512] Second: Injecting faults into the NFV network and monitoring the extent to which business metrics are affected.

[0513] In practice, when injecting faults, faults are injected layer by layer from the infrastructure layer to the VNF layer, according to the NFV network functional architecture; and / or,

[0514] Based on the NFV network functional architecture, faults are injected layer by layer from the VNF layer to the infrastructure layer.

[0515] Specifically, Figure 5 The diagram illustrates the functional architecture of an NFV network. As shown, injecting faults layer by layer from the infrastructure layer to the VNF layer increases the service impact radius and exposes a wider range of problems. Conversely, injecting faults layer by layer from the VNF layer to the infrastructure layer decreases the service impact radius and exposes more focused problems.

[0516] In practice, when a fault is injected, it is in one or a combination of the following aspects:

[0517] The main NFV network service plane, backup NFV network service plane, management plane, network element component level, and infrastructure.

[0518] Faults were injected into the service plane, management plane, and network element components of the primary and backup NFV networks to verify the network's resilience. The main fault scenarios included complete service outages in the primary NFV network, compromised service quality in the primary NFV network, impaired user plane services, and management plane failures.

[0519] During implementation, one or a combination of the following faults may be injected into the NFV network service plane:

[0520] The main NFV network service injection failure was caused by a CloudOS database IP address conflict.

[0521] The main NFV network service injection fault resulted in the complete blocking of VIM and damage to some service virtual machines of network elements.

[0522] The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services.

[0523] A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region.

[0524] The main regional service injection failure was caused by high read / write latency in the CloudOS storage component.

[0525] The service injection failure in the main region was a CloudOS storage component switching failure.

[0526] The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

[0527] The specific implementation can be as follows:

[0528] Scenario of complete disruption of main NFV network services: The injection fault is caused by a CloudOS database IP address conflict, resulting in a complete disruption of VIM (Virtual Infrastructure Management) services, coupled with damage to some virtual machines providing services on network elements, and the data center gateway deleting the S1 interface route, leading to a complete disruption of main NFV network services. This is used to verify the cross-region disaster recovery and failover capabilities of VNFs.

[0529] Scenarios where the quality of business in the main region is affected: injection failures cause batch virtual machine storage disconnections in the main region, large read / write latency of CloudOS storage components, and switching failures. Verify storage bypass capabilities and VNF fault location and handling capabilities.

[0530] Scenario where user plane service quality is impaired: The injected fault is a CloudOS component failure that causes the connection between user plane network elements and control plane network elements to be interrupted, verifying the fault location and handling capabilities.

[0531] During implementation, the OMU fault injected into the management plane was VNF, which prevented access.

[0532] Management plane failure: The injected failure is that the OMU (Operations and Maintenance Unit) of the VNF is inaccessible, verifying the ability to locate and handle the failure.

[0533] In practice, it may further include:

[0534] When a fault is injected, the predetermined indicators of the NFV network are controlled to fluctuate within a preset range.

[0535] In practice, the predetermined indicator is one or a combination of the following indicators:

[0536] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0537] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0538] One or a combination of the following indicators on 4G core network elements:

[0539] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME (Mobility Management Entity);

[0540] Traffic volume and capacity type metrics, success rate type metrics on ServingGW (Serving Gateway);

[0541] On the PGW (PDN Gateway; PDN: Packet Data Network), traffic and capacity type metrics, latency type metrics, and error type metrics are provided.

[0542] Traffic and capacity type metrics, success rate type metrics, and error type metrics on the HSS (Home Subscriber Server);

[0543] Traffic type metrics, success rate type metrics, and error type metrics on the PCRF (Policy and Charging Rules Function);

[0544] One or a combination of the following metrics on 5G core network elements:

[0545] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on AMF (Access and Mobility Management Function);

[0546] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF (Session Management Function);

[0547] Business volume type metrics, success rate type metrics, and error type metrics on UPF (User Plane Function);

[0548] Success rate metrics on UDM (Unified Data Management Entity);

[0549] Traffic type metrics, success rate type metrics, and error type metrics on PCF (Policy Control Function);

[0550] Traffic type metrics, success rate type metrics, and error type metrics on NRF (Network Repository Function);

[0551] Success rate type metrics and error type metrics on NSSF (Network Slice Selection Function).

[0552] In specific implementation, the predetermined indicator is one or a combination of the following indicators:

[0553] One or a combination of the following metrics on 4G base stations: Traffic volume type metrics: number of bytes sent / received via eNB Ethernet interface, number of service bytes sent / received via eNB S1 interface, number of bytes in uplink / downlink PDCP (Packet Data Convergence Protocol) SDU (Service Data Unit), etc. on the cell user plane / control plane; Success rate type metrics: service-related E-RAB (Evolved Radio Access Bearer) establishment success rate, service-related radio connection success rate, RRC (Radio Resource Control)... Resource Control (RRC) connection establishment success rate, handover success rate (including intra-eNB / inter-eNB, same-frequency / different-frequency, handover in / handover out, intra-system / inter-system), latency type indicators: RRC connection average / maximum establishment time, E-RAB average / maximum establishment time, etc., error type indicators: wireless call drop rate (including conversational voice services and conversational live video streaming services), disconnection rate (including real-time gaming services, non-conversational buffered video streaming services, IMS signaling services), and TCP (Transmission Control Protocol) based wireless disconnection rate, uplink / downlink PDCP SDU packet loss rate, number of uplink / downlink transport block errors, and number of paging record drops, etc.

[0554] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics: gNB transmits / receives traffic data from the NG interface (SA networking only), cell user plane uplink / downlink PDCP PDU (Packet Data Unit) bytes (NSA & SA networking), gNB receives / transmits traffic data from the S1 interface (NSA networking only), uplink / downlink transport TB (Transport Block) count (NSA & SA networking), downlink PDSCH (Physical downlink shared channel) PRB (Physical Resource Block) available count (NSA & SA networking), paging record received count (SA networking only), average / maximum number of RRC connections in dual-connectivity and NSA (Non-Stand Alone) SCG (Secondary Cell Group) Split. Bearer (Separate Bearer) type average / maximum ERAB count, success rate indicators: RRC connection establishment success rate (SA network only) = number of successful RRC connection establishments / number of RRC connection establishment requests, Flow establishment success rate (SA network only) = number of successful Flow establishments / number of Flow establishment requests, PDUSESSION (PDU session) establishment success rate (SA network only) = number of successful PDUSESSION establishments / number of PDUSESSION establishment requests, handover success rate (including intra-frequency / inter-frequency, intra-gNB / inter-gNB, handover out / handover in) = number of successful handovers / number of handover requests, NG interface UE-related logical signaling connection establishment success rate (SA network only) = number of successful NG interface UE-related logical signaling connection establishments / number of NG interface UE-related logical signaling connection establishment requests, etc., latency indicators: RRC connection average / maximum establishment time, RLC (Radio Link Control)... LinkControl downlink packet average processing latency (SA network only), average handover time between gNBs (NG) (SA network only), average handover time between gNBs (Xn) (SA network only), Epsfallback (EPS: Evolved Packet System) service handover latency from 5G to 4G (SA network only), average downlink packet processing latency per slice cell (RLC) (SA network only), error type indicators: number of failed flow establishment / modification (SA network only), number of uplink / downlink PDCP packet losses (NSA & SA network only), etc.

[0555] One or a combination of the following indicators on 4G core network elements:

[0556] MME: Traffic type indicators: average / maximum number of MME bearers, number of users in idle / connected state of MME, average / maximum number of attached users of MME, etc.; Success rate type indicators: EPS attachment success rate, authentication success rate, DNS (Domain Name Server) resolution success rate initiated by MME, default / dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate (within / between MMEs, within / between systems), handover success rate (within / between MMEs, within / between systems), etc.; Latency type indicators: average / maximum attachment duration and average / maximum dedicated bearer establishment duration, etc.; Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures, etc.

[0557] ServingGW: Traffic and capacity type metrics: SGW bearer capacity average / peak utilization, user plane uplink / downlink traffic, average GTP (GPRS Tunneling Protocol; GPRS: General Packet Radio Service) uplink / downlink traffic per attached user of ServingGW, average / maximum number of attached users of ServingGW, average / maximum number of bearers of ServingGW, and SGW data throughput capacity utilization, etc. Success rate type metrics: ServingGW default / dedicated bearer establishment success rate, etc.

[0558] PGW: Traffic and capacity type indicators: PGW data throughput capacity utilization, PGW average / peak capacity utilization, PGW S5 / S8 interface uplink / downlink traffic, PGW average / maximum number of attached users, PGW average / peak number of bearers, SGi interface receive / send traffic, etc.; Success rate type indicators: dedicated bearer establishment success rate and CDR transmission success rate, etc.; Latency type indicators: average / maximum duration of dedicated bearer establishment initiated by PGW, etc.; Error type indicators: number of GTP packets dropped due to errors on PGW S5 / S8 interface and number of IP packets dropped due to errors on SGi interface, etc.

[0559] HSS: Service volume and capacity type indicators: number of HSS number releases / active users, HSS authentication capacity utilization rate and HSS static capacity utilization rate, etc.; success rate type indicators: success rate of HSS authentication information query, success rate of HSS location update / cancellation, success rate of HSS user data insertion / deletion, and success rate of HSS UE clearing, etc.; error type indicators: number of location update failures, etc.

[0560] PCRF: Traffic-type metrics: average / peak utilization of Gx session processing capacity, etc.; Success-type metrics: success rate of policy control initiation / update / end, success rate of re-authentication, and success rate of application session authorization, etc.; Error-type metrics: application session call loss rate, etc.

[0561] One or a combination of the following metrics on 5G core network elements:

[0562] AMF: Traffic volume metrics: Number of AMF registered / idle users, number of paging requests, number of first paging responses, and number of second paging responses, etc.; Success rate metrics: Initial registration success rate = number of successful initial registrations / number of initial registration requests; Registration update success rate = number of registration update accepted / number of registration update requests; Handover success rate (including within / between AMFs, within / between systems) = number of successful handovers / number of handover attempts; UECM (UE Context Management). Management) Registration success rate (including AMF-initiated and UDM-initiated) = Number of successful UECM registrations / Number of UECM registration requests; N11 interface session context establishment success rate = Number of successful N11 interface session context establishments / Number of N11 interface session context establishment requests; N11 interface session context update success rate = Number of successful N11 interface session context updates / Number of N11 interface session context update requests; N11 interface session context release success rate = Number of successful N11 interface session context releases / Number of N11 interface session context release requests; N11 interface session context query success rate = Number of successful N11 interface session context queries / Number of N11 interface session context query requests; Business request success rate, etc.; Latency metrics: Average initial registration time, etc.; Error type metrics: Number of authentication parameter errors, Number of authentication rejections, Number of initial registration failures, Number of registration update failures, and Number of business request rejections.

[0563] SMF: Traffic-type metrics: average / maximum number of PDU sessions and average / maximum QoS (Quality of Service) flow count, etc.; Success-type metrics: PDU session establishment success rate = number of successful PDU session establishments / number of PDU session establishment requests; SMF-initiated PDU session modification success rate = number of successful SMF-initiated PDU session modifications / number of SMF-initiated PDU session modification requests; N7 interface creates SM (Session Management). Management policy success rate = N7 interface SM policy creation success count / N7 interface SM policy creation request count; N10 interface UE context registration success rate = N10 interface UE context registration success count / N10 interface UE context registration request count; N7 interface SM policy update success rate = N7 interface SM policy update success count / N7 interface SM policy update request count; N10 interface UE context deregistration success rate = N10 interface UE context deregistration success count / N10 interface UE context deregistration request count; N7 interface SM policy deletion success rate = N7 interface SM policy deletion success count / N7 interface SM policy deletion request count, etc.; latency type indicators: average duration of PDU session establishment process, etc.; error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures, etc.

[0564] UPF: Traffic type metrics: average / maximum QoS flow count, number of GTP packets received / sent on N3 interface, number of GTP packets received / sent on N9a interface, number of bytes received / sent on N6 interface, etc.; Success rate type metrics: PFCP session establishment success rate = number of successful PFCP session establishments / number of PFCP session establishment requests, PFCP session modification success rate = number of successful PFCP session modifications / number of PFCP session modification requests, etc.; Error type metrics: number of failed PFCP session establishment / modifications, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, number of erroneous GTP packets received on N9c interface.

[0565] UDM: Success rate indicators: UECM registration success rate initiated by AMF = number of successful UECM registrations initiated by AMF / number of UECM registration requests initiated by AMF; UECM registration success rate initiated by SMF = number of successful UECM registrations initiated by SMF / number of UECM registration requests initiated by SMF; Registration parameter update success rate = number of successful registration parameter update / number of registration parameter update requests; User data acquisition success rate = number of successful user data acquisition / number of user data acquisition requests; User data subscription success rate = number of successful user data subscriptions / number of user data subscription requests; Unsubscribe user data success rate = number of successful unsubscribe user data subscriptions / number of unsubscribe user data requests, etc.

[0566] PCF: Traffic type metrics: AM (Access and Mobility Management) policy association total average / maximum, SM policy association total average / maximum, etc.; Success rate type metrics: AM policy association establishment success rate = number of successful AM policy association establishment / number of AM policy association establishment requests; AM policy association update success rate = number of successful AM policy association updates / number of AM policy association update requests; AM policy association deletion success rate = number of successful AM policy association deletions / number of AM policy association deletion requests; SM policy association establishment success rate = number of successful SM policy association establishments / number of SM policy association establishment requests; SM policy association update success rate = number of successful SM policy association updates / number of SM policy association update requests; SM policy association deletion success rate = number of successful SM policy association deletions / number of SM policy association deletion requests, etc.; Error type metrics: number of SM policy association establishment failures, number of SM policy association update failures.

[0567] NRF: Traffic-type metrics: number of storage instances, etc.; Success-type metrics: NF (Network Function) update success rate = number of successful NF update attempts / number of NF update requests; NF discovery success rate = number of successful NF discovery attempts / number of NF discovery requests, etc.; Error-type metrics: number of failed NF update attempts and number of failed NF discovery attempts, etc.

[0568] NSSF: Success rate type metric: Network slice selection success rate = Number of successful network slice selections / Number of network slice selection requests; Error type metric: Number of failed network slice selections.

[0569] When the service quality indicators (SMIs) of an NFV network are within the allowable fluctuation range, the service can be considered to be minimally affected. For example: fluctuations in success rate indicators should be less than a%, traffic volume indicators less than b%, latency indicators less than c%, and error indicators less than d. These fluctuation thresholds should be set according to the specific service requirements of the scenario.

[0570] 3. Calculate network and / or network element resilience indicators to measure the reliability of the network and / or network elements against damage.

[0571] In practice, evaluating the NFV network based on the second indicator assesses the overall network service quality's resilience to damage when components in the NFV network fail.

[0572] The technical solutions provided in this embodiment primarily consider the resilience of cloud-based networks. Therefore, network components mainly consider those within the data center, excluding those outside the data center. Based on chaos engineering verification results, an NFV network resilience index is defined, representing the overall network service quality's resilience to failure when major components in the NFV network fail. It is defined as the weighted sum of the proportions of failures in each functional layer component; a higher resilience index value indicates stronger network resilience.

[0573] In practice, the NFV network is evaluated based on the second indicator using one or a combination of the following methods:

[0574] NFV element resilience = α * (a1 * network component failure rate + a2 * compute / storage component failure rate + α) + β * OMU failure rate; where NFV element resilience is an evaluation index for assessing the ability of an NFV element to withstand service quality disruptions when components fail, α represents the user plane weight, β represents the management plane weight, and a1 * network component failure rate + β * OMU failure rate. i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0575] NFV network resilience = α * NFV network element failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability compute and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability compute and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability compute and storage component failure rate) * 100%; where NFV network resilience is an evaluation index for assessing the overall network service quality resilience against damage when NFV network elements, EMS, VNFM, and NFVO functional components fail; α represents the user plane weight, β represents the management plane weight, and b... i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0576] 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0577] 5G wireless network element resilience = (α * (a1 * ∑5G base station high availability network component failure rate + a2 * ∑5G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0578] 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0579] 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0580] 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0581] 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0582] 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0583] 5G core network resilience = (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0584] In practice, the resilience of wireless network elements, the resilience of core network elements, or a combination thereof can also be assessed based on the second indicator in the following ways:

[0585] Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; α and β are set to 0.5, and a1 and a2 are recommended to be set to 1;

[0586] Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0587] In practice, it can be implemented as follows:

[0588] 1. Wireless network element resilience.

[0589] Communication network health dimension: reliability.

[0590] Business requirement: To verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of network elements. It is recommended to perform fault injection monthly.

[0591] Indicator definition: In base station fault injection scenarios, the fluctuation of key performance indicators of the base station is within an acceptable range (e.g., 1% to 3%). Network element resilience is defined as the weighted sum of the proportion of faults occurring in each functional layer component.

[0592] Key performance indicators include: traffic type indicators, success rate type indicators, and error type indicators.

[0593] Calculation formula: (α*(a1*network component failure rate + a2*computing / storage component failure rate) + β*OMU failure rate)*100%, where α, β, a1, and a2 range from [0, 1].

[0594] Data type: Real number.

[0595] Data unit: %.

[0596] Measurement object: network element.

[0597] Statistical period: six months

[0598] Data source: Network administrator statistics.

[0599] 2. Wireless network resilience.

[0600] Communication network health dimension: reliability.

[0601] Business requirement: To verify the cloud-based scalability and fault tolerance of the wireless network service, as well as the impact of maximum uncertainty on network stability. It is recommended to perform fault injection monthly.

[0602] Indicator definition: In each base station fault injection scenario, the fluctuation of each key performance indicator of the wireless network is within an acceptable range (e.g., 1% to 3%). Network resilience is defined as the weighted sum of the proportion of failures of base stations and EMS functional layer components.

[0603] Key performance indicators include: the average traffic type indicators of each network element in the overall network are averaged, the maximum traffic type indicators of each network element are maximum values, the overall network success rate type indicators are averaged, and the overall network error type indicators are summed.

[0604] The calculation formula is: (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0605] Data type: Real number.

[0606] Data unit: %.

[0607] Measurement object: Network.

[0608] Statistical period: six months

[0609] Data source: Network administrator statistics.

[0610] 3. Core network element resilience.

[0611] Communication network health dimension: reliability.

[0612] Business requirement: To verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements. It is recommended to perform fault injection monthly.

[0613] Metric Definition: In core network element fault injection scenarios, the fluctuation of key performance indicators of core network elements is within an acceptable range (e.g., 1% to 3%). Core network element resilience is defined as the weighted sum of the percentage of faults occurring in components across all functional layers. The key performance indicators of core network elements are as follows:

[0614] Traffic type metrics, success rate type metrics, and error type.

[0615] The calculation formula is: (α*(a1*∑failure rate of high-availability network components in core network elements + a2*∑failure rate of high-availability computing and storage components in core network elements) + β*OMU failure rate)*100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, a1, and a2 range from [0, 1].

[0616] Data type: Real number.

[0617] Data unit: %.

[0618] Measurement object: network element.

[0619] Statistical period: six months

[0620] Data source: Network administrator statistics.

[0621] 4. Core network resilience.

[0622] Communication network health dimension: reliability.

[0623] Business requirement: To verify the cloud-based scalability and fault tolerance of core network services, as well as the impact of maximum uncertainty on network stability. Fault injection is recommended monthly.

[0624] Metric Definitions: In core network fault injection scenarios, the fluctuation of key core network performance indicators is within an acceptable range (e.g., 1% to 3%). Core network resilience is defined as the weighted sum of the percentage of failures occurring in core network elements, EMS, VNFM, and NFVO. Key Network Performance Indicators: The average traffic type indicators for each network element in the overall network are averaged, the maximum traffic type indicators for each network element are the maximum values, the overall network success rate type indicators are averaged, and the overall network error type indicators are summed.

[0625] Calculation formula: (α*∑core network element failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5 and b6 are in the range of [0, 1].

[0626] Data type: Real number.

[0627] Data unit: %.

[0628] Measurement object: Network.

[0629] Statistical period: six months

[0630] Data source: Network administrator statistics.

[0631] 5. 4G wireless network element resilience.

[0632] Communication network health dimension: reliability.

[0633] Business requirement: To verify the cloud-based scalability and fault tolerance of 4G wireless network element services, as well as the impact of maximum uncertainty on the steady state of network elements. It is recommended to perform fault injection monthly.

[0634] Indicator definition: In base station fault injection scenarios, the fluctuation of key performance indicators of the base station is within an acceptable range (e.g., 1% to 3%). Network element resilience is defined as the weighted sum of the proportion of faults occurring in each functional layer component of the base station. Key performance indicators include: Service volume type indicators: number of bytes sent / received by the eNB Ethernet interface, number of service bytes sent / received by the eNB S1 interface, number of uplink / downlink PDCP SDU bytes on the cell user plane / control plane, etc.; Success rate type indicators: service-related E-RAB establishment success rate, service-related radio connection success rate, RRC connection establishment success rate, handover success rate (including intra-eNB / inter-eNB, same-frequency / different-frequency, handover in / handover out, intra-system / inter-system), etc.; Latency type indicators: average / maximum establishment time of RRC connection, average / maximum establishment time of E-RAB, etc.; Error type indicators: radio call drop rate (including conversational voice services and conversational live video streaming services), call drop rate (including real-time gaming services, non-conversational buffered video streaming services, IMS signaling services), and TCP-based radio call drop rate, uplink / downlink PDCP SDU packet loss rate, number of uplink / downlink transport block errors, and number of paging record drops, etc.

[0635] The calculation formula is: (α*(a1*∑4G base station high-availability network component failure rate + a2*∑4G base station high-availability computing and storage component failure rate) + β*OMU failure rate)*100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. Taking NICs as an example, the calculation method for the failure rate of high-availability components is: High-availability NIC failure rate = Number of faulty high-availability NICs / Total number of high-availability NICs, where ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, a1, and a2 range from [0, 1].

[0636] Data type: Real number.

[0637] Data unit: %.

[0638] Measurement object: network element.

[0639] Statistical period: six months

[0640] Data source: Network administrator statistics.

[0641] 6. 4G wireless network resilience.

[0642] Communication network health dimension: reliability.

[0643] Business requirement: To verify the cloud-based scalability and fault tolerance of 4G wireless network services, as well as the impact of maximum uncertainty on network stability. Fault injection is recommended monthly.

[0644] Indicator Definition: In base station fault injection scenarios, the overall key performance indicators of available base stations in the wireless network fluctuate within an acceptable range (e.g., 1% to 3%). Network resilience is defined as the weighted sum of the percentage of base station and EMS functional component failures. Key performance indicators include: Service volume type indicators: ∑ eNB Ethernet interface transmitted / received bytes, ∑ eNB S1 interface transmitted / received service bytes, ∑ cell user plane / control plane uplink / downlink PDCP. SDU byte count, etc.; Success rate indicators: AVERAGE (service-related E-RAB establishment success rate), AVERAGE (service-related radio connection success rate), AVERAGE (RRC connection establishment success rate), AVERAGE [handover success rate (including intra-eNB / inter-eNB, same-frequency / different-frequency, handover in / handover out, intra-system / inter-system)], etc.; Latency indicators: AVERAGE (RRC connection average establishment time), MAX (RRC connection maximum establishment time), AVERAGE (E-RAB average establishment time), MAX (E-RAB maximum establishment time), etc.; Error type indicators: AVERAGE [radio call drop rate (including conversational voice services and conversational live video streaming services)], AVERAGE [disconnection rate (including real-time gaming services, non-conversational buffered video streaming services, IMS signaling services)], AVERAGE (TCP-based radio disconnection rate), AVERAGE (uplink / downlink PDCP) The table includes parameters such as SDU packet loss rate, ∑ uplink / downlink transmission block errors, and ∑ paging record discards, where ∑ represents the summation of all base station metrics, AVERAGE represents the average of all base station metrics, and MAX represents the maximum value of all base station metrics.

[0645] The calculation formula is: (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. Taking NICs as an example, the calculation method for the failure rate of high-availability components is: High-availability NIC failure rate = Number of faulty high-availability NICs / Total number of high-availability NICs. ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0646] Data type: Real number.

[0647] Data unit: %.

[0648] Measurement object: Network.

[0649] Statistical period: six months

[0650] Data source: Network administrator statistics.

[0651] 7. 5G wireless network element resilience.

[0652] Communication network health dimension: reliability.

[0653] Business requirement: To verify the cloud-based scalability and fault tolerance of 5G wireless network element services, as well as the impact of maximum uncertainty on the steady state of network elements. It is recommended to perform fault injection monthly.

[0654] Indicator Definition: In base station fault injection scenarios, the fluctuations of key performance indicators of wireless network elements are within an acceptable range (e.g., 1% to 3%). Network element resilience is defined as the weighted sum of the percentage of base station functional layer component failures. Key performance indicators include: traffic volume and capacity type indicators: gNB data volume sent / received from the NG interface (SA networking only), cell user plane uplink / downlink PDCP PDU bytes (NSA & SA networking), gNB data volume received / sent from the S1 interface (NSA networking only), uplink / downlink transmission TB (NSA & SA networking), downlink PDSCH PRB availability (NSA & SA networking), paging record reception count (SA networking only), average / maximum number of RRC connections in dual-connectivity, and NSA SCG Split. Bear type average / maximum ERAB count, success rate indicators: RRC connection establishment success rate (SA network only) = number of successful RRC connection establishments / number of RRC connection establishment requests; Flow establishment success rate (SA network only) = number of successful Flow establishments / number of Flow establishment requests; PDUSESSION establishment success rate (SA network only) = number of successful PDUSESSION establishments / number of PDUSESSION establishment requests; Handover success rate (including same frequency / different frequency, within gNB / between gNBs, handover out / handover in) = number of successful handovers / number of handover requests; NG interface UE-related logical signaling connection establishment success rate (SA network only) = NG interface UE... E-related logical signaling connection establishment success count / NG interface UE-related logical signaling connection establishment request count, etc. Latency type indicators: RRC connection average / maximum establishment time, RLC downlink data packet average processing latency (SA network only), gNB inter-NG handover average time (SA network only), gNB inter-Xn handover average time (SA network only), Epsfallback service handover latency from 5G to 4G (SA network only), RLC downlink data packet average processing latency per slice cell (SA network only), etc. Error type indicators: Flow establishment / modification failure count (SA network only), uplink / downlink PDCP packet loss count (NSA & SA network only), etc.

[0655] The calculation formula is: (α*(a1*∑5G base station high-availability network component failure rate + a2*∑5G base station high-availability computing and storage component failure rate) + β*OMU failure rate)*100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. Taking NICs as an example, the calculation method for the failure rate of high-availability components is: High-availability NIC failure rate = Number of faulty high-availability NICs / Total number of high-availability NICs, where ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, a1, and a2 range from [0, 1].

[0656] Data type: Real number.

[0657] Data unit: %.

[0658] Measurement object: network element.

[0659] Statistical period: six months

[0660] Data source: Network administrator statistics.

[0661] 8. 5G wireless network resilience.

[0662] Communication network health dimension: reliability.

[0663] Business requirement: To verify the cloud-based scalability and fault tolerance of 5G wireless network services, as well as the impact of maximum uncertainty on network steady state. It is recommended to perform fault injection monthly.

[0664] Indicator definition: In each base station fault injection scenario, the overall key performance indicators of the wireless network fluctuate within an acceptable range (e.g., 1% to 3%). Network resilience is defined as the weighted sum of the percentage of available base stations and EMS functional components that fail.Key performance indicators include: Service volume type indicators: ∑gNB sends / receives service data volume from the NG interface (SA networking only), ∑cell user plane uplink / downlink PDCP PDU bytes (NSA & SA networking), ∑gNB receives / sends service data volume from the S1 interface (NSA networking only), ∑uplink / downlink transmission TB (NSA & SA networking), AVERAGE [Number of paging records received (SA networking only)], AVERAGE [Average number of RRC connections in dual-connectivity], MAX [Maximum number of RRC connections in dual-connectivity], AVERAGE [Average number of ERABs for NSA SCG Split Bear type], MAX [Number of ERABs for NSA SCG Split Bear type]. [Maximum number of ERABs for Bear type], etc. Success rate indicators: AVERAGE [RRC connection establishment success rate (SA network only) = number of successful RRC connection establishments / number of RRC connection establishment requests], AVERAGE [Flow establishment success rate (SA network only) = number of successful Flow establishments / number of Flow establishment requests], AVERAGE [PDUSESSION establishment success rate (SA network only) = number of successful PDUSESSION establishments / number of PDUSESSION establishment requests], AVERAGE [Handover success rate (including intra-frequency / inter-frequency, intra-gNB / inter-gNB, handover out / handover in) = number of successful handovers / number of handover requests], AVERAGE [NG interface UE-related logical signaling connection establishment success rate (SA network only) = number of successful NG interface UE-related logical signaling connection establishments / number of NG interface UE-related logical signaling connection establishment requests] [Number of times], etc. Latency type indicators: AVERAGE (average RRC connection duration), MAX (maximum establishment duration), AVERAGE [average RLC downlink packet processing latency (only applicable to SA network)], AVERAGE [average NG handover latency between gNBs (only applicable to SA network)], AVERAGE [average Xn handover latency between gNBs (only applicable to SA network)], AVERAGE [epsfallback service latency from 5G to 4G (only applicable to SA network)], AVERAGE [average RLC downlink packet processing latency per slice cell (only applicable to SA network)], etc. Error type indicators: ∑Flow establishment / modification failure count (only applicable to SA network), ∑Uplink / downlink PDCP packet loss count (applicable to NSA & SA network), etc., where ∑ represents the summation of indicators for each base station, AVERAGE represents the average of indicators for each base station, and MAX represents the maximum value of indicators for each base station.

[0665] The calculation formula is: (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. Taking NICs as an example, the calculation method for the failure rate of high-availability components is: High-availability NIC failure rate = Number of faulty high-availability NICs / Total number of high-availability NICs. ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0666] Data type: Real number.

[0667] Data unit: %.

[0668] Measurement object: Network.

[0669] Statistical period: six months

[0670] Data source: Network administrator statistics.

[0671] 9. 4G core network element resilience.

[0672] Communication network health dimension: reliability.

[0673] Business requirement: To verify the cloud-based scalability and fault tolerance of 4G core network element services, as well as the impact of maximum uncertainty on the steady state of network elements. It is recommended to perform fault injection monthly.

[0674] Indicator Definition: In 4G core network element fault injection scenarios, the fluctuation of key performance indicators of core network elements is within an acceptable range (e.g., 1% to 3%). Core network element resilience is defined as the weighted sum of the percentage of failures in each functional layer component of the network element. The key performance indicators for different 4G core network elements are as follows:

[0675] MME: Traffic type indicators: average / maximum number of MME bearers, number of users in idle / connected state of MME, average / maximum number of attached users of MME, etc.; Success rate type indicators: EPS attachment success rate, authentication success rate, DNS resolution success rate initiated by MME, default / dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate (within / between MMEs, within / between systems), handover success rate (within / between MMEs, within / between systems), etc.; Latency type indicators: average / maximum attachment duration and average / maximum dedicated bearer establishment duration, etc.; Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures, etc.

[0676] ServingGW: Traffic and capacity type metrics: SGW average / peak capacity utilization, user plane uplink / downlink traffic, average GTP uplink / downlink traffic generated per attached user in ServingGW, average / maximum number of attached users in ServingGW, average / maximum number of bearers in ServingGW, and SGW data throughput capacity utilization, etc.; Success rate type metrics: ServingGW default / dedicated bearer establishment success rate, etc.

[0677] PGW: Traffic and capacity type indicators: PGW data throughput capacity utilization, PGW average / peak capacity utilization, PGW S5 / S8 interface uplink / downlink traffic, PGW average / maximum number of attached users, PGW average / peak number of bearers, SGi interface receive / send traffic, etc.; Success rate type indicators: dedicated bearer establishment success rate and CDR transmission success rate, etc.; Latency type indicators: average / maximum duration of dedicated bearer establishment initiated by PGW, etc.; Error type indicators: number of GTP packets dropped due to errors on PGW S5 / S8 interface and number of IP packets dropped due to errors on SGi interface, etc.

[0678] HSS: Service volume and capacity type indicators: number of HSS number / active users, HSS authentication capacity utilization rate and HSS static capacity utilization rate, etc.; success rate type indicators: success rate of HSS authentication information query, success rate of HSS location update / cancellation, success rate of HSS user data insertion / deletion, and success rate of HSS UE clearing, etc.; error type indicators: number of location update failures, etc.

[0679] PCRF: Traffic-type metrics: average / peak utilization of Gx session processing capacity, etc.; Success-type metrics: success rate of policy control initiation / update / end, success rate of re-authentication, and success rate of application session authorization, etc.; Error-type metrics: application session call loss rate, etc.

[0680] The calculation formula is: (α*(a1*∑4G core network element high-availability network component failure rate + a2*∑4G core network element high-availability computing and storage component failure rate) + β*OMU failure rate)*100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. Taking NICs as an example, the calculation method for the failure rate of high-availability components is: High-availability NIC failure rate = Number of faulty high-availability NICs / Total number of high-availability NICs. ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, a1, and a2 range from [0, 1].

[0681] Data type: Real number.

[0682] Data unit: %.

[0683] Measurement object: network element.

[0684] Statistical period: six months

[0685] Data source: Network administrator statistics.

[0686] 10. 4G core network resilience.

[0687] Communication network health dimension: reliability.

[0688] Business requirement: To verify the cloud-based scalability and fault tolerance of 4G core network services, as well as the impact of maximum uncertainty on network stability. Fault injection is recommended monthly.

[0689] Indicator Definition: In 4G core network fault injection scenarios, the overall key performance indicators of the core network fluctuate within an acceptable range (e.g., 1% to 3%). Core network resilience is defined as the weighted sum of the percentage of failures among available core network elements, EMS, VNFM, and NFVO functional components. The key performance indicators for different 4G core network elements are as follows:

[0690] MME: Traffic volume metrics: AVERAGE (average number of MME bearers), MAX (maximum number of MME bearers), AVERAGE (number of idle / connected users in MME), AVERAGE (average number of attached users in MME), MAX (maximum number of attached users), etc.; Success rate metrics: AVERAGE (EPS attach success rate), AVERAGE (authentication success rate), AVERAGE (DNS resolution success rate initiated by MME), AVERAGE (default / dedicated bearer activation success rate), AVERAGE (PDN connection establishment success rate), AVERAGE... (Service request success rate), AVERAGE (paging success rate), AVERAGE [tracking area update success rate (within / between MMEs, within / between systems)], AVERAGE [handover success rate (within / between MMEs, within / between systems)], etc.; latency type indicators: AVERAGE (average attachment time), MAX (maximum attachment time), AVERAGE (average dedicated bearer establishment time), MAX (maximum dedicated bearer establishment time), etc.; error type indicators: ∑ number of authentication parameter errors, ∑ number of UE authentication failures, ∑ number of EPS attachment failures, and ∑ number of tracking area update failures, etc.

[0691] ServingGW: Traffic and capacity type metrics: AVERAGE (SGW average capacity utilization), MAX (SGW peak capacity utilization), ∑ (user surface uplink / downlink traffic), AVERAGE (average GTP uplink / downlink traffic generated per attached user in ServingGW), AVERAGE (average number of attached users in ServingGW), MAX (maximum number of attached users in ServingGW), AVERAGE (average number of bearers in ServingGW), MAX (maximum number of bearers in ServingGW), and AVERAGE (SGW data throughput capacity utilization), etc. Success rate type metrics: AVERAGE (ServingGW default / dedicated bearer establishment success rate), etc.

[0692] PGW: Traffic and capacity type indicators: AVERAGE (PGW data throughput capacity utilization), AVERAGE (PGW bearer capacity average utilization), MAX (PGW bearer capacity peak utilization), ∑ (PGW S5 / S8 interface uplink / downlink traffic), AVERAGE (PGW average number of attached users), MAX (PGW maximum number of attached users), AVERAGE (PGW average number of bearers), MAX (PGW peak number of bearers), ∑ (SGi interface receive / send traffic), etc.; Success rate type indicators: AVERAGE (dedicated bearer establishment success rate) and AVERAGE (CDR transmission success rate), etc.; Latency type indicators: AVERAGE (PGW-initiated dedicated bearer establishment average time), MAX (PGW-initiated dedicated bearer establishment maximum time), etc.; Error type indicators: ∑ number of GTP packets dropped by errors on PGW S5 / S8 interface and ∑ number of IP packets dropped by errors on SGi interface, etc.

[0693] HSS: Business volume and capacity type indicators: ∑HSS number of active users and AVERAGE (HSS authentication capacity utilization rate), AVERAGE (HSS static capacity utilization rate), etc.; Success rate type indicators: AVERAGE (HSS authentication information query success rate), AVERAGE (HSS update / cancel location success rate), AVERAGE (HSS insert / delete user data success rate), and AVERAGE (HSS clear UE success rate), etc.; Error type indicators: ∑ number of failed location updates, etc.

[0694] PCRF: Capacity type indicators: AVERAGE (average utilization of Gx session processing capacity), MAX (peak utilization of Gx session processing capacity), etc.; Success rate type indicators: AVERAGE (success rate of policy control initiation / update / end), AVERAGE (success rate of re-authentication), and AVERAGE (success rate of application session authorization), etc.; Error type indicators: AVERAGE (call loss rate of application session), etc., where ∑ represents the summation of indicators of each network element, AVERAGE represents the average of indicators of each network element, and MAX represents the maximum value of indicators of each network element.

[0695] The calculation formula is: (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high-availability network component failure rate + b2*∑EMS high-availability computing and storage component failure rate + b3*∑VNMF high-availability network component failure rate + b4*∑VNFM high-availability computing and storage component failure rate + b5*∑NFVO high-availability network component failure rate + b6*∑NFVO high-availability computing and storage component failure rate +)*100%, where high-availability components include network components (TOR switches, EOR switches, and network cards, etc.) and computing and storage components (physical machines, virtual machines, and databases, etc.). Taking network cards as an example, the high-availability network card failure rate is calculated as: High-availability network card failure rate = Number of faulty high-availability network cards / Total number of high-availability network cards. ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 range from [0, 1].

[0696] Data type: Real number.

[0697] Data unit: %.

[0698] Measurement object: Network.

[0699] Statistical period: six months

[0700] Data source: Network administrator statistics.

[0701] 11. 5G core network element resilience.

[0702] Communication network health dimension: reliability.

[0703] Business requirement: To verify the cloud-based scalability and fault tolerance of 5G core network element services, as well as the impact of maximum uncertainty on the steady state of network elements. It is recommended to perform fault injection monthly.

[0704] Indicator Definition: In 5G core network element fault injection scenarios, the fluctuation of key performance indicators of core network elements is within an acceptable range (e.g., 1% to 3%). Core network element resilience is defined as the weighted sum of the percentage of faults occurring in each functional layer component of the network element. The key performance indicators for different 5G core network elements are as follows:

[0705] AMF: Traffic volume metrics: Number of AMF registered / idle users, number of paging requests, number of first paging responses, and number of second paging responses, etc.; Success rate metrics: Initial registration success rate = number of successful initial registrations / number of initial registration requests; Registration update success rate = number of registration update accepted / number of registration update requests; Handover success rate (including within / between AMFs, within / between systems) = number of successful handovers / number of handover attempts; UECM deregistration success rate (including AMF-initiated and UDM-initiated) = number of successful UECM deregistrations / number of UECM deregistration requests; N11 interface session context establishment success rate = number of successful N11 interface session context establishments / N11 The metrics include: number of API session context establishment requests, N11 API session context update success rate (Number of successful N11 API session context updates / Number of N11 API session context update requests), N11 API session context release success rate (Number of successful N11 API session context releases / Number of N11 API session context release requests), N11 API session context query success rate (Number of successful N11 API session context queries / Number of N11 API session context query requests), business request success rate, latency metrics such as average initial registration time, and error type metrics such as the number of authentication parameter errors, the number of authentication rejections, the number of initial registration failures, the number of registration update failures, and the number of business request rejections.

[0706] SMF: Traffic volume type metrics: average / maximum number of PDU sessions and average / maximum number of QoS flows, etc.; Success rate type metrics: PDU session establishment success rate = number of successful PDU session establishments / number of PDU session establishment requests; SMF-initiated PDU session modification success rate = number of successful SMF-initiated PDU session modifications / number of SMF-initiated PDU session modification requests; N7 interface SM policy creation success rate = number of successful N7 interface SM policy creations / number of N7 interface SM policy creation requests; N10 interface UE context registration success rate = number of successful N10 interface UE context registrations / number of N10 interface... UE context registration request count, N7 interface SM policy update success rate = N7 interface SM policy update success count / N7 interface SM policy update request count, N10 interface UE context deregistration success rate = N10 interface UE context deregistration success count / N10 interface UE context deregistration request count, N7 interface SM policy deletion success rate = N7 interface SM policy deletion success count / N7 interface SM policy deletion request count, etc.; latency type indicators: average duration of PDU session establishment process, etc.; error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures, etc.

[0707] UPF: Traffic type metrics: average / maximum QoS flow count, number of GTP packets received / sent on N3 interface, number of GTP packets received / sent on N9a interface, number of bytes received / sent on N6 interface, etc.; Success rate type metrics: PFCP session establishment success rate = number of successful PFCP session establishments / number of PFCP session establishment requests, PFCP session modification success rate = number of successful PFCP session modifications / number of PFCP session modification requests, etc.; Error type metrics: number of failed PFCP session establishment / modifications, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, number of erroneous GTP packets received on N9c interface.

[0708] UDM: Success rate indicators: UECM registration success rate initiated by AMF = number of successful UECM registrations initiated by AMF / number of UECM registration requests initiated by AMF; UECM registration success rate initiated by SMF = number of successful UECM registrations initiated by SMF / number of UECM registration requests initiated by SMF; Registration parameter update success rate = number of successful registration parameter update / number of registration parameter update requests; User data acquisition success rate = number of successful user data acquisition / number of user data acquisition requests; User data subscription success rate = number of successful user data subscriptions / number of user data subscription requests; Unsubscribe user data success rate = number of successful unsubscribe user data subscriptions / number of unsubscribe user data requests, etc.

[0709] PCF: Business volume type indicators: AM strategy total number average / maximum, SM strategy total number average / maximum, etc.; Success rate type indicators: AM strategy association establishment success rate = number of successful AM strategy association establishment / number of AM strategy association establishment requests; AM strategy association update success rate = number of successful AM strategy association updates / number of AM strategy association update requests; AM strategy association deletion success rate = number of successful AM strategy association deletions / number of AM strategy association deletion requests; SM strategy association establishment success rate = number of successful SM strategy association establishments / number of SM strategy association establishment requests; SM strategy association update success rate = number of successful SM strategy association updates / number of SM strategy association update requests; SM strategy association deletion success rate = number of successful SM strategy association deletions / number of SM strategy deletion requests, etc.; Error type indicators: number of failed SM strategy association establishments; number of failed SM strategy association updates.

[0710] NRF: Business volume type metrics: number of storage instances, etc.; Success rate type metrics: NF update success rate = number of successful NF update attempts / number of NF update requests; NF discovery success rate = number of successful NF discovery attempts / number of NF discovery requests, etc.; Error type metrics: number of failed NF update attempts and number of failed NF discovery attempts, etc.

[0711] NSSF: Success rate type metric: Network slice selection success rate = Number of successful network slice selections / Number of network slice selection requests; Error type metric: Number of failed network slice selections.

[0712] The calculation formula is: (α*(a1*∑5G core network element high-availability network component failure rate + a2*∑5G core network element high-availability computing and storage component failure rate) + β*OMU failure rate)*100%, where high-availability network components include TOR switches, EOR switches, and network interface cards (NICs), and high-availability computing and storage components include physical machines, virtual machines, and databases. Taking NICs as an example, the calculation method for the failure rate of high-availability components is: High-availability NIC failure rate = Number of faulty high-availability NICs / Total number of high-availability NICs. ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, a1, and a2 range from [0, 1].

[0713] Data type: Real number.

[0714] Data unit: %.

[0715] Measurement object: network element.

[0716] Statistical period: six months

[0717] Data source: Network administrator statistics.

[0718] 12. 5G core network resilience.

[0719] Communication network health dimension: reliability.

[0720] Business requirement: To verify the cloud-based scalability and fault tolerance of 5G core network services, as well as the impact of maximum uncertainty on network stability. Fault injection is recommended monthly.

[0721] Indicator Definition: In 5G core network fault injection scenarios, the fluctuation of key performance indicators of the core network is within an acceptable range (e.g., 1% to 3%). Core network resilience is defined as the weighted sum of the percentage of failures of available core network elements, EMS, VNFM, and NFVO functional components. The key performance indicators for different 5G core network elements are as follows:

[0722] AMF: Traffic volume type metrics: ∑AMF registered / idle user count, ∑paging request count, ∑first paging response count, and ∑second paging response count, etc.; Success rate type metrics: AVERAGE (Initial registration success rate = initial registration success count / initial registration request count), AVERAGE (Registration update success rate = registration update acceptance count / registration update request count), AVERAGE [Handover success rate (including AMF intra / inter-AMF, intra-system / inter-system) = handover success count / handover attempt count], AVERAGE [UECM deregistration success rate (including AMF-initiated and UDM-initiated) = UECM deregistration success count / UECM deregistration request count], AVERAGE (N11 interface session context establishment success rate = N11 interface session context establishment success count / N11 interface... The metrics include: session context establishment request count, AVERAGE (N11 interface session context update success rate = N11 interface session context update success count / N11 interface session context update request count), AVERAGE (N11 interface session context release success rate = N11 interface session context release success count / N11 interface session context release request count), AVERAGE (N11 interface session context query success rate = N11 interface session context query success count / N11 interface session context query request count), AVERAGE (business request success rate), etc.; latency metrics include: AVERAGE (average initial registration time), etc.; and error type metrics include: ∑ number of authentication parameter errors, ∑ number of authentication rejections, ∑ number of initial registration failures, ∑ number of registration update failures, and ∑ number of business request rejections.

[0723] SMF: Traffic volume type metrics: AVERAGE (average number of PDU sessions), MAX (maximum number of PDU sessions), AVERAGE (average number of QoS flows), MAX (maximum number of QoS flows), etc.; Success rate type metrics: AVERAGE (PDU session establishment success rate = number of successful PDU session establishments / number of PDU session establishment requests), AVERAGE (SMF-initiated PDU session modification success rate = number of successful SMF-initiated PDU session modifications / number of SMF-initiated PDU session modification requests), AVERAGE (N7 interface SM policy creation success rate = number of successful N7 interface SM policy creations / number of N7 interface SM policy creation requests), AVERAGE (N10 interface UE context registration success rate = number of successful N10 interface UE context registrations = number of successful N10 interface UE context registrations), AVERAGE (N10 interface UE context registration success rate = number of successful N10 interface UE context registrations), etc. The metrics include: UE context registration success rate / N10 interface UE context registration request rate; AVERAGE (N7 interface SM policy update success rate = N7 interface SM policy update success rate / N7 interface SM policy update request rate); AVERAGE (N10 interface UE context deregistration success rate = N10 interface UE context deregistration success rate / N10 interface UE context deregistration request rate); AVERAGE (N7 interface SM policy deletion success rate = N7 interface SM policy deletion success rate / N7 interface SM policy deletion request rate); latency type metrics: AVERAGE (PDU session establishment process average duration); error type metrics: ∑PDU session establishment failure rate; ∑SMF-initiated PDU session modification failure rate; etc.

[0724] UPF: Traffic type metrics: AVERAGE (average QoS flow count), MAX (maximum QoS flow count), ∑N3 interface received / sent GTP packet bytes, ∑N9a interface received / sent GTP packet bytes, ∑N6 interface received / sent bytes, etc.; Success rate type metrics: AVERAGE (PFCP session establishment success rate = number of successful PFCP session establishments / number of PFCP session establishment requests), AVERAGE (PFCP session modification success rate = number of successful PFCP session modifications / number of PFCP session modification requests), etc.; Error type metrics: ∑PFCP session establishment / modification failures, ∑N3 interface received erroneous GTP packets, ∑N6 interface dropped erroneous IP packets, ∑N9c interface received erroneous GTP packets.

[0725] UDM: Success rate metrics: AVERAGE (UECM registration success rate initiated by AMF = number of successful UECM registrations initiated by AMF / number of UECM registration requests initiated by AMF), AVERAGE (UECM registration success rate initiated by SMF = number of successful UECM registrations initiated by SMF / number of UECM registration requests initiated by SMF), AVERAGE (registration parameter update success rate = number of successful registration parameter update / number of registration parameter update requests), AVERAGE (user data retrieval success rate = number of successful user data retrieval / number of user data retrieval requests), AVERAGE (user data subscription success rate = number of successful user data subscription / number of user data subscription requests), AVERAGE (user data unsubscription success rate = number of successful unsubscription / number of user data unsubscription requests), etc.

[0726] PCF: Business volume type indicators: AVERAGE (average of total AM policy associations), MAX (maximum of total AM policy associations), AVERAGE (average of total SM policy associations), MAX (maximum of total SM policy associations), etc.; Success rate type indicators: AVERAGE (AM policy association establishment success rate = number of successful AM policy association establishments / number of AM policy association establishment requests), AVERAGE (AM policy association update success rate = number of successful AM policy association updates / number of AM policy association update requests), AVERAGE (AM policy association deletion success rate = number of successful AM policy association deletions / number of AM policy association deletion requests), AVERAGE (SM policy association establishment success rate = number of successful SM policy association establishments / number of SM policy association establishment requests), AVERAGE (SM policy association update success rate = number of successful SM policy association updates / number of SM policy association update requests), AVERAGE (SM policy association deletion success rate = number of successful SM policy association deletions / number of SM policy association deletion requests), etc.; Error type indicators: ∑ number of failed SM policy association establishments, ∑ number of failed SM policy association updates.

[0727] NRF: Business volume type metrics: ∑ number of storage instances, etc.; Success rate type metrics: AVERAGE (NF update success rate = number of successful NF updates / number of NF update requests), AVERAGE (NF discovery success rate = number of successful NF discovery / number of NF discovery requests), etc.; Error type metrics: ∑ number of failed NF updates and ∑ number of failed NF discovery, etc.

[0728] NSSF: Success rate type metric: AVERAGE (Network slice selection success rate = Number of successful network slice selections / Number of network slice selection requests), Error type metric: ∑ Number of failed network slice selections.

[0729] Where ∑ represents the summation of the indicators for each network element and / or slice, AVERAGE represents the average of the indicators for each network element and / or slice, and MAX represents the maximum value of the indicators for each network element and / or slice.

[0730] The calculation formula is: (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high-availability network component failure rate + b2*∑EMS high-availability computing and storage component failure rate + b3*∑VNMF high-availability network component failure rate + b4*∑VNFM high-availability computing and storage component failure rate + b5*∑NFVO high-availability network component failure rate + b6*∑NFVO high-availability computing and storage component failure rate +)*100%, where high-availability components include network components (TOR switches, EOR switches, and network cards, etc.) and computing and storage components (physical machines, virtual machines, and databases, etc.). Taking network cards as an example, the high-availability network card failure rate is calculated as: High-availability network card failure rate = Number of faulty high-availability network cards / Total number of high-availability network cards. ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 range from [0, 1].

[0731] Data type: Real number.

[0732] Data unit: %.

[0733] Measurement object: Network.

[0734] Statistical period: six months

[0735] Data source: Network administrator statistics.

[0736] Based on the same inventive concept, this invention also provides a network evaluation device and a computer-readable storage medium. Since the principles by which these devices solve problems are similar to those of the network evaluation method, the implementation of these devices can be referred to the implementation of the method, and repeated details will not be repeated.

[0737] When implementing the technical solutions provided in the embodiments of the present invention, they can be implemented in the following manner.

[0738] Figure 6 The diagram shows the structure of a network evaluation device. The device includes:

[0739] Processor 600 is used to read the program from memory 620 and execute the following procedures:

[0740] Identify the metrics that need to be evaluated in the NFV network;

[0741] Monitor the aforementioned indicators and collect the first indicator of the aforementioned indicators before fault injection;

[0742] Injection failure;

[0743] Monitor the aforementioned indicator and collect a second indicator of the aforementioned indicator after fault injection;

[0744] Restoring the NFV network to the aforementioned metric is the first metric;

[0745] Evaluate NFV networks based on the second metric;

[0746] Transceiver 610 is used to receive and send data under the control of processor 600.

[0747] During implementation, it further includes:

[0748] When a fault is injected, the predetermined indicators of the NFV network are controlled to fluctuate within a preset range.

[0749] In practice, the predetermined indicator is one or a combination of the following indicators:

[0750] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0751] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0752] One or a combination of the following indicators on 4G core network elements:

[0753] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0754] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0755] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0756] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0757] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0758] One or a combination of the following metrics on 5G core network elements:

[0759] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0760] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0761] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0762] Success rate metrics on UDM;

[0763] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0764] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0765] Success rate type metrics and error type metrics on NSSF.

[0766] In practice, the predetermined indicator is one or a combination of the following indicators:

[0767] One or a combination of the following metrics on 4G base stations:

[0768] Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of uplink and / or downlink PDCP SDU bytes on the cell user plane and / or control plane;

[0769] Success rate metrics include: E-RAB establishment success rate (related to services), wireless connection success rate (related to services), RRC connection establishment success rate, and handover success rate.

[0770] Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time;

[0771] Error type metrics: wireless call drop rate, call drop rate, TCP-based wireless call drop rate, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops;

[0772] One or a combination of the following metrics on 5G base stations:

[0773] Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP PDU bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transmission TB, downlink PDSCH PRB available number, paging record received number, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of ERABs for NSA SCG Split Bear type;

[0774] Success rate metrics include: RRC connection establishment success rate, Flow establishment success rate, PDUSESSION establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate.

[0775] Latency type indicators: RRC connection average and / or maximum setup time, RLC downlink data packet average processing latency, gNB inter-NG handover average time, gNB inter-Xn handover average time, Epsfallback service handover latency from 5G to 4G, and RLC downlink data packet average processing latency per slice cell;

[0776] Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses;

[0777] One or a combination of the following indicators on 4G core network elements:

[0778] On MME:

[0779] Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs;

[0780] Success rate metrics include: EPS attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate.

[0781] Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time;

[0782] Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures;

[0783] On ServingGW:

[0784] Traffic volume and capacity type metrics: SGW average and / or peak capacity utilization, user plane uplink and / or downlink traffic, average GTP uplink and / or downlink traffic per attached user of ServingGW, average and / or maximum number of attached users of ServingGW, average and / or maximum number of bearers of ServingGW, SGW data throughput capacity utilization;

[0785] Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate;

[0786] On PGW:

[0787] Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic.

[0788] Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate;

[0789] Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW;

[0790] Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface;

[0791] On HSS:

[0792] Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate;

[0793] Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate;

[0794] Error type indicator: Number of times the position update failed;

[0795] On PCRF:

[0796] Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity;

[0797] Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate;

[0798] Error type metric: Application session call drop rate;

[0799] One or a combination of the following metrics on 5G core network elements:

[0800] On AMF:

[0801] Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses;

[0802] Success rate metrics include: initial registration success rate, registration update success rate, switchover success rate, UECM registration failure success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate.

[0803] Latency-related metrics: Average initial registration time;

[0804] Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections;

[0805] On SMF:

[0806] Traffic volume type metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows;

[0807] Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate;

[0808] Latency type metric: Average duration of PDU session establishment process;

[0809] Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures;

[0810] On UPF:

[0811] Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface;

[0812] Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate;

[0813] Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface;

[0814] On UDM:

[0815] Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate;

[0816] On PCF:

[0817] Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy;

[0818] Success rate type indicators: AM strategy association establishment success rate, AM strategy association update success rate, AM strategy association deletion success rate, SM strategy association establishment success rate, SM strategy association update success rate, SM strategy association deletion success rate;

[0819] Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures;

[0820] On NRF:

[0821] Business volume type metric: Number of storage instances;

[0822] Success rate metrics: NF update initiation success rate, NF discovery success rate;

[0823] Error type indicators: Number of NF update failures and number of NF discovery failures;

[0824] On NSSF:

[0825] Success rate metrics: Network slice selection success rate;

[0826] Error type metric: Number of network slice selection failures.

[0827] During implementation, one or a combination of the following faults may be injected into the NFV network service plane:

[0828] The main NFV network service injection failure was caused by a CloudOS database IP address conflict.

[0829] The main NFV network service injection fault resulted in the complete blocking of VIM and damage to some service virtual machines of network elements.

[0830] The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services.

[0831] A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region.

[0832] The main regional service injection failure was caused by high read / write latency in the CloudOS storage component.

[0833] The service injection failure in the main region was a CloudOS storage component switching failure.

[0834] The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

[0835] During implementation, the OMU fault injected into the management plane was VNF, which prevented access.

[0836] In practice, evaluating the NFV network based on the second indicator assesses the overall network service quality's resilience to damage when components in the NFV network fail.

[0837] During implementation, NFV networks are evaluated according to the second indicator in one or a combination of the following ways:

[0838] NFV element resilience = α * (a1 * network component failure rate + a2 * compute / storage component failure rate + α) + β * OMU failure rate; where NFV element resilience is an evaluation index for assessing the ability of an NFV element to withstand service quality disruptions when components fail, α represents the user plane weight, β represents the management plane weight, and a1 * network component failure rate + β * OMU failure rate. i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0839] NFV network resilience = α * NFV network element failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability compute and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability compute and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability compute and storage component failure rate) * 100%; where NFV network resilience is an evaluation index for assessing the overall network service quality resilience against damage when NFV network elements, EMS, VNFM, and NFVO functional components fail; α represents the user plane weight, β represents the management plane weight, and b... i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0840] 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0841] 5G wireless network element resilience = (α * (a1 * ∑5G base station high availability network component failure rate + a2 * ∑5G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0842] 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0843] 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0844] 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0845] 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0846] 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0847] 5G core network resilience = (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0848] In practice, evaluating an NFV network according to the second metric involves evaluating one or a combination of the following performance characteristics of the NFV network:

[0849] Wireless network element resilience, core network element resilience,

[0850] The wireless network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the base station, and the core network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the core network element.

[0851] In implementation, the aforementioned wireless network element resilience is used to verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of the network element; and / or,

[0852] The core network element resilience is used to verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements.

[0853] During implementation, the resilience of the NFV network's radio elements, core elements, or a combination thereof will be assessed according to the second indicator in the following manner:

[0854] Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; α and β are set to 0.5, and a1 and a2 are recommended to be set to 1;

[0855] Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0856] Among them, Figure 6 In this context, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 600) and memory (memory 620). The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface. Transceiver 610 may be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. Processor 600 is responsible for managing the bus architecture and general processing, and memory 620 may store data used by processor 600 during operation.

[0857] This invention also provides a network evaluation apparatus, comprising:

[0858] The determination module is used to determine the metrics that need to be evaluated in the NFV network;

[0859] The monitoring module is used to monitor the indicators and collect the first indicators of the indicators before fault injection.

[0860] The fault module is used to inject faults.

[0861] The monitoring module is also used to monitor the indicator and collect a second indicator of the indicator after fault injection;

[0862] The fault module is also used to restore the NFV network to the first metric.

[0863] The evaluation module is used to evaluate NFV networks based on a second metric.

[0864] During implementation, it further includes:

[0865] The control module is used to control the fluctuation of predetermined indicators of the NFV network within a preset range when a fault is injected.

[0866] In implementation, the fault module is further used to inject faults layer by layer from the infrastructure layer to the VNF layer, according to the NFV network functional architecture; and / or,

[0867] Based on the NFV network functional architecture, faults are injected layer by layer from the VNF layer to the infrastructure layer.

[0868] In implementation, the fault module is further used to inject faults in one or a combination of the following ways:

[0869] The main NFV network service plane, backup NFV network service plane, management plane, network element component level, and infrastructure.

[0870] In implementation, the fault module is further used to inject one or a combination of the following faults into the NFV network service plane:

[0871] The main NFV network service injection failure was caused by a CloudOS database IP address conflict.

[0872] The main NFV network service injection fault resulted in the complete blocking of VIM and damage to some service virtual machines of network elements.

[0873] The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services.

[0874] A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region.

[0875] The main regional service injection failure was caused by high read / write latency in the CloudOS storage component.

[0876] The service injection failure in the main region was a CloudOS storage component switching failure.

[0877] The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

[0878] In practice, the fault module is further used to inject an OMU fault that is inaccessible in the VNF into the management plane.

[0879] In implementation, the determination module is further used to determine that the indicator is one or a combination of the following NFV network performance indicators:

[0880] One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0881] One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics;

[0882] One or a combination of the following indicators on 4G core network elements:

[0883] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the MME;

[0884] Traffic and capacity type metrics, success rate type metrics on ServingGW;

[0885] Traffic and capacity type metrics, latency type metrics, and error type metrics on the PGW;

[0886] Business volume and capacity type metrics, success rate type metrics, and error type metrics on HSS;

[0887] Traffic type metrics, success rate type metrics, and error type metrics on PCRF;

[0888] One or a combination of the following metrics on 5G core network elements:

[0889] The business volume type indicators, success rate type indicators, latency type indicators, and error type indicators on AMF;

[0890] Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on SMF;

[0891] Business volume type metrics, success rate type metrics, and error type metrics on UPF;

[0892] Success rate metrics on UDM;

[0893] Business volume type metrics, success rate type metrics, and error type metrics on PCF;

[0894] Traffic type metrics, success rate type metrics, and error type metrics on NRF;

[0895] Success rate type metrics and error type metrics on NSSF.

[0896] In practice, the predetermined indicator is one or a combination of the following indicators:

[0897] One or a combination of the following metrics on 4G base stations:

[0898] Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of uplink and / or downlink PDCP SDU bytes on the cell user plane and / or control plane;

[0899] Success rate metrics include: E-RAB establishment success rate (related to services), wireless connection success rate (related to services), RRC connection establishment success rate, and handover success rate.

[0900] Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time;

[0901] Error type metrics: wireless call drop rate, call drop rate, TCP-based wireless call drop rate, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops;

[0902] One or a combination of the following metrics on 5G base stations:

[0903] Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP PDU bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transmission TB, downlink PDSCH PRB available number, paging record received number, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of ERABs for NSA SCG Split Bear type;

[0904] Success rate metrics include: RRC connection establishment success rate, Flow establishment success rate, PDUSESSION establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate.

[0905] Latency type indicators: RRC connection average and / or maximum setup time, RLC downlink data packet average processing latency, gNB inter-NG handover average time, gNB inter-Xn handover average time, Epsfallback service handover latency from 5G to 4G, and RLC downlink data packet average processing latency per slice cell;

[0906] Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses;

[0907] One or a combination of the following indicators on 4G core network elements:

[0908] On MME:

[0909] Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs;

[0910] Success rate metrics include: EPS attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate.

[0911] Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time;

[0912] Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures;

[0913] On ServingGW:

[0914] Traffic volume and capacity type metrics: SGW average and / or peak capacity utilization, user plane uplink and / or downlink traffic, average GTP uplink and / or downlink traffic per attached user of ServingGW, average and / or maximum number of attached users of ServingGW, average and / or maximum number of bearers of ServingGW, SGW data throughput capacity utilization;

[0915] Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate;

[0916] On PGW:

[0917] Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic.

[0918] Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate;

[0919] Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW;

[0920] Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface;

[0921] On HSS:

[0922] Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate;

[0923] Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate;

[0924] Error type indicator: Number of times the position update failed;

[0925] On PCRF:

[0926] Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity;

[0927] Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate;

[0928] Error type metric: Application session call drop rate;

[0929] One or a combination of the following metrics on 5G core network elements:

[0930] On AMF:

[0931] Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses;

[0932] Success rate metrics include: initial registration success rate, registration update success rate, switchover success rate, UECM registration failure success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate.

[0933] Latency-related metrics: Average initial registration time;

[0934] Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections;

[0935] On SMF:

[0936] Traffic volume type metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows;

[0937] Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate;

[0938] Latency type metric: Average duration of PDU session establishment process;

[0939] Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures;

[0940] On UPF:

[0941] Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface;

[0942] Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate;

[0943] Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface;

[0944] On UDM:

[0945] Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate;

[0946] On PCF:

[0947] Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy;

[0948] Success rate type indicators: AM strategy association establishment success rate, AM strategy association update success rate, AM strategy association deletion success rate, SM strategy association establishment success rate, SM strategy association update success rate, SM strategy association deletion success rate;

[0949] Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures;

[0950] On NRF:

[0951] Business volume type metric: Number of storage instances;

[0952] Success rate metrics: NF update initiation success rate, NF discovery success rate;

[0953] Error type indicators: Number of NF update failures and number of NF discovery failures;

[0954] On NSSF:

[0955] Success rate metrics: Network slice selection success rate;

[0956] Error type metric: Number of network slice selection failures.

[0957] In practice, the evaluation module is further used to assess the NFV network's ability to withstand damage to overall network service quality when components in the NFV network fail, based on the second indicator.

[0958] In implementation, the evaluation module is further used to evaluate the NFV network based on the second metric in one or a combination of the following ways:

[0959] NFV element resilience = α * (a1 * network component failure rate + a2 * compute / storage component failure rate + α) + β * OMU failure rate; where NFV element resilience is an evaluation index for assessing the ability of an NFV element to withstand service quality disruptions when components fail, α represents the user plane weight, β represents the management plane weight, and a1 * network component failure rate + β * OMU failure rate. i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0960] NFV network resilience = α * NFV network element failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability compute and storage component failure rate + b3 * ∑VNMF high-availability network component failure rate + b4 * ∑VNFM high-availability compute and storage component failure rate + b5 * ∑NFVO high-availability network component failure rate + b6 * ∑NFVO high-availability compute and storage component failure rate) * 100%; where NFV network resilience is an evaluation index for assessing the overall network service quality resilience against damage when NFV network elements, EMS, VNFM, and NFVO functional components fail; α represents the user plane weight, β represents the management plane weight, and b... i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane;

[0961] 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0962] 5G wireless network element resilience = (α * (a1 * ∑5G base station high availability network component failure rate + a2 * ∑5G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0963] 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0964] 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0965] 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0966] 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1].

[0967] 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0968] 5G core network resilience = (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

[0969] In implementation, the evaluation module is further used to evaluate the NFV network based on the second metric, which is to evaluate one or a combination of the following performance characteristics of the NFV network:

[0970] Wireless network element resilience, core network element resilience,

[0971] The wireless network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the base station, and the core network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the core network element.

[0972] In implementation, the aforementioned wireless network element resilience is used to verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of the network element; and / or,

[0973] The core network element resilience is used to verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements.

[0974] In implementation, the assessment module is further used to assess one or a combination of the following: the resilience of the NFV network's wireless network elements, the resilience of the core network elements, and so on, based on the second metric:

[0975] Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0976] Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1].

[0977] For ease of description, the various parts of the device described above are divided into modules or units according to their functions. Of course, in implementing this invention, the functions of each module or unit can be implemented in one or more software or hardware components.

[0978] This invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program that performs the above-described network evaluation method.

[0979] For specific implementation details, please refer to the implementation of the network evaluation method described above.

[0980] In summary, the technical solution provided by the embodiments of the present invention provides an overall scheme for NFV network chaos engineering experiments, including system framework, operation steps, fault types, and service indicator selection scheme;

[0981] It also provides a definition scheme for NFV network resilience indicators that reflect the ability of NFV networks to withstand damage.

[0982] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0983] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0984] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0985] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0986] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A network evaluation method, characterized in that, include: Identify the metrics that need to be evaluated in Network Functions Virtualization (NFV) networks; Monitor the aforementioned indicators and collect the first indicator of the aforementioned indicators before fault injection; Injection failure; Monitor the aforementioned indicators and collect a second indicator of the aforementioned indicators after fault injection; Restoring the NFV network to the aforementioned metric is the first metric; NFV networks are evaluated based on the second metric.

2. The method as described in claim 1, characterized in that, Further includes: When a fault is injected, the predetermined indicators of the NFV network are controlled to fluctuate within a preset range.

3. The method as described in claim 2, characterized in that, The predetermined indicator is one of the following indicators or a combination thereof: One or a combination of the following metrics on 4G base stations: traffic type metrics, success rate type metrics, latency type metrics, and error type metrics; One or a combination of the following metrics on 5G base stations: traffic volume and capacity type metrics, success rate type metrics, latency type metrics, and error type metrics; One or a combination of the following indicators on 4G core network elements: Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the Mobility Management Entity (MME); Traffic and capacity type metrics, and success rate type metrics on the ServingGW service gateway; Traffic and capacity type metrics, latency type metrics, and error type metrics on the Packet Data Network Gateway (PGW); Traffic and capacity type metrics, success rate type metrics, and error type metrics on the Home Subscriber Server (HSS); Traffic type metrics, success rate type metrics, and error type metrics on the policy and billing rules functional entity PCRF; One or a combination of the following metrics on 5G core network elements: Traffic type metrics, success rate type metrics, latency type metrics, and error type metrics on the Access and Mobility Management Function (AMF); The SMF (Session Management Function) includes metrics for traffic type, success rate type, latency type, and error type. Business volume type metrics, success rate type metrics, and error type metrics on the user plane function UPF; Success rate metrics on the Unified Data Management Entity (UDM); The business volume type metric, success rate type metric, and error type metric on the policy control function PCF; Traffic type metrics, success rate type metrics, and error type metrics on the Network Storage Function (NRF); Success rate type metrics and error type metrics on the Network Slice Selection Function Entity NSSF.

4. The method as described in claim 3, characterized in that, The predetermined indicator is one of the following indicators or a combination thereof: One or a combination of the following metrics on 4G base stations: Service volume type indicators: number of bytes sent and / or received by the eNB Ethernet interface, number of service bytes sent and / or received by the eNB S1 interface, and number of bytes of PDCP Service Data Unit (SDU) for the user plane and / or control plane of the cell. Success rate metrics include: E-RAB establishment success rate (related to services), radio connection success rate (related to services), RRC connection establishment success rate, and handover success rate. Latency type indicators: RRC connection average and / or maximum setup time, E-RAB average and / or maximum setup time; Error type metrics: wireless call drop rate, call drop rate, wireless call drop rate based on Transmission Control Protocol TCP service, uplink and / or downlink PDCP SDU packet loss rate, number of uplink and / or downlink transport block errors, and number of paging record drops; One or a combination of the following metrics on 5G base stations: Traffic volume and capacity type indicators: gNB transmits and / or receives traffic data from the NG interface, cell user plane uplink and / or downlink PDCP packet data unit (PDU) bytes, gNB receives and / or transmits traffic data from the S1 interface, uplink and / or downlink transport blocks (TB), downlink physical shared channel (PDSCH) physical resource blocks (PRB) available, number of paging records received, average and / or maximum number of RRC connections in dual connectivity, average and / or maximum number of E-RABs in non-standalone (NSA) secondary cell group (SCG) split bearer type; Success rate indicators include: RRC connection establishment success rate, Flow establishment success rate, Packet Data Unit Session (PDUSESSION) establishment success rate, handover success rate, and NG interface UE-related logical signaling connection establishment success rate. Latency type indicators: RRC connection average and / or maximum setup time, Radio Link Control (RLC) downlink packet average processing latency, gNB-NG handover average time, gNB-Xn handover average time, Evolved Packet System Fallback (EPSfallback) service latency from 5G to 4G, and per slice cell RLC downlink packet average processing latency. Error type metrics: number of failed flow setup and / or modification attempts, number of uplink and / or downlink PDCP packet losses; One or a combination of the following indicators on 4G core network elements: On MME: Traffic volume type indicators: average and / or maximum number of MMEs, number of users in idle and / or connected state of MMEs, average and / or maximum number of attached users of MMEs; Success rate metrics include: Evolved Packet System (EPS) attachment success rate, authentication success rate, MME-initiated DNS resolution success rate, default and / or dedicated bearer activation success rate, PDN connection establishment success rate, service request success rate, paging success rate, tracking area update success rate, and handover success rate. Delay type indicators: average and / or maximum attachment time, average and / or maximum dedicated bearer establishment time; Error type indicators: number of authentication parameter errors, number of UE authentication failures, number of EPS attachment failures, and number of tracking area update failures; On ServingGW: Traffic volume and capacity type metrics: SGW bearer capacity average and / or peak utilization, user plane uplink and / or downlink traffic, ServingGW average uplink and / or downlink traffic per attached user (GTP), ServingGW average and / or maximum number of attached users, ServingGW average and / or maximum number of bearers, SGW data throughput capacity utilization; Success rate type metrics: ServingGW default and / or dedicated bearer establishment success rate; On PGW: Traffic volume and capacity type metrics: PGW data throughput capacity utilization, PGW average and / or peak capacity utilization, PGW S5 and / or S8 interface uplink and / or downlink traffic, PGW average and / or maximum number of attached users, PGW average and / or peak number of bearers, SGi interface receive and / or transmit traffic. Success rate metrics: Dedicated bearer establishment success rate, CDR transmission success rate; Latency type indicators: average and / or maximum duration of dedicated bearer establishment initiated by PGW; Error type indicators: number of GTP packets dropped due to errors on the PGW S5 and / or S8 interfaces, and number of IP packets dropped due to errors on the SGi interface; On HSS: Business volume and capacity type indicators: HSS number of newly issued and / or active users, HSS authentication capacity utilization rate, HSS static capacity utilization rate; Success rate metrics: HSS authentication information query success rate, HSS update and / or location cancellation success rate, HSS user data insertion and / or deletion success rate, HSS UE clearing success rate; Error type indicator: Number of times the position update failed; On PCRF: Traffic volume type metrics: Average and / or peak utilization of Gx session processing capacity; Success rate metrics: Policy control initiation and / or update and / or termination success rate, re-authentication success rate, application session authorization success rate; Error type metric: Application session call drop rate; One or a combination of the following metrics on 5G core network elements: On AMF: Traffic volume type metrics: Number of AMF registered and / or idle users, number of paging requests, number of first paging responses, and number of second paging responses; Success rate metrics include: initial registration success rate, registration update success rate, handover success rate, UECM deregistration success rate, N11 interface session context establishment success rate, N11 interface session context update success rate, N11 interface session context release success rate, N11 interface session context query success rate, and business request success rate. Latency-related metrics: Average initial registration time; Error type metrics: Number of authentication parameter errors, number of authentication rejections, number of initial registration failures, number of registration update failures, and number of business request rejections; On SMF: Traffic volume metrics: average and / or maximum number of PDU sessions, average and / or maximum number of QoS flows; Success rate type indicators: PDU session establishment success rate, SMF-initiated PDU session modification success rate, N7 interface session management SM policy creation success rate, N10 interface UE context registration success rate, N7 interface SM policy update success rate, N10 interface UE context deregistration success rate, N7 interface SM policy deletion success rate; Latency type metric: Average duration of PDU session establishment process; Error type indicators: number of PDU session establishment failures, number of SMF-initiated PDU session modification failures; On UPF: Traffic volume type metrics: average and / or maximum QoS flow count, number of GTP packets received and / or sent on N3 interface, number of GTP packets received and / or sent on N9a interface, number of bytes received and / or sent on N6 interface; Success rate metrics: PFCP session establishment success rate, PFCP session modification success rate; Error type indicators: number of PFCP session establishment and / or modification failures, number of erroneous GTP packets received on N3 interface, number of IP packets dropped due to errors on N6 interface, and number of erroneous GTP packets received on N9c interface; On UDM: Success rate type metrics: UECM registration success rate initiated by AMF, UECM registration success rate initiated by SMF, registration parameter update success rate, user data acquisition success rate, user data subscription success rate, and user data unsubscription success rate; On PCF: Business volume type indicators: average and / or maximum total number of associations under AM strategy, average and / or maximum total number of associations under SM strategy; Success rate metrics: Access and Mobility AM policy association establishment success rate, AM policy association update success rate, AM policy association deletion success rate, SM policy association establishment success rate, SM policy association update success rate, SM policy association deletion success rate; Error type metrics: Number of SM policy association establishment failures, number of SM policy association update failures; On NRF: Business volume type metric: Number of storage instances; Success rate metrics: Network Functional Entity (NF) update success rate, NF discovery success rate; Error type indicators: Number of NF update failures and number of NF discovery failures; On NSSF: Success rate metrics: Network slice selection success rate; Error type metric: Number of network slice selection failures.

5. The method as described in claim 1, characterized in that, When injecting faults, faults are injected layer by layer from the infrastructure layer to the Virtual Network Function (VNF) layer, according to the NFV network function architecture; and / or, Based on the NFV network function architecture, faults are injected layer by layer from the Virtual Network Function (VNF) layer to the infrastructure layer.

6. The method as described in claim 1, characterized in that, When a fault is injected, it is in one or a combination of the following aspects: The main NFV network service plane, backup NFV network service plane, management plane, network element component level, and infrastructure.

7. The method as described in claim 6, characterized in that, Inject one or a combination of the following faults into the primary NFV network service plane: The fault injected into the main NFV network service was due to an IP address conflict with the CloudOS database. The main NFV network service injection fault resulted in the complete blockage of Virtual Infrastructure Management (VIM) and damage to some virtual machines serving network elements. The main NFV network service injection failure was caused by the data center gateway deleting the S1 interface route, resulting in the complete blockage of main NFV network services. A service injection failure in the main region resulted in the disconnection of a batch of virtual machine storage in the main region. The main regional service injection failure was caused by high read / write latency in the CloudOS storage component. The service injection failure in the main region was a CloudOS storage component switching failure. The user plane service injection failure was caused by a CloudOS component failure, resulting in the interruption of the connection between the user plane network element and the control plane network element.

8. The method as described in claim 6, characterized in that, The management plane is injected with a fault indicating that the Operation and Maintenance Unit (OMU) of the VNF is inaccessible.

9. The method according to any one of claims 1 to 8, characterized in that, The second metric assesses the ability of an NFV network to withstand disruptions to overall network service quality when components fail.

10. The method as described in claim 9, characterized in that, NFV networks are evaluated according to the second metric using one or a combination of the following methods: NFV network element resilience = α *(a1*Network component failure rate + a2*Compute / storage component failure rate +)+ β *OMU failure rate; where NFV network element resilience is an evaluation indicator for assessing the ability of a network element to withstand service quality disruptions when components in an NFV network element fail. α Indicates user face weight, β Indicates the management aspect weight, a i {i∈[1,3]}∈[0,1] represents the weight coefficients of each functional layer in the user plane; NFV network resilience = α *NFV network element failure rate+ β *(b1*∑EMS high-availability network component failure rate + b2*∑EMS high-availability compute and storage component failure rate + b3*∑VNMF high-availability network component failure rate + b4*∑VNFM high-availability compute and storage component failure rate + b5*∑NFVO high-availability network component failure rate + b6*∑NFVO high-availability compute and storage component failure rate)*100%; where NFV network resilience is an evaluation indicator for assessing the overall network service quality's ability to withstand damage when NFV network elements, the Element Management System (EMS), the Virtual Network Function Manager (VNFM), and the Network Function Virtualization Orchestrator (NFVO) functional components fail. α Indicates user face weight, β Indicates the management aspect weight, b i {i∈[1,6]}∈[0,1] represents the weight coefficients of each functional layer in the user plane; 4G wireless network element resilience = (α * (a1 * ∑4G base station high availability network component failure rate + a2 * ∑4G base station high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where the high availability network components include one or a combination of the following components: top rack TOR switch, bottom row EOR switch, network interface card; the high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1]. 5G wireless network element resilience = (α*(a1*∑5G base station high availability network component failure rate + a2*∑5G base station high availability computing and storage component failure rate) + β*OMU failure rate)*100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1]. 4G core network element resilience = (α * (a1 * ∑4G core network element high availability network component failure rate + a2 * ∑4G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1]. 5G core network element resilience = (α * (a1 * ∑5G core network element high availability network component failure rate + a2 * ∑5G core network element high availability computing and storage component failure rate) + β * OMU failure rate) * 100%, where high availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high availability component, and the values ​​of α, β, a1, and a2 are in the range of [0, 1]. 4G wireless network resilience = (α * 4G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1]. 5G wireless network resilience = (α * 5G base station failure rate + β * (b1 * ∑EMS high-availability network component failure rate + b2 * ∑EMS high-availability computing and storage component failure rate)) * 100%, where high-availability network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; high-availability computing and storage components include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents the summation of the failure rates of each high-availability component, and the values ​​of α, β, b1, and b2 range from [0, 1]. 4G core network resilience = (α*∑4G core network element (MME, ServingGW, PGW, HSS, PCRF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1]. 5G core network resilience = (α*∑5G core network element (AMF, SMF, UPF, UDM, PCF, NRF, NSSF) failure rate + β*(b1*∑EMS high availability network component failure rate + b2*∑EMS high availability computing and storage component failure rate + b3*∑VNMF high availability network component failure rate + b4*∑VNFM high availability computing and storage component failure rate + b5*∑NFVO high availability network component failure rate + b6*∑NFVO high availability computing and storage component failure rate +)*100%, where high availability components include one or a combination of the following components: TOR switch, EOR switch, network interface card, and computing and storage include one or a combination of the following components: physical machine, virtual machine, database; ∑ represents summation, and the values ​​of α, β, b1, b2, b3, b4, b5, and b6 are in the range [0, 1].

11. The method according to any one of claims 1 to 8, characterized in that, Evaluating an NFV network according to the second metric involves evaluating one or a combination of the following performance characteristics of the NFV network: Wireless network element resilience, core network element resilience, The wireless network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the base station, and the core network element resilience is used to evaluate the weighted sum of the proportion of failures in each functional layer component of the core network element.

12. The method as described in claim 11, characterized in that, The aforementioned wireless network element resilience is used to verify the cloud-based scalability and fault tolerance of wireless network element services, as well as the impact of maximum uncertainty on the steady state of the network element; and / or, The core network element resilience is used to verify the cloud-based scalability and fault tolerance of core network element services, as well as the impact of maximum uncertainty on the steady state of network elements.

13. The method as described in claim 11, characterized in that, The resilience of wireless network elements, core network elements, or a combination thereof shall be assessed according to the second indicator in the following manner: Wireless network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1]. Core network element resilience = (α * (a1 * network component failure rate + a2 * compute / storage component failure rate) + β * OMU failure rate) * 100%, where network components include one or a combination of the following components: TOR switch, EOR switch, network interface card; compute and storage components include one or a combination of the following components: physical machine, virtual machine, database; the values ​​of α, β, a1, and a2 are in the range of [0, 1]. α Indicates user face weight, β a1 and a2 represent the weights of the management plane and the weight coefficients of the user plane functional layer.

14. A network evaluation device, characterized in that, include: The processor is used to read programs from memory and execute the following procedures: Identify the metrics that need to be evaluated in the NFV network; Monitor the aforementioned indicators and collect the first indicator of the aforementioned indicators before fault injection; Injection failure; Monitor the aforementioned indicators and collect a second indicator of the aforementioned indicators after fault injection; Restoring the NFV network to the aforementioned metric is the first metric; Evaluate NFV networks based on the second metric; A transceiver is used to receive and send data under the control of a processor.

15. A network evaluation device, characterized in that, include: The determination module is used to determine the metrics that need to be evaluated in the NFV network; The monitoring module is used to monitor the indicators and collect the first indicators of the indicators before fault injection. The fault module is used to inject faults. The monitoring module is also used to monitor the indicator and collect a second indicator of the indicator after fault injection; The fault module is also used to restore the NFV network to the first metric. The evaluation module is used to evaluate NFV networks based on a second metric.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that performs the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Reliability evaluation system design method based on software fault injection

    CN103473162A

  • Runtime model based configuration method for fault-tolerant mechanism of cloud computing

    CN105005509A