Network node management methods, devices, and equipment based on network resilience models

CN122293507APending Publication Date: 2026-06-26PURPLE MOUNTAIN LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-06-26

Smart Images

  • Figure CN122293507A_ABST
    Figure CN122293507A_ABST
Patent Text Reader

Abstract

This application relates to the field of communication network technology and discloses a network node management method, apparatus, and device based on a network resilience model. The method involves obtaining resilience identification codes reported by each network node in a faulty network. These codes are generated by the network nodes based on a preset three-layer architecture model. Each resilience identification code includes a network domain code set in the network domain of the network node in the first layer architecture and full-dimensional real-time indicators set in the third layer architecture. Based on the network domain code of each network node, the network domain to which each network node belongs is determined, and a preset first demand threshold is obtained within that network domain. The full-dimensional real-time indicators of each network node are matched with the first demand threshold to calculate a first comprehensive matching degree for each network node. The network node with the lowest first comprehensive matching degree is identified as the faulty node. The beneficial effect is improved network node monitoring efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication network technology, and in particular to a network node management method, apparatus and device based on a network resilience model. Background Technology

[0002] With the rapid development of information and communication technologies, network architecture is evolving towards a cloud-edge-device integrated model, which continuously increases the requirements for network stability and continuity.

[0003] Currently, in a cloud-edge-device network, multiple network nodes are typically divided into multiple network domains such as backbone network and edge network. Furthermore, different rules are usually set for managing different network domains.

[0004] Therefore, the current network suffers from network domain fragmentation, which makes it impossible to perform unified global monitoring and collaborative management, resulting in low efficiency in network node monitoring. Summary of the Invention

[0005] This application provides a network node management method, apparatus, and device based on a network resilience model, which solves the technical problem of low network node monitoring efficiency and achieves the technical effect of improving network node monitoring efficiency.

[0006] To achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, embodiments of this application provide a network node management method based on a network resilience model, the method comprising: Obtain the resilience identification code reported by each network node in the faulty network; the resilience identification code is generated by the network node based on a preset three-layer architecture model, and the resilience identification code includes the network domain code of the network domain where the network node is located in the first layer architecture, and the full-dimensional real-time indicators of the network node in the third layer architecture. Based on the network domain code of each network node, determine the network domain to which each network node belongs, and obtain the preset first demand threshold in the network domain; The first comprehensive matching degree of each network node is calculated by matching the full-dimensional real-time indicators of each network node with the first demand threshold. The network node with the lowest overall matching degree is identified as the faulty node.

[0007] The network node management method based on the network resilience model proposed in this application obtains the resilience identification code of each network node in the faulty network, determines the network domain to which the corresponding network node belongs based on the resilience identification code, obtains the first demand threshold of the network domain, and obtains the first comprehensive matching degree based on the first demand threshold and the real-time indicators of the network node across all dimensions. Finally, the faulty node is determined based on the first comprehensive matching degree. This method achieves accurate monitoring of the faulty network and precise location of the faulty nodes, thereby improving the monitoring efficiency and effectiveness of the faulty network.

[0008] Secondly, embodiments of this application provide a network node management device based on a network resilience model, the device comprising: The acquisition module is used to acquire the elastic identification code reported by each network node in the faulty network. The elastic identification code is generated by the network node based on a preset three-layer architecture model. The elastic identification code includes the network domain code of the network domain where the network node is located in the first layer architecture, and the full-dimensional real-time indicators of the network node in the third layer architecture. The monitoring module is used to determine the network domain to which each network node belongs based on the network domain code of each network node, and obtain a preset first demand threshold in the network domain; match the full-dimensional real-time indicators of each network node with the first demand threshold to calculate the first comprehensive matching degree of each network node; and determine the network node with the lowest first comprehensive matching degree as the fault node.

[0009] Thirdly, embodiments of this application provide a computer device, including: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method described in any of the above embodiments.

[0010] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions, which are used to cause a computer to perform the method described in any one of the above embodiments.

[0011] Fifthly, embodiments of this application provide a computer program product, including computer instructions, which are used to cause a computer to perform the method described in any of the above embodiments. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating a network node management method based on a network resilience model, provided in this application embodiment; Figure 2 This is a schematic diagram of a tree structure for a flexible identification code provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a network node management system based on a network resilience model provided in an embodiment of this application; Figure 4 A structural diagram of a network node management device based on a network resilience model provided in this application embodiment; Figure 5 This is a structural diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] With the rapid development of information and communication technologies, network architecture is evolving towards an integrated "cloud-edge-device" architecture. The services it supports have expanded from traditional data transmission to multiple fields such as industrial control, telemedicine, autonomous driving, and financial transactions. These services place extremely high demands on network stability and continuity. Network failures (such as link interruptions, malicious attacks, or sudden traffic surges) can lead to significant economic losses and even security risks. Therefore, network resilience (i.e., the network's ability to withstand failures, maintain services, recover quickly, and adaptively optimize) has become a core indicator for ensuring business continuity.

[0016] However, current network resilience management faces three core problems: First, the metrics are fragmented. Different network domains (such as backbone networks and edge networks) and different equipment vendors have inconsistent definitions and measurement metrics for network resilience. For example, some network domains only focus on "availability," while others additionally include "attack resistance capability." This situation makes it difficult to compare and coordinate network resilience management across different domains.

[0017] Secondly, the modeling methods are heterogeneous. Existing network resilience modeling is mostly designed for specific scenarios such as data centers and industrial control networks. However, different scenarios lack a unified model framework. Therefore, existing networks cannot achieve a unified mapping of network-wide resilience capabilities, increasing the complexity of cross-domain collaboration.

[0018] Third, the identification system is lacking. Existing networks lack a standardized mechanism for identifying resilience status, making it impossible to accurately pinpoint the weak points in each node's resilience. Consequently, it is difficult to quickly trace the root cause when a fault occurs, leading to low fault handling efficiency, indiscriminate resource scheduling, and impacting network recovery speed.

[0019] Furthermore, as cybersecurity shifts from "risk prevention and control" to "enhancing network resilience," existing technological solutions are no longer sufficient to meet the intelligent and standardized requirements of next-generation networks for resilience management. Therefore, there is an urgent need to build a complete technical system encompassing "unified modeling, accurate identification, and efficient operation" to address the aforementioned management inefficiencies.

[0020] To address the management inefficiencies caused by fragmented network resilience metrics, heterogeneous modeling methods, and lack of an identification system in existing networks, this application provides a full-cycle intelligent management method applicable to integrated "cloud-edge-device" network scenarios such as telecommunications backbone networks, edge computing networks, and financial data center networks. This method can standardize network resilience metrics, unify models, and make states identifiable, thereby improving the efficiency of resilience management and the accuracy of fault handling in heterogeneous network environments.

[0021] First, this application constructs a unified description framework for network resilience covering the entire process of "prevention-endurance-recovery-adaptation," and designs 16 quantitative measurement indicators to form a comprehensive measurement model. These 16 quantitative indicators cover the entire lifecycle of resilience capabilities, solve the problem of indicator fragmentation, and enable cross-domain resilience state comparison.

[0022] Secondly, based on the three-layer structure of "domain-node-capability," this application constructs a unified modeling model and designs an elastic identification code using a 64-bit binary tree structure to achieve precise binding between indicators and nodes. The three-layer "domain-node-capability" model can better adapt to the integrated "cloud-edge-device" network, achieving model unification and providing a unified mapping for elastic capabilities across the entire network, reducing the complexity of cross-domain collaboration. Furthermore, the 64-bit tree identification code enables end-to-end positioning of "domain-node-capability-indicator," achieving state identifiability and supporting one-click querying, precise positioning, and traceability of elastic states.

[0023] Furthermore, this application establishes a full-cycle status monitoring, fault location, and dynamic priority recovery scheduling mechanism based on identification codes. This code-based operating mechanism reduces fault location time by more than 50% and improves recovery scheduling accuracy by more than 40%, significantly enhancing network resilience management efficiency and achieving highly efficient management.

[0024] Furthermore, this application achieves identifiable, queryable, and traceable network resilience status, improving fault handling efficiency and resource scheduling accuracy in heterogeneous network environments. It can be widely applied in scenarios such as telecommunications backbone networks, edge networks, and financial data center networks, providing technical support for the construction of an intelligent, highly resilient network ecosystem. Additionally, this application establishes a dynamic priority recovery scheduling mechanism based on identification codes, ranking backup nodes according to resilience index matching degree, thereby improving recovery accuracy.

[0025] Finally, this application has a wide range of application scenarios. Specifically, it can be applied to integrated "cloud-edge-device" networks such as telecommunications backbone networks, edge computing networks, financial data center networks, and industrial control networks. Furthermore, this application is based on existing network equipment (routers, firewalls), monitoring tools (Zabbix, Prometheus), and blockchain technology, eliminating the need for developing entirely new hardware and keeping costs under control.

[0026] According to an embodiment of this application, a network node management method based on a network resilience model is provided. This network node management method based on a network resilience model is applied to a management node. A network management platform may run on this management node. Specifically, the management node may be a computer device. In the following embodiments, the computer device is described as the executing entity. Optionally, the computer device may be a mobile terminal, a personal computer, a server, etc.

[0027] It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed on a computer device via a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0028] Figure 1A flowchart of a network node management method based on a network resilience model provided in this application embodiment is shown below. Figure 1 As shown, with a computer device as the execution subject, the process includes the following steps: S101. Obtain the resilience identification code reported by each network node in the faulty network. The resilience identification code includes the network domain code of the network domain where the network node is located, as well as the real-time indicators of the network node across all dimensions.

[0029] For example, a computer device can obtain the resilient identification code uploaded by each network node from the network nodes of a faulty network.

[0030] In one implementation, the computer device can send a Resilience Code retrieval command to each network node. Upon receiving the command, each network node can extract the Resilience Code from its own data. Each network node can then transmit the Resilience Code back to the computer device via the network. The computer device receives and collects these Resilience Codes.

[0031] In another implementation, the computer device can obtain the elastic identification code periodically uploaded by each network node. Each network node can extract the elastic identification code from its own data according to a preset reporting period and then send it to the computer device.

[0032] In one implementation, the resilience identifier is a string of codes used to identify specific information about a network node. The resilience identifier may include the network domain code of the network domain in which the network node resides. Furthermore, the resilience identifier may include real-time metrics of various dimensions extracted from the network node's data.

[0033] Optionally, the flexible identification code can be a fixed-length code. For example, the fixed length can be 32 bits, 48 ​​bits, 64 bits, 80 bits, 128 bits, etc.

[0034] Optionally, each bit in this flexible identification code can have a pre-defined meaning. For example, bits 0-16 can correspond to network domain coding. Similarly, bits 33-64 can correspond to full-dimensional real-time metrics.

[0035] In one implementation, a network domain code is a code used to uniquely identify a network domain. A network may contain one or more network domains. A network domain code can be assigned to each network domain. For example, a network may include multiple network domains such as a backbone network, edge network, and data center network.

[0036] In one implementation, the multi-dimensional real-time metrics are data reflecting the network resilience of a network node across multiple dimensions. For example, these could include four core dimensions: prevention, resilience, recovery, and adaptation. Each dimension can correspond to one or more specific data metrics. For instance, the prevention dimension could include two metrics: mean time between failures (MTTR) and attack prevention rate. Similarly, the resilience dimension could include eight metrics: fault detection latency, fault detection accuracy, mean diagnostic time, core service availability, functional degradation level, network speed decrease rate during events, throughput decrease rate, and latency increase level. The recovery dimension could include three metrics: mean time to repair (MTTR), recovery point objective (RPO) achievement rate, and recovery time objective (RTO) achievement rate. And the adaptation dimension could include three metrics: resource reallocation latency, resilience policy update frequency, and service continuity assurance rate.

[0037] In one implementation, the resilient identifier may further include a node code for the network node. Optionally, the node code may be a code with a preset length. Optionally, each network node may have a unique node code. For example, the data length of the node code may be 8 bits.

[0038] In one implementation, network nodes can report to computer devices via an IPsec encrypted channel, thereby enabling the network management platform running on the computer device to obtain the elastic identification code of each network node.

[0039] S102. Based on the network domain code of each network node, determine the network domain to which each network node belongs, and obtain the preset first demand threshold in the network domain.

[0040] For example, the computer device parses the collected elastic identification codes of each network node and extracts the network domain code. The computer device can determine the network domain to which the network node belongs based on the network domain code. Furthermore, the computer device can obtain a preset first demand threshold for the network domain based on the network domain code.

[0041] In one implementation, the first demand threshold preset for the network domain can be pre-stored locally. The computer device can then retrieve the threshold based on the network domain code after determining the network domain code.

[0042] In another implementation, the first demand threshold can be stored on a specified network node within the network domain. The computer device can then request and retrieve the first demand threshold from the specified network node within that network domain after determining the network domain to which the network node belongs.

[0043] In one implementation, the first requirement threshold is a threshold standard that is pre-set in the network domain and used to match and compare with the real-time indicators of the network nodes across all dimensions.

[0044] In one implementation, since the network field code has a fixed position in the flexible identification code, the computer device can directly extract a portion of data from the fixed position of the flexible identification code according to its fixed format and parse it to obtain the network field code.

[0045] In one implementation, the computer device maintains a local mapping table that maps network domain codes to network domains. The computer device can determine the network domain corresponding to the parsed network domain code by looking up this mapping table.

[0046] For example, when the elastic identification code includes four dimensions of real-time indicators, the first demand threshold acquired by the computer device can also include thresholds in four dimensions. This first demand threshold can be denoted as D=(D1,D2,D3,D4). Here, D1,D2,D3,D4 can be vectors or matrices. D1,D2,D3,D4 can include thresholds for multiple specific indicators.

[0047] In one implementation, different threshold values ​​can be set for the same metric for different network domains. For example, the core domain MTTR ≤ 30s and the edge domain MTTR ≤ 60s.

[0048] S103. Use the full-dimensional real-time indicators of each network node to match with the first demand threshold, and calculate the first comprehensive matching degree of each network node.

[0049] For example, after obtaining the full-dimensional real-time metrics of each network node and the first demand threshold of the network domain to which each network node belongs, the computer device can calculate the matching between the full-dimensional real-time metrics and the first demand threshold by comparing the full-dimensional real-time metrics with the first demand threshold. Finally, the computer device can calculate the first comprehensive matching degree of each network node.

[0050] In one implementation, the real-time metrics of a network node across all dimensions can be denoted as R=(R1,R2,R3,R4).

[0051] In one implementation, the computer device can compare the real-time indicators R=(R1,R2,R3,R4) of a network node with a first demand threshold D=(D1,D2,D3,D4) one by one to determine the matching degree between the real-time indicators of the network node and the first demand threshold. That is, for each dimension, R1 is compared with D1, R2 with D2, R3 with D3, and R4 with D4. Furthermore, the computer device can comprehensively calculate the final first comprehensive matching degree based on the matching degree of each dimension. In one implementation, the first comprehensive matching degree is a comprehensive evaluation value obtained by matching the real-time indicators of the network node across all dimensions with the first demand threshold. This value is used to measure the degree of matching between the network node and the threshold standard.

[0052] S104. Determine the network node with the lowest overall matching degree as the faulty node.

[0053] For example, the computer device compares the calculated first comprehensive matching scores of each network node and finds the one with the smallest value. Then, the computer device can obtain the network node corresponding to the smallest value. The computer device can then identify the network node corresponding to the smallest value as the faulty node.

[0054] In one implementation, within a faulty network, the network node with the lowest overall matching degree typically indicates that the gap between its real-time metrics and the required threshold is the largest. In this case, it can be determined that the probability of this network node being abnormal is the highest.

[0055] In one implementation, once a faulty node is identified in the faulty network, it can be locked as the root node. Optionally, in subsequent processing, the computer device can infer the network domain where the root node is located based on the root node and perform replacement based on the root node.

[0056] In this embodiment, by obtaining the elastic identification code of each network node in the faulty network, determining the network domain to which the corresponding network node belongs based on the elastic identification code, and then obtaining the first demand threshold of the network domain, and obtaining the first comprehensive matching degree based on the first demand threshold and the real-time indicators of the network node across all dimensions, the faulty node is finally determined based on the first comprehensive matching degree. This method enables accurate monitoring of the faulty network and precise location of the faulty nodes, thereby improving the monitoring efficiency and effectiveness of the faulty network.

[0057] In one example, in step S103 above, taking a network node as an example, the computer device can parse and obtain the full-dimensional real-time indicators of the network node after obtaining its elastic identification code. The full-dimensional real-time indicators include at least one real-time indicator for at least one dimension. For example, the full-dimensional real-time indicators may include four dimensions. Each dimension may correspond to multiple real-time indicators.

[0058] Computer equipment can calculate the first comprehensive matching degree of each network node based on the network node's full-dimensional real-time indicators. This process may include: S1031. If the real-time indicator is within the range indicated by the first demand threshold, then the indicator matching degree of the real-time indicator of the determined dimension is the highest preset matching value.

[0059] For example, the computer device can process each real-time indicator individually. After acquiring a real-time indicator, the computer device can obtain the corresponding demand threshold from a first demand threshold. The computer device can compare the real-time indicator with the range indicated by the demand threshold. If the real-time indicator value falls within the range indicated by the demand threshold, the computer device sets the indicator matching degree of that real-time indicator to the highest preset matching value. This highest preset matching value indicates that the real-time indicator of that dimension fully meets the requirements of the first demand threshold. The highest preset matching value can be 1.

[0060] In one implementation, the real-time metric may include both positive and negative metrics.

[0061] Optionally, a positive indicator is typically one where a higher real-time value better meets the demand. The demand threshold typically indicates a range greater than or equal to the demand threshold and less than or equal to the maximum value of the real-time indicator.

[0062] Optionally, negative indicators are typically those where a smaller real-time value better reflects the demand. The range indicated by this demand threshold is usually less than or equal to the demand threshold and greater than or equal to the minimum value of this real-time indicator.

[0063] In one implementation, the metric matching degree is used to quantify the degree of conformity between real-time metrics and required thresholds. The value of the metric matching degree typically ranges from 0 to 1. 1 indicates that the real-time metric in that dimension fully meets expectations. 0 indicates that the real-time metric does not exist, or that the real-time metric did not take a value in this instance. Within the range of 0 to 1, the closer the metric matching degree is to 1, the more the real-time metric conforms to the requirements.

[0064] S1032. If the real-time indicator is not within the range indicated by the first demand threshold, the indicator matching degree of the real-time indicator is calculated based on the difference between the real-time indicator and the first demand threshold.

[0065] For example, when a computer device detects that a real-time metric value is outside the range indicated by the corresponding demand threshold in a first demand threshold, the computer device can calculate the difference between the real-time metric and the corresponding demand threshold. Then, the computer device can perform a calculation based on this difference to obtain a value between 0 and 1, and use this value as the metric matching degree of the real-time metric. This value reflects the degree of deviation between the real-time metric and the demand threshold. For example, this calculation method can be exponentialization, normalization, etc.

[0066] In one implementation, the computer device can exponentialize the difference using an exponential function. For example, the exponential function could be e raised to the power of x, 2 raised to the power of x, etc. The difference is represented by x in the exponential function.

[0067] In one implementation, the computer device may also replace the indexation process with other normalization methods to map the difference to a value between 0 and 1.

[0068] S1033. Calculate the weighted sum of the indicator matching degrees of multiple real-time indicators for each dimension to obtain the dimension matching degree of each dimension. For example, after calculating the indicator matching degree of multiple real-time indicators for each dimension according to the above steps S1031 and S1032, the computer device can perform a weighted sum of the indicator matching degrees of multiple real-time indicators for each dimension to obtain the dimension matching degree of each dimension.

[0069] In one implementation, a weight can be assigned to each real-time metric within each dimension. The sum of the weights of multiple real-time metrics within a dimension can be 1.

[0070] One implementation method involves pre-setting weights based on the importance of each real-time indicator within that dimension. A larger weight value indicates greater importance of that real-time indicator in the evaluation of that dimension.

[0071] S1034. Calculate the weighted sum of the dimensional matching degrees of each network node to obtain the first comprehensive matching degree of each network node.

[0072] For example, after calculating the dimensional matching degree of a network node according to step S1033 above, the computer device can calculate the dimensional matching degree of the network node and the weighted sum of the dimensional matching degree of the network node to obtain the first comprehensive matching degree of the network node.

[0073] In one implementation, a weight can be preset for each dimension. The sum of the weights of all dimensions of a network node is 1. For example, when there are 4 dimensions, the sum of the weights of the 4 dimensions is 1.

[0074] One implementation involves pre-setting weights as numerical values ​​based on the importance of each dimension in the overall evaluation. A larger weight value indicates a greater influence of that dimension on the overall evaluation.

[0075] In one example, when the full-dimensional real-time metrics in a network node's resilience identifier are denoted as R=(R1,R2,R3,R4), and the first demand threshold is denoted as D=(D1,D2,D3,D4), the formula for calculating the first comprehensive metric can be: Among them, M(R,D) is the first comprehensive index. The weights are for four dimensions. And, . Let be the dimension matching degree of the i-th dimension. The value of i can be from 1 to 4.

[0076] In this example, by matching the real-time metrics of network nodes across all dimensions with the first demand threshold, a comprehensive quantitative evaluation of network node metrics is achieved, providing a foundation for unified monitoring and management of network nodes and improving network monitoring efficiency.

[0077] In one example, in step S1032 above, when the real-time indicator is not within the range indicated by the first demand threshold, the computer device can calculate the indicator matching degree of the dimension based on the difference between the real-time indicator and the first demand threshold. This process may include: Obtain the difference between the real-time metric and the first requirement threshold. Calculate the percentage of this difference within the range of the real-time metric values ​​for that dimension, obtaining the power. Use this power to exponentiate the natural constant, yielding the metric matching degree for that dimension.

[0078] For example, the process can be represented by the following formula: in, Let be the j-th real-time metric in the i-th dimension. The threshold value is the demand threshold corresponding to the j-th real-time indicator in the i-th dimension. This represents the maximum value of the j-th real-time indicator in the i-th dimension. This represents the minimum value of the j-th real-time indicator in the i-th dimension.

[0079] In this example, by using an indicator function with base e to calculate the matching degree between each real-time indicator and the first demand threshold, the accuracy of indicator matching degree is achieved, providing a basis for network monitoring and improving the monitoring efficiency of each network node in the network.

[0080] In one implementation, the calculation formula for the above positive index can be: in, Let be the j-th real-time metric in the i-th dimension. The threshold value is the demand threshold corresponding to the j-th real-time indicator in the i-th dimension. This represents the maximum value of the j-th real-time indicator in the i-th dimension. This represents the minimum value of the j-th real-time indicator in the i-th dimension. This indicates that the real-time indicator does not exist.

[0081] In one implementation, the calculation formula for the aforementioned negative index can be: in, Let be the j-th real-time metric in the i-th dimension. The threshold value is the demand threshold corresponding to the j-th real-time indicator in the i-th dimension. This represents the maximum value of the j-th real-time indicator in the i-th dimension. This represents the minimum value of the j-th real-time indicator in the i-th dimension. This indicates that the real-time metric does not exist. For example, for core business operations, , .

[0082] In one example, the indicator matching degree of the real-time indicator is calculated based on the difference between the real-time indicator and the first demand threshold, including: Obtain the difference between the real-time metric and the first requirement threshold. Calculate the percentage of this difference within the range of the real-time metric values ​​for that dimension. Based on this percentage, determine the metric matching degree for that dimension.

[0083] For example, the process can be represented by the following formula: in, Let be the j-th real-time metric in the i-th dimension. The threshold value is the demand threshold corresponding to the j-th real-time indicator in the i-th dimension. This represents the maximum value of the j-th real-time indicator in the i-th dimension. This represents the minimum value of the j-th real-time indicator in the i-th dimension.

[0084] For example, this method can linearly calculate the matching degree of indicators, which greatly improves the computational efficiency.

[0085] In one example, after the computer device has located the faulty node, it can also optimize that node. This process may include: S201. Based on the network domain code of the faulty node, determine the network domain to which the faulty network belongs, and obtain the backup node and the preset second demand threshold in the network domain to which the faulty node belongs.

[0086] For example, the computer device can read the network domain code of the faulty node. Based on the network domain code, the computer device determines the network domain in which the faulty node resides. Furthermore, based on the network domain code, the computer device can determine a backup node and a preset second demand threshold within the network domain in which the faulty node resides.

[0087] In one implementation, network domain coding is used to uniquely identify a specific network domain within the network. Examples include edge network domains and backbone network domains. Another example is a financial data center domain.

[0088] In one implementation, the backup node is a redundant network node serving as a backup in the network domain. Optionally, the backup node can be used to replace a failed node. Optionally, the backup node can be used to supplement other nodes in the network domain when they cannot meet the demand.

[0089] In one implementation, the second requirement threshold is a performance or status indicator threshold set for the standby node, used to evaluate whether the standby node meets the conditions for replacing the faulty node.

[0090] In one implementation, the computer device can determine the network domain corresponding to the network domain code by querying a mapping table stored locally.

[0091] In one implementation, the computer device can also determine the service type corresponding to the network domain code by querying locally stored data. For example, the service type could be a core transaction service.

[0092] In one implementation, the information about the backup node and the second demand threshold can be stored locally. The computer device can then obtain the backup node in the network domain by looking up the backup node information table stored locally, based on the network domain encoding. Furthermore, the computer device can obtain the second demand threshold by looking up the threshold information table stored locally, based on the network domain encoding. For example, the second demand threshold may include core service availability. The core service availability threshold can be ≥99.95%. As another example, the second demand threshold may include MTTR. The MTTR threshold can be ≤30s.

[0093] In one implementation, the information about the backup node and the second demand threshold can be stored on a network node within the network domain. Then, the computer device can locate the network node in the network domain locally and access that network node to obtain the backup node and the second demand threshold for that network domain.

[0094] S202. If the real-time metrics of the standby node are not within the range indicated by the second demand threshold, then delete the standby node.

[0095] For example, after determining the backup nodes, the computer device can also obtain the elasticity identification code of each backup node. The computer device can then parse the elasticity identification code of each backup node to obtain its full-dimensional real-time metrics. Next, the computer device can filter backup nodes by comparing the full-dimensional real-time metrics and the second demand threshold of each backup node.

[0096] In one implementation, the computer device can compare the full-dimensional real-time metrics of each backup node with the range indicated by a second demand threshold. If the full-dimensional real-time metrics of the backup node are within the range indicated by the second demand threshold, the backup node is determined as a candidate node, and further screening is performed on the backup node in subsequent steps. If the full-dimensional real-time metrics of the backup node are not within the range indicated by the second demand threshold, the backup node is deleted to reduce the computational load of subsequent steps.

[0097] In one implementation, when there are no backup nodes for the full-dimensional real-time indicators within the range indicated by the second demand threshold, the computer device can relax the filtering conditions and select backup nodes whose number of real-time indicators that do not meet the second demand threshold is less than or equal to the filtering threshold. For example, when the filtering threshold is 3, the computer device can filter backup nodes whose number of real-time indicators that do not meet the second demand threshold is less than or equal to 3.

[0098] S203. Use the real-time indicators of the backup node across all dimensions to match the second demand threshold, and calculate the second comprehensive matching degree of the backup node.

[0099] For example, the computer device can calculate the matching degree of the remaining backup nodes after screening in step S202 above to obtain a second comprehensive matching degree. The method of calculating this matching degree is similar to the specific process of calculating the first comprehensive matching degree in step S103 above.

[0100] In one implementation, the computer device can first match each real-time indicator in the full-dimensional real-time indicators with the value in the second demand threshold to calculate the indicator matching degree. Then, the computer device can calculate the dimension matching degree based on the indicator matching degrees of multiple real-time indicators for each dimension. Finally, the computer device calculates the second comprehensive matching degree based on the dimension matching degrees of each dimension of the backup node.

[0101] In one implementation, the computer device can determine the matching degree of the indicator as the highest preset matching value when the real-time indicator is within the range indicated by the second demand threshold. The highest preset matching value can be 1. Furthermore, when the real-time indicator is not within the range indicated by the second demand threshold, the computer device can perform indexation based on the difference between the real-time indicator and the second demand threshold to obtain the final indicator matching degree.

[0102] In one implementation, the second comprehensive matching degree is a quantitative indicator used to represent the overall degree of matching between the standby node and the second requirement threshold across all key performance and status indicators.

[0103] S204. Schedule the backup node with the highest overall matching degree to replace the faulty node.

[0104] For example, the computer device selects the backup node with the highest matching degree as the best candidate to replace the faulty node based on the second comprehensive matching degree of each backup node calculated in step S203. Then, the computer device performs a scheduling operation to replace the faulty node with the selected backup node.

[0105] In one implementation, the computer device can sort the backup nodes according to the second comprehensive matching degree. A higher second comprehensive matching degree indicates that the backup node better meets the requirements of the second demand threshold.

[0106] In one implementation, the computer device can replace the call to the faulty node with the call to the backup node in network management according to specific policies and conditions, thereby achieving the replacement of the faulty node.

[0107] In one implementation, the computer device can also forward the data stored in the faulty node to the backup node, so that the backup node can replace the faulty node.

[0108] In another implementation, the standby node can be a backup node for the failed node. In this case, the computer equipment can directly migrate function calls to the standby node to replace it.

[0109] One implementation method is that computer equipment can issue a combination of "service switching + resource configuration + policy optimization" commands to achieve the effect of a backup node replacing a faulty node.

[0110] For example, the instruction could include: switching the core transaction business to a backup node, or adjusting the backup cycle from 1 hour to 30 minutes.

[0111] In this example, by identifying the resilience of a faulty node, a backup node for that faulty node is determined. Based on the resilience identification code of the backup node, methods for users to replace the backup node of the faulty node are selected, thereby enabling rapid maintenance of faulty nodes in the network, ensuring the normal operation of the network, improving monitoring efficiency, and enhancing network management efficiency.

[0112] In one example, after replacing the faulty node, the computer equipment can continuously monitor the replaced standby node to determine whether the standby node is functioning correctly. This process may include: S205. After the backup node completes the first replacement duration, obtain the elastic identification code of the backup node.

[0113] For example, after the standby node starts up and completes the operation of replacing the faulty node, the computer device can obtain a new resilience identifier from the standby node that replaced the faulty node within a first period of time.

[0114] In one implementation, the computer device can request a new elastic identification code from the backup node after the first period of replacement is completed.

[0115] In another implementation, the computer device can obtain the new resilient identification code uploaded by the backup device after the first period of replacement is completed. The backup device can then upload the resilient identification code according to a preset upload period after the replacement is complete. This upload period can be less than or equal to the first period of replacement.

[0116] In one implementation, the first duration can be a pre-set time length. For example, it can be 10s, 20s, 30s, etc.

[0117] S206. If the full-dimensional real-time indicators in the elastic identification code of the standby node are not within the range indicated by the first demand threshold corresponding to the network domain to which the standby node belongs, then the standby node replacement is determined to have failed. Otherwise, the standby node replacement is determined to have succeeded.

[0118] For example, after obtaining the resilience identifier of the backup node, the computer device can parse it to obtain the full-dimensional real-time indicators. The computer device can compare the first demand indicator of the network domain corresponding to the backup node with the full-dimensional real-time indicators to determine whether the full-dimensional real-time indicators of the backup node are within the range indicated by the first demand threshold. If they are, it means that the backup node has been replaced. Otherwise, if any one or more indicators are found to be outside the threshold range, it means that the replacement has failed.

[0119] In one implementation, if the replacement of the backup node fails, the computer equipment can trigger subsequent operations such as notifying maintenance personnel.

[0120] In another implementation, if the backup node replacement fails, the computer equipment can trigger a secondary scheduling. The computer equipment can select the backup node with the highest second comprehensive index among the unselected backup nodes based on the second comprehensive index, and use that backup node to replace the failed node.

[0121] In one implementation, if the backup node is successfully replaced, the computer device can trigger subsequent operations such as updating the network topology.

[0122] In this example, after the backup node is replaced, its elastic identification code is obtained after a first period of time. Based on the comparison results of real-time indicators across all dimensions and the first demand threshold, an accurate method is used to determine whether the backup node replacement was successful. This achieves continuous network detection, improves network detection efficiency, and ensures network effectiveness.

[0123] In one example, the computer device can also agree with each network node on the update efficiency of the resilient identifier. This process may include: S301. Calculate the probability of failure of each network node within a preset time period.

[0124] For example, the computer device first determines a preset duration. This preset duration can be set according to actual needs, such as one day, one week, or one month. Then, the computer device collects fault record information of each network node within this preset duration and calculates the probability of failure of that network node.

[0125] In one implementation, the computer device can statistically analyze the fault records of each network node and calculate the number of times each node fails within a preset time period. The computer device can then divide the number of failures by the preset time period to obtain the probability of each network node failing within the preset time period.

[0126] In one implementation, the computer device can also set different weights for different fault types.

[0127] S302. Configure the elastic identification code reporting cycle of network nodes according to the probability of network node failure.

[0128] For example, after obtaining the failure probability of each network node, the computer device assigns different elastic identification code reporting periods to network nodes with different failure probabilities according to a preset configuration strategy. Then, the computer device can send this reporting period to different network nodes, enabling these network nodes to automatically report at different periods.

[0129] Generally, network nodes with a higher probability of failure have shorter reporting cycles for their Resilience Identifiers (RCIs) to obtain their status information more promptly. Conversely, network nodes with a lower probability of failure have relatively longer reporting cycles to reduce unnecessary network overhead.

[0130] In one implementation, the configuration strategy can include a mapping table of fault occurrence probability and reporting cycle.

[0131] In one implementation, the computer device can also integrate information such as the probability of the fault occurrence with business frequency, access frequency, and business sensitivity to determine the reporting cycle.

[0132] For example, the reporting cycle for core network nodes can be 10 seconds per report, while the reporting cycle for ordinary network nodes can be 60 seconds per report.

[0133] In this example, the computer device can dynamically adjust the reporting frequency of status information based on the node failure risk and optimize the allocation of network monitoring resources by statistically analyzing the probability of failure of each network node within a preset time period and configuring the elastic identification code reporting cycle accordingly, thereby further improving the network monitoring efficiency.

[0134] In one example, the computer device can also parse, store, and display the resilient identifier. Specifically, this includes: S401. Analyze the network domain encoding, node encoding, and full-dimensional real-time indicators in the elastic identification code. Then, the computer equipment can classify and store each real-time indicator according to its dimension, establishing a mapping database of "elastic identification code - real-time indicators".

[0135] S402. Generate the "Network Resilience Lifecycle Dashboard". This "Network Resilience Lifecycle Dashboard" can display the stored indicators in four dimensions: prevention, resilience, recovery, and adaptation.

[0136] For example, the prevention zone in this "Network Resilience Lifecycle Dashboard" can be used to display MTBF trends and real-time attack prevention rate values.

[0137] The availability zone can be marked with colors to indicate the status of indicators. For example, green indicates that the target has been met, yellow indicates a warning, and red indicates that the threshold has been exceeded. For instance, when the value of the core business availability indicator is less than 99.9%, the corresponding position of the indicator in the availability zone can be marked in yellow.

[0138] The recovery area can display graphs generated by tracking MTTR and RPO / RTO achievement rates, and then show the recovery progress based on these graphs.

[0139] The adaptive area can be used to display resource reallocation latency, policy update frequency, and to show the evaluation results of long-term resilience.

[0140] S403. Based on this "Network Resilience Full-Lifecycle Dashboard", trend analysis and threshold warnings are performed on various dimensions.

[0141] For example, if a certain indicator is outside the security threshold range, the computer device can trigger a graded alarm and push it to the operation and maintenance team based on the security level corresponding to the indicator and the degree of difference between the indicator and the security threshold.

[0142] For example, the security threshold for attack prevention rate can be 99%. When the attack prevention rate is less than 99%, it can be determined that the attack prevention rate is not within the security threshold range.

[0143] In one implementation, the computer device can calculate the difference between the attack prevention rate and 99%. The computer device can then project this difference onto the corresponding anomaly level to determine different handling methods. For example, the anomaly level can include normal, minor anomaly, and major anomaly. A minor anomaly might correspond to a yellow alert, while a major anomaly might correspond to a red alert. For instance, a red alert might be pushed to the core operations team, while a yellow alert might be pushed to the regional operations team.

[0144] S404. Using the node code in the elasticity identifier, determine the network node corresponding to the metric that is outside the security threshold range, and issue a temporary control command to that network node. This temporary control command is used to instruct optimization of the dimension corresponding to the metric. For example, if the metric outside the security threshold range is the attack prevention rate, this temporary control command can be used to instruct the firewall rules to be upgraded.

[0145] This example demonstrates how to establish a mapping database by parsing elastic identification codes, generate a full-cycle network elasticity dashboard, perform trend analysis and threshold warnings, and identify corresponding network nodes based on abnormal indicators and issue temporary control commands. This enables network visualization monitoring, hierarchical warnings, and targeted optimization control based on the full-cycle network elasticity, improving network monitoring efficiency, network management efficiency, and interactive efficiency during the monitoring process. It also enhances the user experience for administrators and reduces management difficulty.

[0146] In one example, the resilient identifier of the network node can be a three-layer unified modeling model of "domain-node-capability". The top layer is the network domain layer, used to configure network domain coding. The middle layer is the network node layer, used to configure network node coding. The bottom layer is the resilient capability layer, used to write real-time metrics for the network node.

[0147] For example, the network domain layer is used to macroscopically divide the network, assigning a 16-bit binary network domain code to each network domain. This network domain code is a unique "domain ID." This network domain code is used to distinguish the flexible management boundaries of different network types, such as backbone domains, edge domains, and data center domains. For example, the backbone domain code could be 0001000000000001.

[0148] The network node layer is used to write the network node code for this network node. This network node code can be a unique 16-bit binary "node ID". This network node code is used to distinguish several physical or logical nodes contained in each network domain. This network node code serves as the basic carrier for resilient state acquisition and management. For example, the code for router A is 0010000000000001.

[0149] The resilience capability layer is used to encode the dimension of each dimension included in the full-dimensional real-time metrics written to each network node. This dimension code is called the "capability sub-ID". Each dimension code can uniquely correspond to one dimension. Each dimension code can be 8 bits. For example, when four dimensions are included, the dimensions can be encoded as follows: prevention dimension is encoded as 00000001, resilience dimension is encoded as 00000010, recovery dimension is encoded as 00000100, and adaptation dimension is encoded as 00001000.

[0150] In one implementation, the three-layer unified modeling model of "domain-node-capability" can be represented by the following formula: in, For network domain encoding. Encoding for network nodes. R1-R4 correspond to four sets of metrics across four dimensions. to These correspond to the dimensional codes for prevention, resilience, recovery, and adaptation. MTBF represents Mean Time Between Failures (MTBF). AttackProtectionRate represents the attack prevention rate. FaultDetectionLatency represents the fault detection latency. MTTD represents the Mean Time to Diagnose (MTTD). MTTR represents the Mean Time to Recover (MTTR). RPOAchievementRate represents the recovery point target achievement rate. ResourceReallocationLatency represents the resource reallocation latency.

[0151] In one implementation, the three-layer unified modeling model of the elastic identification code—"domain-node-capability"—can be a 64-bit binary tree structure. The root node is the network domain ID, the first layer is the network node ID, the second layer is four elastic capability sub-IDs, and the third layer is the quantitative indicator corresponding to each capability sub-ID, forming a full-link positioning relationship of "domain-node-capability-indicator". This structure supports "one-click traceability," allowing for layer-by-layer positioning of the domain, node, and capability dimensions to ultimately obtain specific indicator values. Optionally, this tree structure can be as follows: Figure 2 As shown.

[0152] In one implementation, the resilient identifier includes an identifier space and a metadata space, which are linked by a "capability sub-ID". The identifier space stores 64-bit data as shown in Table 1. The metadata space stores specific real-time metrics. Computer devices can then use this... to This corresponds to a specific location in the metadata space, thereby obtaining the real-time metrics of a network node.

[0153] In one implementation, the identifier space is responsible for locating network nodes and their full-dimensional real-time metrics. The allocation of the 64-bit flexible identifier bits within the identifier space can be shown in Table 1.

[0154] Table 1 The first 16 digits are the network domain ID ( ), used to distinguish different network domains; Bits 17-32: Network Node ID ( (), used to distinguish different physical or logical nodes within a domain; Positions 33-40: Sub-ID of preventative capability ( ), locate pre-failure indicators such as MTBF and attack prevention rate; 41-48: Sub-ID of tolerance ( ), locate fault detection delays, core business availability and other fault indicators; 49th-56th digits: Sub-ID of recovery capability ( ), locate post-fault indicators such as MTTR and RPO achievement rate; Bits 57-64: Adaptive capability sub-ID ( ), and optimize indicators such as resource reallocation delay and policy update frequency after recovery; In one implementation, each real-time metric in the metadata space can correspond to a storage area. The metadata space can store the specific metric measurement value corresponding to each capability sub-ID. For example, IDres1 is associated with MTBF=1000h, attack prevention rate=99.5%, etc. If the real-time metric is not collected or is not used, it will not be stored in that storage area, but the storage area for the real-time metric will be reserved to ensure the compatibility and scalability of the elastic identification code.

[0155] In one implementation approach, the full-dimensional real-time metrics can specifically include four core dimensions: prevention, resilience, recovery, and adaptation. Prevention dimension metrics include: Mean Time Between Failures (MTBF) and attack prevention rate. Resilience dimension metrics include: fault detection latency, fault detection accuracy, Mean Time To Diagnostic (MTTD), core business availability, functional degradation level, network speed decrease rate during an event, throughput decrease rate, and latency increase level. Recovery dimension metrics include: Mean Time To Recovery (MTTR), Recovery Point Objective (RPO) achievement rate, and Recovery Time Objective (RTO) achievement rate. Adaptation dimension metrics include: resource reallocation latency, elastic policy update frequency, and business continuity assurance rate.

[0156] In one implementation, all metrics are obtained through a three-stage process: “device log collection - monitoring system reporting - simulation test verification”, to ensure data accuracy.

[0157] The data metrics and their acquisition methods are shown in Table 2, including: In one implementation, the process for acquiring these comprehensive real-time metrics may include: device log collection, monitoring system reporting, and simulation testing verification. Device log collection can obtain raw logs (such as fault records and traffic data) from devices such as routers, firewalls, and servers. Monitoring system reporting can collect real-time metrics (such as service availability and latency) using monitoring tools such as Zabbix and Prometheus. Simulation testing verification can verify the authenticity of the metrics through fault injection (such as simulating link interruptions) and functional testing (such as iperf speed testing and OpenVPN verification).

[0158] Figure 3A schematic diagram of a network node management system based on a network resilience model is provided as an embodiment of this application, such as... Figure 3 As shown, this network node management system based on a network resilience model runs on a computer device. The system includes: The unified modeling module 501 is used to construct a three-layer model of "domain-node-capability" and a quantitative measurement index system, and outputs the modeling results to the identification code generation module.

[0159] The identification code generation module 502 can generate a 64-bit binary tree-structured elastic identification code based on the modeling results of the three-layer unified modeling of "domain-node-capability". This modeling result can be associated with node indicator data and sent to the operation management module and data storage module.

[0160] The operation management module 503 may include a status monitoring submodule 5031, a fault location submodule 5032, and a recovery scheduling submodule 5033.

[0161] Among them, the status monitoring submodule 5031 can be deployed with a flexible identification code parsing unit, a cross-dimensional indicator storage unit, and a hierarchical alarm unit to realize indicator collection, parsing, display, and alarm.

[0162] In one implementation, the status monitoring submodule 5031 supports "one-click query of full-cycle elastic status". After inputting the elastic identification code of the target node, the status monitoring submodule 5031 can use the elastic identification parsing unit to parse the domain ID and node ID to locate the node, the cross-dimensional indicator storage unit to retrieve related indicators, generate an "elastic status dashboard" containing indicator trends and threshold comparisons, and mark indicators exceeding the threshold.

[0163] The built-in index matching degree calculation unit in the fault location submodule 5032 is used to extract the indexes in the identification code data, calculate the comprehensive matching degree, and lock the fault root node.

[0164] The recovery scheduling submodule 5033 includes a backup node screening unit, a priority sorting unit, and an instruction issuance and verification unit, which realizes dynamic priority recovery scheduling.

[0165] In one implementation, the instruction issuance and verification unit of the recovery scheduling submodule 5033 supports multiple types of instructions, including core business switching instructions, redundant resource configuration instructions, and elastic policy update instructions. During verification, a closed-loop process of "real-time indicator collection - threshold comparison - result feedback" is adopted to ensure that the recovery effect meets the standards.

[0166] The data storage module 504 uses a distributed database (such as HBase) to store three types of data. These three types of data include "identification code-metric" mapping data, "identification code-resource" mapping data, and historical node elasticity metric data. Furthermore, this module can also support data querying and retrieval by the operation management module.

[0167] The security enhancement module 505 is used to store network resilience identification codes and node resilience indicators using blockchain technology. It utilizes the immutability of blockchain to prevent identification code forgery or indicator tampering, thereby improving data credibility.

[0168] In this example, the adaptable cloud-edge-device integrated network architecture based on network elasticity supports deployment in multiple scenarios such as backbone networks, edge networks, and data center networks. It can work in conjunction with network devices through a software-defined networking (SDN) controller to achieve automatic distribution and execution of elastic policies. This mechanism covers three core processes: full-cycle status monitoring, fault location, and dynamic priority recovery scheduling, realizing a closed loop of elastic management.

[0169] One implementation example is a "financial data center network." This network comprises three domains: a core domain (carrying core transaction services), an edge domain (carrying user access services), and a storage domain (carrying data backup services), totaling 50 nodes. These 50 nodes can consist of 2 core routers, 10 edge gateways, and 38 storage servers. This application enables elastic, full-lifecycle network management based on this scenario.

[0170] S601, Administrator executes unified modeling implementation.

[0171] First, divide the network into domains and assign domain IDs. The core domain ID is 0001000000000001 (16-bit binary); the edge domain ID is 0010000000000001; and the storage domain ID is 0011000000000001.

[0172] Secondly, assign node IDs to each domain node: Core domain router A: 0001000000000001 (16-bit binary); Edge domain gateway B: 0002000000000001; Storage domain server C: 0003000000000001.

[0173] Finally, assign capability sub-IDs. Prevention capability sub-ID: 00000001; Endurance capability sub-ID: 00000010; Recovery capability sub-ID: 00000100; Adaptive capability sub-ID: 00001000.

[0174] S602. Network nodes can collect and verify metrics. A network node can obtain MTBF=1500h (device fault log statistics) from router A; monitor the core service availability of gateway B using Zabbix to achieve 99.98%; and test the throughput degradation rate of server C using iPerf3 to achieve 2% (fault injection simulating hard drive failure).

[0175] S603, Network nodes perform elastic identification code generation.

[0176] For example, generate a 64-bit identifier for router A: Bits 1-16: 0001000000000001 (core domain ID); Bits 17-32: 0001000000000001 (node ​​ID); Bits 33-40: 00000001 (prevention capability sub-ID); Bits 41-48: 00000010 (endurance capability sub-ID); Bits 49-56: 00000100 (recovery capability sub-ID); Bits 57-64: 00001000 (adaptive capability sub-ID).

[0177] The metadata space can be associated with metrics such as MTBF=1500h, attack prevention rate=99.8%, and fault detection latency=80ms.

[0178] S604. Computer equipment can perform full-cycle status monitoring.

[0179] The computer device can obtain the Resilience Identifier (RII) reported by Router A every 10 seconds. The computer device can parse the RII. The computer device can display the MTBF trend (stable for 1500 hours) in the "Prevention Zone" of the dashboard, and mark the fault detection latency in the "Tolerance Zone" as 80ms (green, compliant, threshold ≤100ms).

[0180] S605, Fault Location.

[0181] When gateway B's service is interrupted at a certain moment, the computer device can extract the full-dimensional real-time indicators R from the last acquired elasticity identifier: (attack prevention rate = 98.5%, fault detection latency = 150ms, MTTR = 40s, resource reallocation latency = 600ms). Additionally, the computer device can obtain the first demand thresholds D for each network: (≥99%, ≤100ms, ≤30s, ≤500ms).

[0182] Computer devices can calculate the matching degree in various dimensions: Prevention dimension (positive indicator): R1=98.5% <D1=99%,max(R1)=100%,min(R1)=99%, Tolerance dimension (negative index): R² = 150ms > D² = 100ms, max(R²) = 200ms. Recovery dimension (negative index): R3=40s>D3=30s, M(R3,D3)≈0.731; Adaptive dimension (negative index): R4=600ms>D4=500ms, M(R4,D4)≈0.741; The computer device can calculate the first comprehensive matching degree of network nodes in each dimension: M(R,D)=0.15×0.606+0.3×0.716+0.4×0.731+0.15×0.741≈0.705.

[0183] When the first comprehensive matching degree of the network node is the minimum value of the edge domain, the network node can be regarded as a fault node and the fault node can be identified as the root node of the fault.

[0184] S606, Dynamic Priority Recovery Scheduling.

[0185] The computer equipment can parse the elastic identification code of gateway B and determine that its service type is "user access service". The computer equipment can obtain the second demand threshold D of gateway B (≥99%, ≤100ms, ≤30s, ≤500ms).

[0186] The computer can select two backup nodes (Gateway D and Gateway E) from the database. The computer can then calculate a second comprehensive score for each backup node. Gateway D has a score of 0.85, and Gateway E has a score of 0.78. The computer can prioritize scheduling Gateway D to replace network B.

[0187] S607. Issue instructions and switch the user access service to gateway D, and adjust firewall rules to improve the attack prevention rate.

[0188] S608, after 10 seconds, retrieved the gateway D's identification code metadata, obtaining an attack prevention rate of 99.2%, a fault detection latency of 90ms, an MTTR of 25s, and a resource reallocation latency of 450ms. After confirming that all real-time metrics for gateway D met the standards, the fault was deemed resolved, and recovery verification was completed.

[0189] Figure 4 A structural diagram of a network node management device based on a network resilience model provided in this application embodiment is shown below. Figure 4 As shown, the network node management device 700 based on the network resilience model includes: The acquisition module 701 is used to acquire the elastic identification code reported by each network node in the faulty network; the elastic identification code includes the network domain code of the network domain where the network node is located, as well as the real-time indicators of the network node in all dimensions. The monitoring module 702 is used to determine the network domain to which each network node belongs based on the network domain code of each network node, and obtain the preset first demand threshold in the network domain; to match the full-dimensional real-time indicators of each network node with the first demand threshold to calculate the first comprehensive matching degree of each network node; and to determine the network node with the lowest first comprehensive matching degree as the fault node.

[0190] In one example, the full-dimensional real-time metrics include at least one real-time metric for at least one dimension; monitoring module 702 is used for: If the real-time indicator is within the range indicated by the first demand threshold, then the indicator matching degree of the real-time indicator is determined to be the highest preset matching value; if the real-time indicator of a dimension is not within the range indicated by the first demand threshold, then the indicator matching degree of the real-time indicator is calculated based on the difference between the real-time indicator and the first demand threshold; the weighted sum of the indicator matching degrees of multiple real-time indicators of each dimension is calculated to obtain the dimension matching degree of each dimension; the weighted sum of the dimension matching degrees of each dimension of the network node is calculated to obtain the first comprehensive matching degree of the network node.

[0191] In one example, monitoring module 702 is used for: Obtain the difference between the real-time metric and the first requirement threshold; calculate the proportion of the difference within the range of the real-time metric values ​​for the dimension, and obtain the power; use the power to exponentiate the natural constant to obtain the metric matching degree for the dimension.

[0192] In one example, monitoring module 702 is used for: Based on the network domain code of the faulty node, determine the network domain to which the faulty network belongs, and obtain the backup node and the preset second demand threshold in the network domain to which the faulty node belongs; use the real-time indicators of the backup node in all dimensions to match with the second demand threshold, and calculate the second comprehensive matching degree of the backup node; schedule the backup node with the highest second comprehensive matching degree to replace the faulty node.

[0193] In one example, monitoring module 702 is used for: If the real-time metrics of the standby node are not within the range indicated by the second demand threshold, then the standby node is deleted.

[0194] In one example, monitoring module 702 is used for: After the backup node completes the first replacement duration, obtain the elasticity identification code of the backup node; if the full-dimensional real-time indicators in the elasticity identification code of the backup node are not within the range indicated by the first demand threshold corresponding to the network domain to which the backup node belongs, then the backup node replacement is determined to have failed; otherwise, if the full-dimensional real-time indicators of the backup node are within the range indicated by the first demand threshold, then the backup node replacement is determined to have succeeded.

[0195] In one example, module 701 is used for: Calculate the probability of failure for each network node within a preset time period; configure the elastic identification code reporting cycle for network nodes based on the probability of failure.

[0196] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0197] Figure 5 A structural diagram of a computer device provided in an embodiment of this application, such as... Figure 5 As shown, the computer device 800 includes one or more processors 801 and a memory 802. The processor 801 may be a central processing unit, a network processor, or a combination thereof. The memory 802 stores instructions executable by at least one processor 801 to cause the at least one processor 801 to perform the methods shown in the above embodiments. The computer device also includes a communication interface 803 for communicating with other network nodes.

[0198] This application also provides a computer-readable storage medium. The methods described in this application embodiment can be implemented in hardware or firmware, or exist in the form of computer code. They can be recorded on the storage medium or downloaded from a remote medium to a local storage location. Storage media include magnetic disks, optical disks, various types of storage memory, hard disks, and combinations thereof. It is understood that hardware such as computers contains storage components, and when software or code is accessed and executed by it, the methods shown in the above embodiments are implemented.

[0199] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any embodiment of this application.

[0200] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A network node management method based on a network resilience model, characterized in that, Applied to a management node, the method includes: Obtain the resilience identification code reported by each network node in the faulty network; the resilience identification code is generated by the network node based on a preset three-layer architecture model, and the resilience identification code includes the network domain code of the network domain where the network node is located in the first layer architecture, and the full-dimensional real-time indicators of the network node in the third layer architecture. Based on the network domain code of each network node, determine the network domain to which each network node belongs, and obtain the preset first demand threshold in the network domain; The first comprehensive matching degree of each network node is calculated by matching the full-dimensional real-time indicators of each network node with the first demand threshold. The network node with the lowest overall matching degree is identified as the faulty node.

2. The method according to claim 1, characterized in that, The full-dimensional real-time indicators include at least one real-time indicator for at least one dimension; the full-dimensional real-time indicators of each network node are matched with the first demand threshold to calculate the first comprehensive matching degree of each network node, including: If the real-time indicator is within the range indicated by the first demand threshold, then the indicator matching degree of the real-time indicator is determined to be the highest preset matching value. If the real-time indicator of the dimension is not within the range indicated by the first demand threshold, the indicator matching degree of the real-time indicator is calculated based on the difference between the real-time indicator and the first demand threshold. Calculate the weighted sum of the indicator matching degrees of multiple real-time indicators for each dimension to obtain the dimension matching degree of each dimension; The weighted sum of the dimensional matching degrees of each dimension of the network node is calculated to obtain the first comprehensive matching degree of the network node.

3. The method according to claim 2, characterized in that, The indicator matching degree of the real-time indicator is calculated based on the difference between the real-time indicator and the first demand threshold, including: Obtain the difference between the real-time indicator and the first demand threshold; Calculate the percentage of the difference within the range of values ​​of the real-time indicator of the dimension, and obtain the power. The natural constant is exponentialized using the power to obtain the index matching degree of the dimension.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: Based on the network domain code of the faulty node, determine the network domain to which the faulty network belongs, and obtain the backup node and the preset second demand threshold in the network domain to which the faulty node belongs; The second comprehensive matching degree of the backup node is calculated by matching the full-dimensional real-time indicators of the backup node with the second demand threshold. The backup node with the highest overall matching degree is scheduled to replace the faulty node.

5. The method of claim 4, wherein, Before calculating the second comprehensive matching degree of the backup node by matching the full-dimensional real-time indicators of the backup node with the second demand threshold, the method further includes: If the real-time metrics of the backup node are not within the range indicated by the second demand threshold, then the backup node is deleted.

6. The method of claim 4, wherein, The method includes: After the backup node completes the first replacement period, obtain the elastic identification code of the backup node; If the full-dimensional real-time indicators in the elastic identification code of the backup node are not within the range indicated by the first demand threshold corresponding to the network domain to which the backup node belongs, then the replacement of the backup node is determined to have failed. Otherwise, the replacement of the backup node is confirmed to be successful.

7. The method according to any one of claims 1-3, characterized in that, The method includes: Calculate the probability of failure for each of the network nodes within a preset time period; Configure the elastic identification code reporting cycle of the network node based on the failure probability of the network node.

8. A network node management device based on a network resilience model, characterized in that, Applied to a management node, the device includes: The acquisition module is used to acquire the elastic identification code reported by each network node in the faulty network. The elastic identification code is generated by the network node based on a preset three-layer architecture model. The elastic identification code includes the network domain code of the network domain where the network node is located in the first layer architecture, and the full-dimensional real-time indicators of the network node in the third layer architecture. The monitoring module is used to determine the network domain to which each network node belongs based on the network domain code of each network node, and obtain a preset first demand threshold in the network domain; match the full-dimensional real-time indicators of each network node with the first demand threshold to calculate the first comprehensive matching degree of each network node; and determine the network node with the lowest first comprehensive matching degree as the fault node.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.

10. A computer program product, characterised in that, Includes computer instructions for causing a computer to perform the method of any one of claims 1 to 7.