Intelligent machine room network rapid recovery method and system

By collecting network status information and constructing temporary transmission paths, the problem of unavailable preset backup paths was solved, enabling rapid and reliable recovery of the smart data center network under complex conditions and ensuring the continuous and stable operation of the network.

CN120856545BActive Publication Date: 2026-05-19SHENZHEN HUACHUANG INTELLIGENT ENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HUACHUANG INTELLIGENT ENG TECH CO LTD
Filing Date
2025-09-12
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing network fast recovery methods may fail when the network topology changes unexpectedly, as the preset backup path may become unavailable, leading to prolonged network recovery time or failure, and thus failing to guarantee the fast and reliable recovery of the smart data center network.

Method used

Collect operational status information from the smart data center network, determine the location and type of fault, construct a temporary network transmission path, and switch network traffic to the temporary path to achieve fast and reliable network recovery.

Benefits of technology

When the network topology changes, new transmission paths are constructed to avoid prolonged recovery time or failure, thus ensuring the continuous and stable operation of the smart data center network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856545B_ABST
    Figure CN120856545B_ABST
Patent Text Reader

Abstract

The application provides a kind of intelligent machine room network fast recovery method and system, the method comprises: the running state information of network equipment in intelligent machine room network is collected;If network failure occurs based on running state information, determine the network area range affected by network failure based on the fault location information and fault type information of network failure;Collect the connection relationship information between all network equipment in network area range and the available resource information of all network equipment;Transmission path planning is carried out based on network area range, connection relationship information and available resource information, to construct network temporary transmission path;The network traffic affected by network failure is switched from original transmission path to network temporary transmission path, to recover the network of intelligent machine room.The application realizes the fast and reliable recovery of intelligent machine room network under complex conditions, and guarantees the continuous and stable operation of intelligent machine room network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and specifically to a method and system for rapid network recovery in smart data centers. Background Technology

[0002] With the rapid development of information technology, data center networks are playing an increasingly critical role in various business operations. Network failures can lead to serious consequences such as business interruptions and data loss, causing significant losses to enterprises and organizations. Therefore, ensuring the stable operation of data center networks and enabling rapid recovery in the event of failures has become an important research direction in the field of network management.

[0003] Existing network rapid recovery methods are mainly based on preset backup paths. These methods plan and store multiple backup paths in advance when the network is running normally. When a network failure is detected, data transmission is quickly switched to the preset backup path to achieve rapid network recovery. Compared to traditional network recovery methods, which require recalculating the recovery path after a failure, making the process cumbersome and time-consuming, and unable to meet the business requirements for network recovery speed, the preset backup path method can effectively shorten the recovery time and improve the efficiency of network recovery.

[0004] However, when the network topology undergoes unexpected changes, the preset backup paths in the preset backup path method may become unavailable. In this case, it is necessary to find a new available path, which will lead to a longer network recovery time, or even recovery failure. This will result in the smart data center network being unable to recover quickly and reliably under complex conditions, thus failing to guarantee the continuous and stable operation of the smart data center network. Summary of the Invention

[0005] This invention provides a method and system for rapid recovery of smart data center networks, which enables rapid and reliable recovery of smart data center networks under complex conditions and ensures the continuous and stable operation of smart data center networks.

[0006] In a first aspect, the present invention provides a method for rapid network recovery in smart data centers, comprising:

[0007] Collect operational status information of network devices in the smart data center network;

[0008] If a network failure is determined based on the operational status information, then the network area affected by the network failure is determined based on the failure location information and failure type information.

[0009] Collect connection relationship information between all network devices within the network area and available resource information of all network devices;

[0010] Based on the network area range, the connection relationship information, and the available resource information, a transmission path is planned, and a temporary network transmission path is constructed.

[0011] The network traffic affected by the network failure is switched from the original transmission path to the temporary network transmission path to restore the network of the smart data center.

[0012] Secondly, the present invention also provides a smart data center network rapid recovery system, applied to the smart data center network rapid recovery method as described in the first aspect; the smart data center network rapid recovery system includes:

[0013] The first data acquisition module is used to collect the operating status information of network devices in the smart data center network;

[0014] The fault range prediction module is used to determine the network area affected by the network fault based on the fault location information and fault type information if the network fault is determined to have occurred based on the operating status information.

[0015] The second acquisition module is used to acquire connection relationship information between all network devices within the network area and available resource information of all network devices.

[0016] The transmission path planning module is used to plan transmission paths based on the network area, the connection relationship information, and the available resource information, and to construct temporary network transmission paths.

[0017] The transmission path switching module is used to switch network traffic affected by network failure from the original transmission path to the temporary network transmission path in order to restore the network of the smart data center.

[0018] Thirdly, the present invention also provides an electronic device, comprising: a memory for storing computer software programs; and a processor for reading and executing the computer software programs, thereby realizing the intelligent data center network rapid recovery method as described above.

[0019] Fourthly, the present invention also provides a non-transitory computer-readable storage medium storing a computer software program, which, when executed by a processor, implements the intelligent data center network rapid recovery method described above.

[0020] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the intelligent data center network rapid recovery method as described above.

[0021] The intelligent data center network rapid recovery method provided in this invention quickly locates the network area affected by the network fault based on the fault location and fault type information. Then, based on the connection relationship information between all network devices within the network area and the available resource information of all network devices, a new temporary network transmission path is constructed. Finally, the network traffic affected by the fault is switched through the temporary network transmission path to complete the rapid network recovery. This method enables the construction of an effective transmission path based on the real-time network status when unexpected changes in the network topology render the original preset backup path unavailable. It avoids prolonged recovery time or recovery failure due to reliance on preset paths, achieving rapid and reliable recovery of the intelligent data center network under complex conditions and ensuring the continuous and stable operation of the intelligent data center network. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating the intelligent data center network rapid recovery method provided in an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram of the structure of the intelligent data center network rapid recovery system provided in an embodiment of the present invention;

[0024] Figure 3 An embodiment diagram of the electronic device provided in this invention;

[0025] Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments.

[0028] The following description is provided to enable any person skilled in the art to implement and use the present invention. In this description, details are set forth for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be implemented without these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope accorded to the principles and features disclosed herein.

[0029] Optional, see below Figure 1 , Figure 1 This is a flowchart illustrating the intelligent data center network rapid recovery method provided by the present invention. In this embodiment, the executing entity of the intelligent data center network rapid recovery method is a network management system. The network management system can be understood as a manifestation of the intelligent data center network rapid recovery system. Therefore, the intelligent data center network rapid recovery method includes:

[0030] Step 10: Collect the operating status information of network devices in the smart data center network.

[0031] Optionally, the network management system uses monitoring agents, network traffic probes, and built-in management interfaces (such as Simplified Network Management Protocol) deployed in the smart data center network to collect real-time data from all network devices, including routers, switches, servers, and firewalls. The collected operational status information covers multiple key dimensions, including CPU utilization, memory usage, port traffic data (input and output rates), port connection status (connected, disconnected, or not connected), device temperature, and device runtime. The collection frequency can be set according to the importance of the devices and network stability requirements; generally, core devices are collected every 5 seconds, and non-core devices every 30 seconds, ensuring timely capture of changes in device operational status.

[0032] In one embodiment, assume the smart data center network includes a core router A, an aggregation switch B, an access switch C, and a server D. The network management system establishes connections with these devices through a Simple Network Management Protocol (SMMP) interface and collects information in real time. At a certain moment, the network management system collects data showing that the core router A has a CPU utilization of 85% (normal threshold is ≤70%) and a memory utilization of 65%; the input traffic rate of port 1 of aggregation switch B (connected to core router A) is 1.2Gbps, the output traffic rate is 1.1Gbps, and the port status shows as connected; port 3 of access switch C (connected to server D) suddenly shows as disconnected; the CPU utilization of server D is 30%, the memory utilization is 40%, and the device temperature is 42℃.

[0033] Step 20: If a network failure is determined based on the operating status information, then the network area affected by the network failure is determined based on the failure location information and failure type information.

[0034] Furthermore, the network management system analyzes the operational status information collected in step 10 and determines whether a network fault has occurred based on preset fault judgment rules (such as CPU utilization exceeding a threshold for 5 minutes, port connection being disconnected for more than 10 seconds, traffic rate being 0 and no data transmission on the port, etc.). Once a network fault is confirmed, the network management system extracts the fault location information (such as which port of which device) and fault type information (such as port connection interruption fault, device overload fault, link congestion fault, etc.). Continuing with the above embodiment, the network management system, through analysis, discovers that port 3 of access switch C is disconnected for more than 10 seconds. Combined with the information that server D cannot communicate with other devices, it determines that a "port connection interruption fault" has occurred.

[0035] Furthermore, the network management system invokes a pre-built fault structure tree, which is constructed based on network fault domain knowledge. The nodes in the tree include various fault types (such as port faults, device hardware faults, link faults, etc.) and network units (such as core network areas, aggregation network areas, access network areas, specific device locations, etc.). The edges between nodes represent the causal relationship between fault types (such as device hardware faults may lead to port faults) or the subordinate relationship between faults and network units (such as port faults being subordinate to the network unit corresponding to the device where the port is located).

[0036] Furthermore, the network management system performs matching and tracing in the fault structure tree based on the fault location information and fault type information. By analyzing the impact range of the fault type and the relationship between the network unit where the fault is located and other network units, the network area affected by the network fault is finally determined, as in steps 201 to 204.

[0037] Step 30: Collect connection relationship information between all network devices within the network area and available resource information of all network devices.

[0038] Furthermore, based on the network area affected by the network fault determined in step 20, the network management system performs targeted information collection on the connection relationship information between all network devices within that network area and the available resource information of all network devices.

[0039] For connection relationship information, such as collecting the physical and logical connections between devices, including whether there is a direct connection between intermediate network devices (such as other switches and routers in the area) and source devices (the data sending device before the failure, determined according to the original transmission path) and destination devices (the data receiving device before the failure); the number of historical connection interruptions and the total number of connections between intermediate network devices and source devices; and the number of historical connection interruptions and the total number of connections between intermediate network devices and destination devices.

[0040] For available resource information, such as the storage space utilization rate (used storage space / total storage space), computing power utilization rate (current CPU utilization rate / maximum tolerable CPU utilization rate) and port utilization rate (number of used ports / total number of ports) of each network device.

[0041] In one embodiment, the determined network fault-affected area is the access network area where access switch C is located. Network devices within this area include access switch C, server D, and intermediate switch E. The original transmission path is server D → access switch C → aggregation switch B, where the source device is server D and the destination device is aggregation switch B. The network management system collects connection relationship information: intermediate switch E has a direct connection with server D (through port 2) and a direct connection with aggregation switch B (through port 5); the historical connection interruption count between intermediate switch E and server D is 2, with a total of 100 connections; the historical connection interruption count between intermediate switch E and aggregation switch B is 1, with a total of 120 connections. Available resource information is collected: intermediate switch E has a storage space utilization rate of 30%, a computing power utilization rate of 25%, and a port utilization rate of 40%; access switch C, except for the faulty port 3, is not considered for the time being (because it is located in the core of the fault-affected area and port 3 is faulty).

[0042] Step 40: Based on the network area range, connection relationship information and available resource information, perform transmission path planning and construct temporary network transmission paths.

[0043] Furthermore, the network management system first determines whether the network area range determined in step 20 is larger than the preset range area (the preset range area can be set according to the network scale, such as including 5 or more core devices or covering 3 or more subnets is larger than the preset range area, otherwise it is smaller than the preset range area).

[0044] When the network area is smaller than the preset area, the network management system uses a single relay point method for transmission path planning. That is, it selects a qualified device from the intermediate network devices within the network area as a relay point and constructs a temporary network transmission path from the source device to the relay point to the destination device.

[0045] When the network area is larger than the preset area, the network management system adopts a multiple relay point method. Based on the network topology and device distribution, the transmission path is divided into multiple segments. Each segment selects a suitable relay point to ensure stable connection and sufficient resources between segments, thus constructing a temporary network transmission path.

[0046] Step 50: Switch the network traffic affected by the network failure from the original transmission path to a temporary network transmission path to restore the network of the smart data center.

[0047] Furthermore, after constructing the temporary network transmission path, the network management system generates a traffic switching instruction, which contains detailed routing information of the temporary transmission path (such as source device IP address, destination device IP address, relay device IP address, port mapping relationship, etc.).

[0048] Furthermore, the network management system sends instructions to all network devices (source devices, relay devices, and destination devices) involved in the temporary transmission path through network control protocols (such as Open Stream Protocol), configures the routing tables and forwarding rules of these devices, and enables the devices to identify and forward data according to the temporary transmission path.

[0049] Furthermore, during the switchover process, the network management system monitors the traffic switching in real time, ensuring that traffic on the original transmission path gradually decreases while traffic on the temporary transmission path gradually increases, thus preventing traffic interruptions or data loss. Once it is confirmed that traffic on the temporary transmission path is transmitting stably and that traffic on the original transmission path has been completely switched to the temporary transmission path, the smart data center network is restored.

[0050] In one embodiment, the network management system generates a traffic switching instruction, which includes the IP address of the source device server D (192.168.1.100), the IP address of the destination device aggregation switch B (192.168.2.1), the IP address of the relay device intermediate switch E (192.168.1.200), and port mapping relationships (port 8080 of server D is mapped to port 2 of intermediate switch E, and port 5 of intermediate switch E is mapped to port 1 of aggregation switch B). The network management system sends the instruction to server D, intermediate switch E, and aggregation switch B via the Open Flow Protocol. After receiving the instruction, server D redirects the traffic originally destined for access switch C to port 2 of intermediate switch E; intermediate switch E configures its routing table to forward the traffic received from port 2 to aggregation switch B via port 5; aggregation switch B receives the traffic from intermediate switch E and forwards it according to the normal route. Real-time monitoring by the network management system shows that after one minute, the traffic from server D to aggregation switch B has been completely transmitted through the temporary transmission path, and the data transmission is stable with no packet loss, and the smart data center network has returned to normal operation.

[0051] This invention enables rapid location of the network area affected by a network fault based on the fault location and fault type information. Then, based on the connection relationships between all network devices within the affected area and the available resources of all network devices, a new temporary network transmission path is constructed. Finally, the network traffic affected by the fault is switched through this temporary transmission path, achieving rapid recovery of the smart data center network. This allows for the construction of an effective transmission path based on real-time network conditions when unexpected changes in the network topology render the original backup path unavailable. This avoids prolonged recovery time or recovery failure due to reliance on preset paths, achieving rapid and reliable recovery of the smart data center network under complex conditions and ensuring its continuous and stable operation.

[0052] In one embodiment, steps 201 to 205 include:

[0053] Step 201: Match the fault type information with the nodes in the pre-built fault structure tree to obtain the target fault node.

[0054] Optionally, the network management system compares fault type information (such as port connection interruption fault, device hardware fault, link congestion fault, etc.) with nodes in the fault structure tree one by one through string matching or a preset fault type mapping relationship. When the fault type descriptions of the two are completely consistent or conform to the preset mapping rules, the node can be determined as the target fault node corresponding to the current fault.

[0055] Continuing with the above embodiment, the determined network fault type information is "port connection interruption fault". There is a "port connection interruption fault" node in the fault structure tree. The fault type information is matched with the node in the tree. Since the two are completely consistent, the "port connection interruption fault" node is determined to be the target fault node.

[0056] Step 202: Based on the initial branch path of the target fault node in the fault structure tree, start the same branch fault traversal. The traversal direction is to extend from the target fault node to the root node and leaf node of the fault structure tree to obtain the same branch traversal path of the target fault node in the fault structure tree.

[0057] Furthermore, after identifying the target fault node, the network management system initiates a same-branch fault traversal, starting from the initial branch path of the target fault node in the fault structure tree. During the traversal, the system extends simultaneously from the target fault node towards both the root node and leaf nodes of the fault structure tree. Traversing towards the root node allows tracing the higher-level fault category to which the fault type belongs, while traversing towards the leaf nodes allows obtaining more specific fault manifestations contained within the fault type. Through bidirectional traversal, the complete same-branch traversal path of the target fault node in the fault structure tree is finally obtained.

[0058] Continuing with the above embodiment, the target fault node is "port connection interruption fault," and its initial branch path in the fault structure tree is "network fault → link fault → port fault → port connection interruption fault." The network management system traverses towards the root node, from "port connection interruption fault" to "port fault," then to "link fault," and finally to "network fault"; traversing towards the leaf nodes, there are no more specific leaf nodes under "port connection interruption fault." Therefore, the traversal path within the same branch is "network fault → link fault → port fault → port connection interruption fault."

[0059] Step 203: Based on the network units associated with each node in the same branch traversal path and the fault location information, filter to obtain the units in the same branch that are directly affected by the fault.

[0060] Furthermore, the network management system obtains the network units associated with each node in the same branch traversal path based on the association relationship between nodes and network units. The association relationship is preset based on network fault domain knowledge when constructing the fault structure tree. That is, each fault type node corresponds to a specific network unit that it may affect (such as the area where the device is located, the link coverage area, etc.).

[0061] Furthermore, the network management system, in conjunction with the fault location information determined in step 10 (specifically, which port of which device experienced the fault), filters these associated network units. The filtering rule is: retain network units directly related to the fault location information, and remove network units not directly related to the fault location, ultimately obtaining the affected units on the same branch that are directly affected by the fault.

[0062] Continuing with the above embodiment, the network unit associated with the node "Port Connection Interruption Failure" in the same branch traversal path is "the subnet unit where port 3 of access switch C is located," the network unit associated with "Port Failure" is "the access network unit where access switch C is located," the network unit associated with "Link Failure" is "the link coverage unit between access switch C and server D," and the network unit associated with "Network Failure" is "the entire smart data center network unit." Combining the fault location information "port 3 of access switch C" for filtering, "the entire smart data center network unit," which is too broad and has no direct association, is removed. The "subnet unit where port 3 of access switch C is located," "the access network unit where access switch C is located," and "the link coverage unit between access switch C and server D" are retained; these are the units affected by the same branch.

[0063] Step 204: Based on the traversal path of the same branch and the edges in the fault structure tree, perform cross-branch association positioning to obtain the cross-branch affected area under the fault.

[0064] Furthermore, the network management system performs cross-branch association positioning based on the same branch traversal path and the edges in the fault structure tree to obtain the cross-branch affected area under the fault, as specifically in steps 2041 to 2043.

[0065] Step 205: Based on the fusion of the affected units in the same branch and the affected areas across branches, the network area affected by the network fault is obtained.

[0066] Furthermore, the network management system merges the affected units within the same branch and the affected areas across branches. During the merging process, duplicate network units are removed. For network units with an inclusion relationship, units with more accurate ranges are retained. For adjacent or connected network units, they are integrated into a continuous area to obtain the network area affected by the network fault.

[0067] In one embodiment, the affected units within the same branch are "the subnet unit where port 3 of access switch C is located", "the access network unit where access switch C is located", and "the link coverage unit between access switch C and server D". The affected areas across branches are "the transmission area unit between the access network where access switch C is located and the aggregation network where aggregation switch B is located" and "the downlink connection area unit of aggregation switch B". After merging and removing duplicate parts, the resulting network area is: the entire area centered on the access network unit where access switch C is located, including the subnet unit where port 3 of access switch C is located, the link coverage unit between access switch C and server D, the transmission area unit between the access network where access switch C is located and the aggregation network where aggregation switch B is located, and the downlink connection area unit of aggregation switch B.

[0068] The embodiments of the present invention can start from the fault type and use node matching, path traversal, cross-branch positioning, and fusion of the same branch and cross-branch to accurately identify the units directly affected by the fault in the same branch and the areas affected by the fault in the middle of the cross-branch, thus avoiding the omission of the fault-affected areas.

[0069] In one embodiment, steps 2041 to 2043 include:

[0070] Step 2041: Traverse the fault tree structure based on the same branch traversal path and the edges in the fault tree structure to determine the cross-related nodes in the fault tree structure that have cross-related relationships with the same branch nodes.

[0071] Optionally, the network management system traverses a pre-built fault structure tree based on all nodes in the same branch traversal path (such as network faults, link faults, port faults, port connection interruption faults, etc.), focusing on the relationships represented by the edges between nodes in the fault structure tree. These edges include causal relationship edges between fault types and subordinate relationship edges between faults and network units.

[0072] Therefore, the network management system checks each node in the same branch traversal path one by one to determine whether there is a direct connection edge (i.e., a cross connection edge) between it and other non-branch nodes. When a non-branch node has a direct connection edge with a node in the same branch, that non-branch node is a cross connection node that has a cross connection with a node in the same branch.

[0073] Continuing with the above example, the nodes in the same branch traversal path are "Network Failure," "Link Failure," "Port Failure," and "Port Connection Interruption Failure." The network management system, traversing the fault structure tree, finds a causal relationship between the "Port Connection Interruption Failure" node and the "Data Transmission Delay Failure" node (port connection interruption causes data retransmission, leading to delay), and a causal relationship between the "Link Failure" node and the "Insufficient Network Bandwidth Failure" node (link failure may cause some links to become unavailable, resulting in insufficient bandwidth). Since neither "Data Transmission Delay Failure" nor "Insufficient Network Bandwidth Failure" belongs to the nodes in the same branch traversal path, these two nodes are identified as cross-related nodes.

[0074] Step 2042: Based on each cross-related node as the starting node of the cross-branch, traverse the fault structure tree to determine the cross-branch path corresponding to each cross-branch starting node, and filter out the effective paths with actual fault propagation possibility based on each cross-branch path to obtain the cross-branch propagation path.

[0075] Furthermore, the network management system uses each cross-related node as a cross-branch starting node and initiates a traversal process in the fault structure tree. The traversal direction extends from the cross-branch starting node towards the leaf nodes of the fault structure tree to obtain more specific fault type nodes contained under that starting node, forming the initial cross-branch path.

[0076] Furthermore, the network management system formulates a fault propagation probability judgment rule based on network fault domain knowledge and historical fault data (e.g., if a certain fault type node has a historical probability of actually occurring due to the failure of the starting node ≥10%, it is considered to have actual propagation probability). Based on this rule, each initial cross-branch path is screened, retaining paths with actual fault propagation probability and eliminating paths without propagation probability, thus finally obtaining the cross-branch propagation path.

[0077] Continuing with the above embodiment, the cross-related nodes "data transmission delay fault" and "insufficient network bandwidth fault" serve as the starting nodes for cross-branch propagation. Traversing with "data transmission delay fault" as the starting node yields the initial cross-branch path: "data transmission delay fault → critical business data delay fault"; traversing with "insufficient network bandwidth fault" as the starting node yields the initial cross-branch path: "insufficient network bandwidth fault → non-critical business bandwidth limited fault". Based on historical data, the actual probability of "port connection interruption fault" leading to "data transmission delay fault" further causing "critical business data delay fault" is 25% (>10%); while the actual probability of "link fault" leading to "insufficient network bandwidth fault" causing "non-critical business bandwidth limited fault" is 5% (<10%). Therefore, the cross-branch propagation path obtained after filtering is "data transmission delay fault → critical business data delay fault".

[0078] Step 2043: Based on the network regions associated with each node in the cross-branch diffusion path and the fault location information, determine the cross-branch affected area under the cross-branch.

[0079] Furthermore, the network management system obtains the network regions associated with each node in the cross-branch propagation path. These associations are preset when constructing the fault structure tree, meaning that each fault type node corresponds to a specific network region (such as a transmission area, service area, etc.) that it may affect. Further, the network management system analyzes these associated network regions based on the fault location information determined in step 10 (which port of which device experienced the fault). During the analysis, the network management system retains network regions with direct or indirect connections to the fault location and removes network regions with no connection to the fault location, ultimately determining the cross-branch affected areas under the cross-branch structure.

[0080] Continuing with the above embodiment, the cross-branch propagation path is "data transmission delay fault → critical business data delay fault," where the network area associated with the "data transmission delay fault" is the "transmission area unit between the access network where access switch C is located and the aggregation network where aggregation switch B is located," and the network area associated with the "critical business data delay fault" is the "area unit where the critical business server connected to aggregation switch B is located." Combining the fault location information "port 3 of access switch C," the network management system analysis reveals that the "transmission area unit between the access network where access switch C is located and the aggregation network where aggregation switch B is located" is directly connected to the fault location, while the "area unit where the critical business server connected to aggregation switch B is located" is indirectly connected to the fault location through aggregation switch B. Therefore, the determined cross-branch impact areas are the "transmission area unit between the access network where access switch C is located and the aggregation network where aggregation switch B is located" and the "area unit where the critical business server connected to aggregation switch B is located."

[0081] The embodiments of the present invention can accurately identify cross-related nodes starting from the same branch node, filter out cross-branch paths with actual diffusion potential, and determine the actual affected cross-branch area in combination with the fault location. Therefore, it takes into account the potential diffusion paths and correlations of the fault between different branches, and avoids the omission or false judgment of the cross-branch affected area.

[0082] In one embodiment, steps 401 to 403 include:

[0083] Step 401: Determine the source device and destination device based on the original transmission path.

[0084] Optionally, the network management system analyzes the original transmission path, identifies the initial sending device as the source device, and the final receiving device as the destination device.

[0085] In one embodiment, it is known that before the network failure, the original transmission path was "Server D → Access Switch C → Aggregation Switch B". Therefore, the network management system identifies the starting sending device as Server D as the source device and the final receiving device as Aggregation Switch B as the destination device.

[0086] Step 402: If the network area is less than or equal to the preset fault range, then a transmission path is planned for a single transit point based on the source device, destination device, connection relationship information, and available resource information, and a temporary network transmission path is constructed.

[0087] Furthermore, if the network area is less than or equal to the preset fault range, the network management system will plan the transmission path of a single transit point based on the source device, destination device, connection relationship information, and available resource information, and construct a temporary network transmission path, as described in steps 4021 to 4024.

[0088] Step 403: If the network area is larger than the preset fault range, then based on the source device, destination device, connection relationship information and available resource information, a transmission path planning with multiple transit points is performed to construct a temporary network transmission path.

[0089] Furthermore, if the network area is larger than the preset fault range, the network management system will plan the transmission path for multiple transit points based on the source device, destination device, connection relationship information, and available resource information, and construct a temporary network transmission path, as described in steps 4031 to 4035.

[0090] The embodiments of the present invention can accurately determine the starting point and ending point of transmission, and flexibly select a single relay point or multiple relay point path planning method according to the size of the network failure affected area. This allows for the full integration of connection relationship information and available resource information to select stable and resource-sufficient relay points, construct a reliable temporary network transmission path, effectively realize network transmission path replacement under different failure ranges, achieve rapid and reliable recovery of the smart data center network under complex conditions, and ensure the continuous and stable operation of the smart data center network.

[0091] In one embodiment, steps 4021 to 4024 include:

[0092] Step 4021: Based on the connection relationship information, determine the candidate network devices that have direct connections with both the source device and the destination device, and based on the connection relationship information, determine the first historical connection interruption count and the first total connection count between each candidate network device and the source device, as well as the second historical connection interruption count and the second total connection count between each candidate network device and the destination device.

[0093] Optionally, the network management system can filter out intermediate network devices that have direct connections with both the source and destination devices based on the connection relationship information, and determine them as candidate network devices.

[0094] Furthermore, the network management system extracts the number of historical connection interruptions between each candidate network device and the source device from the connection relationship information, denoted as the first historical connection interruption count; it also extracts the total number of connections between each candidate network device and the source device, denoted as the first total connection count. Simultaneously, it extracts the number of historical connection interruptions between each candidate network device and the destination device, denoted as the second historical connection interruption count; and it also extracts the total number of connections between each candidate network device and the destination device, denoted as the second total connection count.

[0095] Continuing with the above embodiment, the source device is server D, and the destination device is aggregation switch B. The network management system filters out intermediate network devices from the connection relationship information, identifying intermediate switch E as the intermediate network device that has a direct connection relationship with both server D and aggregation switch B. Therefore, intermediate switch E is the candidate network device. The network management system extracts the following: the first historical connection interruption count between intermediate switch E and server D is 2, and the first total connection count is 100; the second historical connection interruption count between intermediate switch E and aggregation switch B is 1, and the second total connection count is 120.

[0096] Step 4022: Based on the first historical connection interruption count and the first total connection count of each candidate network device, determine the first connection stability coefficient between each candidate network device and the source device, and based on the second historical connection interruption count and the second total connection count of each candidate network device, determine the second connection stability coefficient between each candidate network device and the destination device.

[0097] Furthermore, for each candidate network device, the network management system calculates the connection stability coefficient with both the source and destination devices using the connection stability coefficient calculation formula. Therefore, the first connection stability coefficient = (first total number of connections - first historical connection interruption count) / first total number of connections, where the first connection stability coefficient reflects the stability of the connection between the candidate network device and the source device. The second connection stability coefficient = (second total number of connections - second historical connection interruption count) / second total number of connections, where the second connection stability coefficient reflects the stability of the connection between the candidate network device and the destination device.

[0098] Continuing with the above embodiment, for the candidate network device intermediate switch E, the network management system calculates the first connection stability coefficient: (100-2) / 100=0.98, that is, the first connection stability coefficient between intermediate switch E and source device server D is 0.98. The second connection stability coefficient is calculated: (120-1) / 120≈0.99, that is, the second connection stability coefficient between intermediate switch E and destination device aggregation switch B is 0.99.

[0099] Step 4023: Based on the available resource information and the first and second connection stability coefficients of each candidate network device, determine the optimal relay network device.

[0100] Furthermore, the network management system determines the optimal relay network device based on the available resource information and the first and second connection stability coefficients of each candidate network device, as described in steps 40231 to 40234.

[0101] Step 4024: Based on the connection relationship between the source device, the optimal relay network device, and the destination device, construct a temporary network transmission path.

[0102] Furthermore, the network management system clarifies the connection relationships between the source device, the optimal relay network device, and the destination device. That is, the source device is directly connected to the optimal relay network device, and the optimal relay network device is directly connected to the destination device. Data transmission links are constructed in the order of "source device → optimal relay network device → destination device" to form a temporary network transmission path. The temporary network transmission path clarifies the complete transmission route of data from the source device, through the optimal relay network device, to the destination device.

[0103] Continuing with the above embodiment, the source device is server D, the optimal relay network device is intermediate switch E, and the destination device is aggregation switch B. It is known that server D is directly connected to intermediate switch E, and intermediate switch E is directly connected to aggregation switch B. Therefore, the temporary network transmission path constructed by the network management system is "server D → intermediate switch E → aggregation switch B".

[0104] The embodiments of the present invention can accurately screen candidate relay devices from both the perspectives of connection relationship and resource status, calculate the connection stability coefficient and determine the optimal relay point in combination with resource occupancy, and finally construct a single relay point network temporary transmission path with stable connection and sufficient resources, ensuring the reliability and effectiveness of the temporary transmission path, realizing the rapid and reliable recovery of the smart data center network under complex conditions, and ensuring the continuous and stable operation of the smart data center network.

[0105] In one embodiment, steps 40231 to 40234 include:

[0106] Step 40231: Based on the storage space utilization rate, computing power utilization rate and port utilization rate of each candidate network device determined by the available resource information, determine the resource load margin of each candidate network device.

[0107] Optionally, the network management system extracts the storage space utilization rate, computing power utilization rate, and port utilization rate of each candidate network device from the available resource information. Based on the storage space utilization rate, computing power utilization rate, and port utilization rate, the resource load margin is determined. The resource load margin is used to reflect the remaining resource carrying capacity of the candidate network device. The calculation method is as follows: storage space load margin = 1 - storage space utilization rate, computing power load margin = 1 - computing power utilization rate, and port load margin = 1 - port utilization rate.

[0108] Continuing with the above embodiment, the candidate network device is intermediate switch E, whose available resource information shows a storage space utilization rate of 30%, a computing power utilization rate of 25%, and a port utilization rate of 40%. The network management system calculates that: storage space load margin = 1 - 30% = 70%, computing power load margin = 1 - 25% = 75%, and port load margin = 1 - 40% = 60%.

[0109] Step 40232: Determine the first connection bandwidth between each candidate network device and the source device, and the second connection bandwidth between each candidate network device and the destination device, based on the available resource information.

[0110] Furthermore, the network management system extracts the real-time connection bandwidth between each candidate network device and the source device from the connection relationship information and determines it as the first connection bandwidth, whereby the first connection bandwidth reflects the data transmission capability between the candidate device and the source device. Simultaneously, it extracts the real-time connection bandwidth between each candidate network device and the destination device and determines it as the second connection bandwidth, whereby the second connection bandwidth reflects the data transmission capability between the candidate device and the destination device.

[0111] Continuing with the above embodiment, the real-time connection bandwidth between the candidate network device intermediate switch E and the source device server D is 0.8Gbps, therefore the first connection bandwidth is 0.8Gbps; the real-time connection bandwidth between intermediate switch E and the destination device aggregation switch B is 1.0Gbps, therefore the second connection bandwidth is 1.0Gbps.

[0112] Step 40233: Determine the adaptation value for each candidate network device based on the first connection stability coefficient, the second connection stability coefficient, the resource load margin, the first connection bandwidth, and the second connection bandwidth.

[0113] Furthermore, the network management system comprehensively considers various indicators of each candidate network device to determine the adaptation value. The adaptation value is a comprehensive indicator reflecting the degree of suitability of the candidate device as a relay point. The calculation logic is as follows: the adaptation value must simultaneously meet the following requirements: first connection stability coefficient ≥ 0.90, second connection stability coefficient ≥ 0.90, storage space load margin ≥ 40%, computing capacity load margin ≥ 50%, port load margin ≥ 30%, first connection bandwidth ≥ 0.5Gbps, and second connection bandwidth ≥ 0.5Gbps. When all indicators meet the requirements, the adaptation value is recorded as "qualified"; if any indicator is not met, the adaptation value is recorded as "unqualified".

[0114] Continuing with the above embodiment, the candidate network device intermediate switch E has a first connection stability coefficient of 0.98 (>0.90) and a second connection stability coefficient of 0.99 (>0.90); storage space load margin is 70% (>40%), computing capacity load margin is 75% (>50%), and port load margin is 60% (>30%); the first connection bandwidth is 0.8Gbps (>0.5Gbps), and the second connection bandwidth is 1.0Gbps (>0.5Gbps). All indicators meet the requirements, therefore the adaptation value of intermediate switch E is "qualified".

[0115] Step 40234: The candidate network devices that are in normal operation and whose adaptation value is greater than the preset adaptation threshold are determined as the optimal relay network devices.

[0116] Furthermore, the network management system checks the operational status of each candidate network device and filters out candidate devices that are in normal operating condition (no hardware failures, no software errors, and not offline). Further, from these normally operating candidate devices, the network management system selects devices with a "qualified" adaptation value (i.e., an adaptation value greater than a preset adaptation threshold, which is "unqualified").

[0117] If multiple devices meet the criteria, the device with the highest first connection stability coefficient and the highest second connection stability coefficient will be selected first; if the first connection stability coefficient and the second connection stability coefficient are the same, the device with the higher resource load margin will be selected first, and the device will be determined as the optimal relay network device.

[0118] Continuing with the above embodiment, the intermediate switch E, the candidate network device, is in normal operation, without any faults or offline status, and its adaptation value is "qualified" (greater than the preset adaptation threshold "unqualified"). Since there is only one candidate network device, intermediate switch E, the network management system determines intermediate switch E as the optimal relay network device.

[0119] This invention comprehensively evaluates candidate relay devices from multiple dimensions, including resource load capacity, connection transmission capacity, and connection stability, and accurately selects devices with high adaptability and stable operation as the optimal relay point. This fully ensures the reliability and carrying capacity of the relay devices, provides key support for building a stable and efficient temporary network transmission path, and ensures smooth data transmission switching in case of failure.

[0120] In one embodiment, steps 4031 to 4035 include:

[0121] Step 4031: Divide the network area into multiple sub-areas and determine the source sub-area and destination sub-area where the source device and destination device are located, respectively.

[0122] Optionally, the network management system divides the network area into multiple sub-regions based on the network topology, device function types, and physical location distribution. The division principle is that devices within each sub-region have similar functional attributes (e.g., access layer devices, aggregation layer devices, core layer devices), and the devices are physically close and closely connected. After division, the network management system identifies the sub-region where the source device is located and designates it as the source sub-region; it identifies the sub-region where the destination device is located and designates it as the destination sub-region; the remaining sub-regions are designated as intermediate sub-regions.

[0123] In one embodiment, the network area encompasses an access network area, an aggregation network area, and a core network area. The network management system divides it into access sub-area A (containing access layer devices such as access switches and servers), aggregation sub-area B (containing aggregation layer devices such as aggregation switches), and core sub-area C (containing core layer devices such as core routers and core switches). The source device, server I, is located in access sub-area A; therefore, the source sub-area is access sub-area A. The destination device, core switch J, is located in core sub-area C; therefore, the destination sub-area is core sub-area C. Aggregation sub-area B is an intermediate sub-area.

[0124] Step 4032: Based on the total connection bandwidth and number of devices between any two intermediate sub-regions, determine the connection tightness index between any two intermediate sub-regions, and based on the location distribution of each intermediate sub-region and the connection tightness index between any two intermediate sub-regions, plan the intermediate rotor region from the source sub-region to the destination sub-region.

[0125] Furthermore, the network management system extracts the connection bandwidth of all devices between any two intermediate sub-regions from the connection relationship information, and sums them to obtain the total connection bandwidth between the two intermediate sub-regions; at the same time, it counts the number of devices in each intermediate sub-region, and calculates the connection tightness index between any two intermediate sub-regions based on the total connection bandwidth and the number of devices between them. The formula for calculating the connection tightness index is: Connection tightness index = Total connection bandwidth between the two intermediate sub-regions / (Sum of the number of devices in the two intermediate sub-regions).

[0126] Furthermore, based on the connection tightness index between any two intermediate sub-regions and the location distribution of each intermediate sub-region (the approximate directional order from the source sub-region to the destination sub-region), the network management system prioritizes intermediate sub-regions with high connection tightness index and continuous location distribution, and plans the intermediate rotor region sequence from the source sub-region to the destination sub-region, ensuring that the intermediate rotor regions are tightly connected and the paths are continuous.

[0127] In one embodiment, the intermediate sub-regions include convergence sub-regions B1, B2, and B3. The calculated connections are as follows: the total bandwidth between convergence sub-regions B1 and B2 is 5Gbps, with a total of 10 devices, resulting in a connection tightness index of 5 / 10 = 0.5; the total bandwidth between convergence sub-regions B2 and B3 is 6Gbps, with a total of 12 devices, resulting in a connection tightness index of 6 / 12 = 0.5; and the total bandwidth between convergence sub-regions B1 and B3 is 3Gbps, with a total of 11 devices, resulting in a connection tightness index of 3 / 11 ≈ 0.27. Considering the location distribution, the directional order from source sub-region A to destination sub-region C is B1→B2→B3. Therefore, the planned intermediate sub-regions are convergence sub-regions B1, B2, and B3.

[0128] Step 4033: For each first network device in each intermediate sub-region, determine candidate network devices based on the first connection bandwidth with the second network device in the preceding sub-region and the second connection bandwidth with the third network device in the following sub-region, which are determined by the connection relationship information.

[0129] Furthermore, for each first network device in each central rotor region, the network management system extracts the real-time connection bandwidth between the device and the second network device in the preceding sub-region (i.e., the sub-region before the current central rotor region, with the source sub-region being the preceding sub-region of the first central rotor region) from the connection relationship information, as the first connection bandwidth; at the same time, it extracts the real-time connection bandwidth between the device and the third network device in the following sub-region (i.e., the sub-region after the current central rotor region, with the destination sub-region being the following sub-region of the last central rotor region), as the second connection bandwidth.

[0130] Furthermore, the network management system selects a first network device with a first connection bandwidth ≥ 0.5Gbps and a second connection bandwidth ≥ 0.5Gbps, and determines it as a candidate network device for the central rotor area.

[0131] Continuing with the above embodiment, the first network devices in the aggregation sub-area B2 of the central rotor region include aggregation switches G1 and G2. Aggregation switch G1 has a first connection bandwidth of 1.2Gbps with the second network device in the preceding sub-area B1 and a second connection bandwidth of 1.5Gbps with the third network device in the following sub-area B3. Aggregation switch G2 has a first connection bandwidth of 0.4Gbps with the second network device in the preceding sub-area B1 and a second connection bandwidth of 1.3Gbps with the third network device in the following sub-area B3. Since both the first and second connection bandwidths of aggregation switch G1 are greater than 0.5Gbps, it is determined to be a candidate network device; aggregation switch G2 is not listed as a candidate network device because its first connection bandwidth does not meet the requirement.

[0132] Step 4034: Based on the storage space utilization rate, computing power utilization rate and port utilization rate of each candidate network device determined by the available resource information, and combined with the first connection bandwidth and the second connection bandwidth, determine the cross-regional transmission capability value of each candidate network device.

[0133] Furthermore, the network management system extracts the storage space utilization rate, computing power utilization rate, and port utilization rate of each candidate network device from the available resource information, and calculates the resource load margin (calculation method is the same as step 40231). The evaluation logic for the cross-regional transmission capacity value is as follows: the storage space load margin must be ≥40%, the computing power load margin must be ≥50%, the port load margin must be ≥30%, the first connection bandwidth must be ≥0.5Gbps, and the second connection bandwidth must be ≥0.5Gbps. When all conditions are met, the cross-regional transmission capacity value is recorded as "good"; if any condition is not met, it is recorded as "insufficient".

[0134] Continuing with the above embodiment, the candidate network device aggregation switch G1 has a storage space utilization rate of 35%, a computing power utilization rate of 40%, and a port utilization rate of 50%. The calculated storage space load margin is 65%, the computing power load margin is 60%, and the port load margin is 50%, all meeting the resource load margin requirements. The first connection bandwidth is 1.2Gbps, and the second connection bandwidth is 1.5Gbps, both greater than 0.5Gbps. Therefore, the cross-regional transmission capability of aggregation switch G1 is "good".

[0135] Step 4035: The candidate network devices that are in normal operation and have a cross-regional transmission capacity value greater than the preset capacity threshold in each intermediate rotor region are determined as the optimal relay network devices in each intermediate rotor region, and a temporary network transmission path is constructed based on the connection relationship between the source device, the destination device and the optimal relay network device in each intermediate rotor region.

[0136] Furthermore, the network management system checks the operational status of candidate network devices in each intermediate rotor area, filters out devices that are in normal operating condition (no hardware failures, no software errors, and not offline), and then selects devices from these devices whose cross-regional transmission capability value is "good" (i.e., greater than the preset capability threshold "insufficient"). If multiple devices meet the criteria, the device with the highest first and second connection bandwidth is selected as the optimal relay network device for the intermediate rotor area. Further, the network management system constructs temporary network transmission paths according to the sequence "source device → first optimal relay device for the intermediate rotor area → second optimal relay device for the intermediate rotor area → ... → destination device," combined with the connection relationships between the devices.

[0137] Continuing from point sixteen, the optimal relay device for sub-region B1 is aggregation switch F, for B2 it is aggregation switch G1, and for B3 it is core router H. The source device, server I, is located in source sub-region A, and the destination device, core switch J, is located in destination sub-region C. The temporary network transmission path constructed by the network management system is "Server I → Aggregation Switch F → Aggregation Switch G1 → Core Router H → Core Switch J".

[0138] This invention, by comprehensively considering regional connectivity, device transmission capacity, and resource load status, constructs a temporary transmission path consisting of multiple stable and reliable relay points. This ensures that even in the event of a large network area failure, data can be transmitted efficiently through the optimal relay link, enabling rapid and reliable recovery of the smart data center network under complex conditions and guaranteeing the continuous and stable operation of the smart data center network.

[0139] Furthermore, the intelligent data center network rapid recovery system provided by the present invention will be described below. The intelligent data center network rapid recovery system described below and the intelligent data center network rapid recovery method described above can be referred to in correspondence.

[0140] Optional, refer to Figure 2 , Figure 2 This is a schematic diagram of the intelligent data center network rapid recovery system provided by the present invention. The intelligent data center network rapid recovery system includes:

[0141] The first acquisition module 210 is used to collect the operating status information of network devices in the smart data center network;

[0142] The fault range prediction module 220 is used to determine the network area affected by the network fault based on the fault location information and fault type information if the network fault is determined to have occurred based on the operating status information.

[0143] The second acquisition module 230 is used to acquire connection relationship information between all network devices within the network area and available resource information of all network devices.

[0144] The transmission path planning module 240 is used to plan transmission paths based on network area range, connection relationship information and available resource information, and to construct temporary network transmission paths.

[0145] The transmission path switching module 250 is used to switch network traffic affected by network failure from the original transmission path to a temporary network transmission path in order to restore the network of the smart data center.

[0146] This invention enables rapid location of the network area affected by a network fault based on the fault location and fault type information. Then, based on the connection relationships between all network devices within the affected area and the available resources of all network devices, a new temporary network transmission path is constructed. Finally, the network traffic affected by the fault is switched through this temporary transmission path, achieving rapid recovery of the smart data center network. This allows for the construction of an effective transmission path based on real-time network conditions when unexpected changes in the network topology render the original backup path unavailable. This avoids prolonged recovery time or recovery failure due to reliance on preset paths, achieving rapid and reliable recovery of the smart data center network under complex conditions and ensuring its continuous and stable operation.

[0147] Please see Figure 3 , Figure 3 An embodiment diagram of an electronic device provided in accordance with the present invention. For example... Figure 3 As shown, this embodiment of the invention provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it performs the following steps:

[0148] Collect operational status information of network devices in the smart data center network;

[0149] If a network failure is determined based on the operational status information, then the network area affected by the network failure is determined based on the failure location information and failure type information.

[0150] Collect information on the connectivity relationships between all network devices within the network area and the available resource information of all network devices;

[0151] Based on network area range, connectivity information, and available resource information, transmission path planning is performed to construct temporary network transmission paths;

[0152] The network traffic affected by the network failure is switched from the original transmission path to a temporary network transmission path in order to restore the network of the smart data center.

[0153] Please see Figure 4 , Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with an embodiment of the present invention is shown. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, it performs the following steps:

[0154] Collect operational status information of network devices in the smart data center network;

[0155] If a network failure is determined based on the operational status information, then the network area affected by the network failure is determined based on the failure location information and failure type information.

[0156] Collect information on the connectivity relationships between all network devices within the network area and the available resource information of all network devices;

[0157] Based on network area range, connectivity information, and available resource information, transmission path planning is performed to construct temporary network transmission paths;

[0158] The network traffic affected by the network failure is switched from the original transmission path to a temporary network transmission path in order to restore the network of the smart data center.

[0159] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the intelligent data center network fast recovery method provided by the above methods, the method including:

[0160] Collect operational status information of network devices in the smart data center network;

[0161] If a network failure is determined based on the operational status information, then the network area affected by the network failure is determined based on the failure location information and failure type information.

[0162] Collect information on the connectivity relationships between all network devices within the network area and the available resource information of all network devices;

[0163] Based on network area range, connectivity information, and available resource information, transmission path planning is performed to construct temporary network transmission paths;

[0164] The network traffic affected by the network failure is switched from the original transmission path to a temporary network transmission path in order to restore the network of the smart data center.

[0165] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0166] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus necessary general-purpose hardware platforms, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for rapid network recovery in a smart data center, characterized in that, include: Collect operational status information of network devices in the smart data center network; If a network failure is determined based on the operational status information, then the network area affected by the network failure is determined based on the failure location information and failure type information. Collect connection relationship information between all network devices within the network area and available resource information of all network devices; Based on the network area range, the connection relationship information, and the available resource information, a transmission path is planned, and a temporary network transmission path is constructed. The network traffic affected by the network failure is switched from the original transmission path to the temporary network transmission path to restore the network of the smart data center; The step of planning transmission paths and constructing temporary network transmission paths based on the network area range, the connection relationship information, and the available resource information includes: The source and destination devices are determined based on the original transmission path; If the network area is less than or equal to the preset fault range, then based on the source device, the destination device, the connection relationship information, and the available resource information, a transmission path is planned for a single transit point to construct the network temporary transmission path; If the network area is larger than the preset fault range, then based on the source device, the destination device, the connection relationship information, and the available resource information, a transmission path with multiple transit points is planned to construct the network temporary transmission path. The step of planning a transmission path with multiple relay points based on the source device, the destination device, the connection relationship information, and the available resource information, and constructing the temporary network transmission path, includes: The network area is divided into multiple sub-regions, and the source sub-region and destination sub-region where the source device and the destination device are located are determined respectively. Based on the total connection bandwidth and number of devices between any two intermediate sub-regions, determine the connection tightness index between any two intermediate sub-regions, and plan the intermediate rotor region from the source sub-region to the destination sub-region based on the location distribution of each intermediate sub-region and the connection tightness index between any two intermediate sub-regions. For each first network device in each mid-rotor region, a candidate network device is determined based on the first connection bandwidth with the second network device in the preceding sub-region and the second connection bandwidth with the third network device in the following sub-region, as determined by the connection relationship information. Based on the available resource information, the storage space utilization rate, computing power utilization rate, and port utilization rate of each candidate network device are determined, and combined with the first connection bandwidth and the second connection bandwidth, the cross-regional transmission capability value of each candidate network device is determined. Candidate network devices that are in normal operation and have a cross-regional transmission capacity value greater than a preset capacity threshold in each intermediate rotor region are identified as the optimal relay network device in each intermediate rotor region. If there are multiple candidate network devices that meet the conditions, the candidate network device with the highest first connection bandwidth and the highest second connection bandwidth is selected as the optimal relay network device in the intermediate rotor region. The temporary network transmission path is constructed based on the connection relationship between the source device, the destination device, and the optimal relay network device in each relay region.

2. The method for rapid network recovery in a smart data center according to claim 1, characterized in that, The step of planning a transmission path for a single relay point based on the source device, the destination device, the connection relationship information, and the available resource information, and constructing the temporary network transmission path, includes: Based on the connection relationship information, candidate network devices that have direct connections with both the source device and the destination device are determined, and based on the connection relationship information, the first historical connection interruption count and the first total connection count of each candidate network device with the source device, as well as the second historical connection interruption count and the second total connection count of each candidate network device with the destination device are determined. Based on the first historical connection interruption count and the first total connection count of each candidate network device, a first connection stability coefficient between each candidate network device and the source device is determined, and based on the second historical connection interruption count and the second total connection count of each candidate network device, a second connection stability coefficient between each candidate network device and the destination device is determined. Based on the available resource information and the first and second connection stability coefficients of each candidate network device, the optimal relay network device is determined. Based on the connection relationship between the source device, the optimal relay network device, and the destination device, the temporary network transmission path is constructed.

3. The method for rapid network recovery in a smart data center according to claim 2, characterized in that, The step of determining the optimal relay network device based on the available resource information and the first and second connection stability coefficients of each candidate network device includes: Based on the available resource information, the storage space utilization rate, computing power utilization rate and port utilization rate of each candidate network device are determined, and the resource load margin of each candidate network device is determined. Based on the available resource information, determine the first connection bandwidth between each candidate network device and the source device, and the second connection bandwidth between each candidate network device and the destination device; Based on the first connection stability coefficient, second connection stability coefficient, resource load margin, first connection bandwidth and second connection bandwidth of each candidate network device, determine the adaptation value of each candidate network device; Candidate network devices that are in normal operation and whose adaptation value is greater than the preset adaptation threshold are identified as the optimal relay network device. If there are multiple candidate network devices that meet the conditions, the candidate network device with the highest first connection stability coefficient and the highest second connection stability coefficient is selected first. If the first connection stability coefficient and the second connection stability coefficient are the same, the candidate network device with the higher resource load margin is selected first, and finally determined as the optimal relay network device.

4. The method for rapid recovery of a smart data center network according to any one of claims 1 to 3, characterized in that, The determination of the network area affected by the network fault based on the fault location information and fault type information includes: The target fault node is obtained by matching the fault type information with the nodes in the pre-built fault structure tree. Based on the initial branch path of the target fault node in the fault structure tree, start the same branch fault traversal. The traversal direction is to extend from the target fault node to the root node and leaf node of the fault structure tree to obtain the same branch traversal path of the target fault node in the fault structure tree. Based on the network units associated with each node in the same branch traversal path and the fault location information, the affected units in the same branch that are directly affected by the fault are obtained. Cross-branch association positioning is performed based on the same branch traversal path and the edges in the fault structure tree to obtain the cross-branch influence area affected by the fault under the cross branch. The network area affected by the network fault is obtained by fusing the same branch influence unit and the cross branch influence region. The fault structure tree is constructed based on fault types and their corresponding hierarchical relationships and network units and their corresponding association relationships in the knowledge of network fault domains. The nodes in the fault structure tree are fault types and network units, and the edges represent the causal relationships between fault types or the subordinate relationships between faults and network units.

5. The method for rapid network recovery in a smart data center according to claim 4, characterized in that, The method of cross-branch association positioning based on the same branch traversal path and the edges in the fault structure tree to obtain the cross-branch affected area under the fault includes: Traverse the fault tree structure based on the same branch traversal path and the edges in the fault tree structure to determine the cross-related nodes in the fault tree structure that have cross-related relationships with the same branch nodes. Based on each cross-related node as the starting node of the cross-branch, the fault structure tree is traversed to determine the cross-branch path corresponding to each cross-branch starting node, and effective paths with actual fault propagation potential are selected based on each cross-branch path to obtain the cross-branch propagation path. Based on the network regions associated with each node in the cross-branch diffusion path and the fault location information, the cross-branch affected area under the cross-branch is determined.

6. A smart data center network rapid recovery system, characterized in that, The system is applied to the intelligent data center network rapid recovery method as described in any one of claims 1 to 5; the system includes: The first data acquisition module is used to collect the operating status information of network devices in the smart data center network; The fault range prediction module is used to determine the network area affected by the network fault based on the fault location information and fault type information if the network fault is determined to have occurred based on the operating status information. The second acquisition module is used to acquire connection relationship information between all network devices within the network area and available resource information of all network devices. The transmission path planning module is used to plan transmission paths based on the network area range, the connection relationship information, and the available resource information, and to construct temporary network transmission paths. The transmission path switching module is used to switch network traffic affected by network failure from the original transmission path to the temporary network transmission path in order to restore the network of the smart data center. The step of planning transmission paths and constructing temporary network transmission paths based on the network area range, the connection relationship information, and the available resource information includes: The source and destination devices are determined based on the original transmission path; If the network area is less than or equal to the preset fault range, then based on the source device, the destination device, the connection relationship information, and the available resource information, a transmission path is planned for a single transit point to construct the network temporary transmission path; If the network area is larger than the preset fault range, then based on the source device, the destination device, the connection relationship information, and the available resource information, a transmission path with multiple transit points is planned to construct the network temporary transmission path. The step of planning a transmission path with multiple relay points based on the source device, the destination device, the connection relationship information, and the available resource information, and constructing the temporary network transmission path, includes: The network area is divided into multiple sub-regions, and the source sub-region and destination sub-region where the source device and the destination device are located are determined respectively. Based on the total connection bandwidth and number of devices between any two intermediate sub-regions, determine the connection tightness index between any two intermediate sub-regions, and plan the intermediate rotor region from the source sub-region to the destination sub-region based on the location distribution of each intermediate sub-region and the connection tightness index between any two intermediate sub-regions. For each first network device in each mid-rotor region, a candidate network device is determined based on the first connection bandwidth with the second network device in the preceding sub-region and the second connection bandwidth with the third network device in the following sub-region, as determined by the connection relationship information. Based on the available resource information, the storage space utilization rate, computing power utilization rate, and port utilization rate of each candidate network device are determined, and combined with the first connection bandwidth and the second connection bandwidth, the cross-regional transmission capability value of each candidate network device is determined. Candidate network devices that are in normal operation and have a cross-regional transmission capacity value greater than a preset capacity threshold in each intermediate rotor region are identified as the optimal relay network device in each intermediate rotor region. If there are multiple candidate network devices that meet the conditions, the candidate network device with the highest first connection bandwidth and the highest second connection bandwidth is selected as the optimal relay network device in the intermediate rotor region. The temporary network transmission path is constructed based on the connection relationship between the source device, the destination device, and the optimal relay network device in each relay region.

7. An electronic device, comprising: Memory, used to store computer software programs; A processor for reading and executing the computer software program, characterized in that, when the processor executes the computer software program, it implements the smart data center network rapid recovery method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium, wherein a computer software program is stored therein, characterized in that, When the computer software program is executed by the processor, it implements the smart data center network fast recovery method as described in any one of claims 1 to 5.