Network topology change prediction method and device combined with big data analysis
By constructing a coupled network topology and combining it with big data analysis, the physical and logical layers are mapped in real time, which solves the problem of insufficient coupling between the physical and logical layers in network topology change prediction. This enables accurate identification of fault propagation paths and impact ranges, and improves the fault tolerance and service reliability of the network system.
Patent Information
- Application Number
- CN202511470304.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing technologies fail to adequately consider the dynamic coupling between the physical and logical layers in predicting network topology changes, resulting in an inability to accurately assess the propagation and impact of faults between the two layers, thus affecting the accuracy and timeliness of network fault prediction.
By constructing a coupled network topology, the physical and logical topology layers are dynamically mapped in real time. Edge fault prediction nodes are deployed at the physical layer. Topology cascade propagation analysis and accompanying entity mapping are used to identify fault propagation paths and their impact range. Combined with big data analysis, business service dependency chain shutdown mapping is performed to output fault risks.
It improves the real-time performance and accuracy of network fault prediction, enabling early warnings before faults occur, preventing network outages and equipment downtime, and enhancing the reliability and stability of network services.
Smart Images

Figure CN120935049B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network topology prediction, in particular to a network topology change prediction method and device combined with big data analysis. BACKGROUND
[0002] With the increasing complexity of modern network systems, especially in the case of the coexistence of physical devices and virtualized network resources, how to efficiently predict network topology changes has become an important research topic. Existing network topology change prediction is mainly based on the analysis of static topology structure of the physical layer network or the virtualization layer. However, the dynamic mapping and interaction relationship between the physical layer and the virtual layer is not fully considered, which leads to the inability to accurately predict the propagation path and impact range of faults in the virtualization environment, which makes it impossible to accurately assess the propagation and impact of faults between the two, thereby affecting the prediction accuracy and timeliness of network faults. SUMMARY
[0003] The present application provides a network topology change prediction method and device combined with big data analysis, aiming to solve the technical problem that the existing network topology change prediction is mainly based on static topology structure, ignoring the dynamic coupling of physical topology layer and logical topology layer, resulting in the inability to accurately assess the propagation and impact of faults between the two, thereby affecting the prediction accuracy of network faults.
[0004] The first aspect of the present application provides a network topology change prediction method combined with big data analysis, the method comprising: pre-building a coupled network topology, wherein the coupled network topology comprises a physical topology layer and a logical topology layer connected in real-time dynamic mapping; deploying a plurality of edge fault prediction nodes locally in a plurality of physical devices in the physical topology layer; after a first edge fault identification node analyzes and outputs a first hop prediction fault according to a first real-time operating condition, performing topology cascade propagation analysis and positioning X associated fault devices according to the out-degree connection relationship of the physical topology layer; the first edge fault identification node sends the first hop prediction fault to X edge fault identification nodes of the X associated fault devices to trigger a cooperative prediction mechanism, performs fault front prediction update, and outputs X front prediction faults; performing concomitant entity mapping according to the first hop prediction fault and the X front prediction faults, and outputting a global logical impact domain; projecting the global logical impact domain to the logical topology layer in the coupled network topology to locate network topology change distribution; and performing business service dependency chain stall mapping according to the network topology change distribution, and outputting business service interruption risk.
[0005] In a second aspect, the application discloses a network topology change prediction device combined with big data analysis, which is used for the network topology change prediction method combined with big data analysis, and comprises a network topology construction module, a prediction node deployment module, a fault device positioning module, a pre-prediction update module, a concomitant entity mapping module, a change distribution positioning module and a risk output module.
[0006] The one or more technical solutions provided in the application have at least the following beneficial effects:
[0007] By constructing the physical topology layer and the logical topology layer, and performing dynamic mapping, the relationship between the physical network and the virtual network can be comprehensively displayed, which provides a global perspective for network management and ensures that the state changes between physical devices and virtual resources can be synchronized in time. By deploying multiple edge fault prediction nodes in the physical topology layer, data analysis can be performed near the devices, reducing delay and bandwidth consumption, which helps to improve the real-time performance of fault prediction, so that device failures can be identified and responded to in the first time. The collaborative prediction mechanism between multiple edge nodes can optimize the accuracy of fault pre-position prediction through the collective prediction results. Through topology cascade propagation analysis, the fault propagation path and its impact range can be accurately identified according to the out-degree connection relationship of the physical topology layer, which means that not only local faults can be detected, but also which other devices will be affected by the fault can be predicted, so that associated devices can be located and risk assessment can be performed in advance. Through the accompanying entity mapping, the impact domain of the logical layer, i.e. the impact range of the fault on the virtualized resources, can be quickly located based on the physical layer fault, so that more comprehensive fault management can be performed. By analyzing the business service dependency chain shutdown mapping, it can be identified in advance which services will be affected and may cause shutdown when the network fails, and the potential business impact of fault propagation can be accurately evaluated using network topology change distribution. By closely combining the physical and logical layers, combined with big data analysis and prediction algorithms, the entire method significantly improves the fault tolerance of the network system when facing faults. It can provide early warning before the fault occurs to prevent network interruption and device downtime, and improve the reliability and stability of network services.
[0008] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the following specific embodiments of the present application can be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 The network topology change prediction method combined with big data analysis provided by the embodiments of the present application is shown in the flowchart.
[0010] Figure 2 The network topology change prediction device structure combined with big data analysis provided by the embodiments of the present application is shown in the schematic diagram.
[0011] Explanation of reference numerals: network topology construction module 10, prediction node deployment module 20, fault device positioning module 30, pre-position prediction update module 40, accompanying entity mapping module 50, change distribution positioning module 60, risk output module 70. DETAILED DESCRIPTION
[0012] The embodiment of the application provides a network topology change prediction method and device combined with big data analysis, and solves the technical problem that the network topology change prediction in the prior art is mainly based on a static topology structure, the dynamic coupling of a physical topology layer and a logical topology layer is ignored, the propagation and influence of a fault between the two layers cannot be accurately evaluated, and therefore the prediction accuracy of a network fault is affected.
[0013] After introducing the basic principle of the application, various non-limiting embodiments of the application will be specifically introduced in combination with the accompanying drawings of the specification. It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.
[0014] Embodiment one, as shown in the figure, the embodiment of the application provides a network topology change prediction method combined with big data analysis, which comprises the following steps: Figure 1
[0015] Pre-constructing a coupled network topology, wherein the coupled network topology comprises a physical topology layer and a logical topology layer which are dynamically mapped in real time.
[0016] The physical topology layer contains the connection relationship of all physical devices, and the devices form a network structure through hardware connection in this layer, and can be multiple physical devices such as sensors, computing nodes, switches and the like, each device has an out-degree of connection, that is, the number of devices connected outward; the logical topology layer describes the connection relationship between virtual network entities (such as virtual machines, virtual network connections and the like), and the logical topology layer usually depends on virtualization technology, and represents the mutual relationship of different network entities in the virtualization environment. The mapping between the physical topology and the logical topology is dynamic, that is, the relationship between them is not static, but is dynamically updated with the running state of the network, the configuration change, the device fault and the like, for example, when the physical device changes or fails, the mapping relationship of the logical topology layer will also be adjusted accordingly.
[0017] A plurality of edge fault prediction nodes are locally deployed in the plurality of physical devices in the physical topology layer.
[0018] The edge fault prediction node is responsible for obtaining real-time running data such as temperature, vibration, load, current and the like from the local device, and performing fault prediction based on the data, in actual deployment, the edge fault prediction node of each physical device is connected to other nodes through a network, forming a distributed system, and the state of the physical device is monitored in real time by the edge fault prediction node.
[0019] When the first edge fault identification node analyzes and outputs the first-hop predicted fault according to the first real-time running condition, the topological cascade propagation analysis is performed according to the out-degree connection relationship of the physical topology layer to locate X associated fault devices.
[0020] The first edge fault identification node analyzes whether there is a fault, such as temperature anomaly or current exceeding the standard, according to the first real-time running condition obtained by the first edge fault identification node, and outputs a first-hop predicted fault, which means that a certain physical device in the network is experiencing or about to experience a fault. After predicting the fault, the fault is propagated in the physical topology layer through the connection relationship. The out-degree connection relationship represents the network topology structure between a device and other devices connected to it. Based on this relationship, the first node finds the devices affected by the fault through topology cascade propagation analysis. This analysis is performed through graph propagation algorithms, cascade propagation algorithms, etc. According to the propagation result, X associated fault devices are located, which have direct or indirect connection relationship with the original fault node in the topology structure.
[0021] The first edge fault identification node sends the first-hop predicted fault to the X edge fault identification nodes of the X associated fault devices to trigger a cooperative prediction mechanism, performs fault front prediction update, and outputs X front prediction faults.
[0022] At the same time of sending the first-hop predicted fault, the cooperative prediction mechanism is activated. The function of the cooperative prediction mechanism is to synchronize fault prediction on multiple devices in a cooperative manner to improve the overall prediction accuracy and reliability. Based on the received first-hop predicted fault and the current device running condition data, the edge fault identification nodes of the X associated fault devices update the local fault front prediction, which means that the device not only predicts the impending fault, but also identifies potential fault risks in advance. In this process, each associated device outputs a front prediction fault, and finally, X front prediction faults are obtained. These front prediction faults provide more accurate prediction information for subsequent fault handling and risk assessment.
[0023] According to the first-hop predicted fault and the X front prediction faults, a companion entity mapping is performed to output a global logical impact domain.
[0024] The companion entity mapping is a process of associating the predicted fault with other entities in the network. Specifically, the first-hop predicted fault and the X front prediction faults are mapped to logical entities in the network. These logical entities are virtual components, services, resources or other forms of network elements in the network. Between the physical network and the logical network, the fault not only affects the physical device itself, but also affects the service quality and structure of the entire network. Therefore, the companion entity mapping is to propagate the fault to the logical layer and evaluate its impact on the entire network architecture. The global logical impact domain represents the propagation range and impact of the fault in the logical topology layer. According to the result of the companion entity mapping, the logical entities that may be affected by the fault are located. The connection relationship between these logical entities and the fault propagation path help to predict and identify the potential impact area in the entire network.
[0025] The coupling network topology projects the global logical influence domain to the logical topology layer, and locates the network topology change distribution.
[0026] The global logical influence domain is projected to the logical topology layer, which mainly involves the relationship between virtual network entities and services, and the mapping process transmits the information of physical fault influence to the virtual layer. In the projection process, according to the dynamic mapping rules between the physical topology layer and the logical topology layer, the influence area and the fault propagation path are converted into the influence intensity, resource damage and other information of the logical layer, and finally the network topology change distribution is located, that is, which part of the logical topology is affected or changed. The network topology change usually includes the virtualization state change of the device, the failure of the logical node or the change of the resource scheduling, etc.
[0027] According to the network topology change distribution, the service service dependency chain stall mapping is performed, and the service service interruption risk is output.
[0028] The service service dependency chain refers to the dependency relationship between different services, applications or resources in the system. In a complex network environment, there is usually a hierarchical dependency structure between services, for example, the normal operation of an application may depend on multiple underlying services or virtual resources. Stall mapping refers to identifying changes in the network topology to find out which services or applications will be affected, even stalled. The service interruption risk reflects the severity of service interruption caused by topology changes. The risk is usually determined by multiple factors, including the number of affected services, the importance of services, the dependency relationship between services and the possible impact on business processes. According to the mapping result of the dependency chain stall, the interruption risk of each affected service is evaluated, and a risk value is output, for example, the stall of a key service will cause serious service interruption, while some less important services may only cause slight impact. In this way, it can provide key decision basis for managers and help to evaluate the risk and potential service interruption events that need to be handled first.
[0029] Further, when the first edge fault identification node analyzes and outputs the first hop prediction fault according to the first real-time running condition, the topology cascade propagation analysis is performed according to the out-degree connection relationship of the physical topology layer to locate X associated fault devices, and the method comprises:
[0030] With the fault intervention response time as a constraint, the networking calls the first physical device in multiple sample device failure scenarios multiple sample failure precursor feature datasets; after constructing multiple heterogeneous failure recognition models based on the multiple sample failure precursor feature dataset, the first failure recognition unit is constructed by parallel connection of the multiple heterogeneous failure recognition models; with the fault intervention response time as the segmentation window, 1 / 8 of the fault intervention response time as the overlap scale, the first real-time operating condition is slid and segmented to obtain the first operating condition time sequence slice; the first operating condition time sequence slice is sequentially input into the first failure recognition unit for failure precursor perception until the first hop predicted failure is output; with the first physical device as the cascade propagation source starting point, topological cascade propagation analysis is performed according to the out-degree connection relationship of the physical topology layer to locate the X associated fault devices.
[0031] The fault intervention response time refers to the time interval from the detection of a failure signal by a device to the adoption of corrective measures. In the process of failure prediction and intervention, the shorter the response time, the more timely the measures can be taken, reducing the risk of device downtime or service interruption. Based on this time delay, relevant sample failure precursor feature datasets are extracted from multiple sample device failure scenarios. These feature datasets include various operating state data of the device before failure, such as temperature, vibration, load, current, etc.
[0032] The multiple heterogeneous failure recognition models are constructed using machine learning algorithms such as support vector machines, decision trees, neural networks, etc., and are trained based on multiple sample failure precursor feature datasets. They can more accurately identify various types of failure precursors. The multiple heterogeneous failure recognition models are connected in parallel to form an integrated first failure recognition unit. The parallel structure means that these models work simultaneously, and each model independently makes predictions based on the input data. Through this parallel method, the advantages of each model can be combined to compensate for the shortcomings of a single model, improving the accuracy and robustness of failure prediction. The output results of each model can be fused through weighted averaging, voting mechanism or other methods to obtain the final prediction result.
[0033] The fault intervention response time is used as the segmentation window, i.e., the time series data of the real-time operating condition is divided into multiple time periods, each time period is called a time sequence slice, which is used to analyze the state changes of the device in different time periods. In order to improve the accuracy of prediction and the continuity of time series data, the overlap scale is set to 1 / 8 of the fault intervention response time, which means that there is a certain overlap between each slice and the previous slice, ensuring that no key information is missed, while avoiding information disruption between slices. According to the fault intervention response time and the overlap scale, the real-time operating condition data is sliced to generate multiple operating condition time sequence slices, each slice containing the device operating data within a period of time.
[0034] The first operation condition time sequence slice is sequentially input into a first fault identification unit. The fault precursor awareness refers to identifying abnormal or irregular patterns in the operation state of the device before the device has a complete failure. These patterns indicate the occurrence of device failure. The model captures possible trends or abnormal points by analyzing these slices, thereby identifying the failure in advance. When the processing is completed, the first fault identification unit outputs a first-hop predicted failure. The first-hop predicted failure refers to the "first-hop" failure of the device in the fault propagation network, meaning that it is the first device in the fault chain, and subsequent devices may be affected.
[0035] The first physical device is taken as the starting point of the cascading propagation source. This device is the first to detect the failure, so its failure will become the basis for subsequent analysis and prediction. Based on the out-degree connection relationship of the devices in the physical topology layer, that is, which other devices it connects to from this device, the analysis of how to propagate the failure through these connections is started. Topology cascading propagation analysis refers to simulating the propagation process of the failure through the connection relationship of the devices in the physical topology layer. By calculating these connection relationships, the path of the failure propagation from the source device to the downstream device can be predicted. This propagation analysis usually uses graph theory algorithms or propagation models to calculate which devices will be affected according to the dependency relationship and connection strength between devices, and further locate the specific failure propagation path. Based on the failure propagation analysis, X associated failure devices related to the first-hop failure device are identified, X being a positive integer. These devices have direct or indirect connection relationships with the source device in the topology structure, so they may also be affected by the source device failure or be the result of the source device failure.
[0036] Further, the topology cascading propagation analysis is performed based on the out-degree connection relationship of the physical topology layer with the first physical device as the starting point of the cascading propagation source, and the X associated failure devices are located. The method comprises:
[0037] The multiple historical failure logs of the multiple physical devices are obtained interactively. The fault linkage evaluation is performed based on the multiple historical failure logs, and the device failure linkage matrix is constructed. Y candidate failure devices are output according to the row vector correlation degree analysis of the first physical device in the device failure linkage matrix. After the topology cascading propagation analysis is performed based on the out-degree connection relationship of the physical topology layer with the first physical device as the starting point of the cascading propagation source, C initial failure devices are located, and the C initial failure devices are verified for failure conduction path by using the Y candidate failure devices, and the X associated failure devices are output.
[0038] Obtain, from a plurality of physical devices via a network interface or a fault management system, historical fault logs related to faults, the historical fault logs recording past fault events of the devices, and the historical fault logs usually including key information such as a fault time of the devices, a fault mode, a repair history, a fault frequency, and the like.
[0039] The fault linkage evaluation refers to analyzing the fault correlation between a plurality of physical devices, and there can be direct or indirect dependency relationships between the devices, for example, a fault of one device can affect other connected devices, or a plurality of devices can have linkage faults due to sharing the same resources. In the evaluation process, the timing correlation of each device when a fault occurs is analyzed, that is, whether the device faults have consistency or sequence in time, for example, a fault of one device can be a precursor of a fault of another device, or two devices can be affected by a common fault source.
[0040] According to the fault linkage evaluation, a device fault linkage matrix is constructed, the rows and columns of the matrix respectively represent different physical devices, and each element in the matrix represents the fault correlation degree between two devices, and the correlation degree is usually calculated based on timing data of fault occurrence, dependency relationships between devices, and historical fault logs. The content of the device fault linkage matrix can reflect the fault propagation path between devices.
[0041] The row vector correlation degree refers to the fault correlation degree between a device (that is, a first physical device) and other devices in the device fault linkage matrix. The row vector of each device records the fault correlation information of all other devices related to the device. By analyzing the correlation degrees, the correlation between the row vector of the first physical device in the device fault linkage matrix and other devices is found, for example, if the fault of the first physical device frequently causes faults of other devices, then the correlation degree of the row vector of the first physical device with these devices is high. According to the analysis result, Y candidate fault devices most relevant to the first physical device are identified, Y being a positive integer. These devices are sorted according to their fault linkage and correlation degree with the first device, and are usually devices with high correlation degrees. The Y candidate fault devices are selected to ensure that these devices will become the focus of subsequent analysis when a fault occurs.
[0042] The first physical device is taken as a source device of fault propagation, and a topology cascade propagation analysis is started. The fault of the device is the starting point of the entire propagation chain, and the fault will be propagated from the device to the devices directly connected to it through the out-degree connection relationship in the physical topology layer, i.e., the connection mode between devices. The topology cascade propagation analysis is first located to C initial fault devices, C being a positive integer. These devices are devices directly connected to the first physical device through a propagation path, and they are "initial fault" nodes in the fault propagation chain and are the first wave of influence objects of fault propagation. After finding the C initial fault devices, the Y alternative fault devices are used to verify the fault conduction path of the initial fault devices. The goal of this step is to confirm whether the fault will continue to affect the Y alternative fault devices after the fault is propagated from the first physical device to the C initial fault devices. After verification, X associated fault devices are output. These devices are key devices in the entire fault propagation path, and they may be affected by the fault of the first physical device.
[0043] Further, the first-hop predicted fault is accompanied by a fault device identifier, a predicted fault type code, and a fault time window marker.
[0044] The fault device identifier is a unique identifier for marking a device that has failed. The identifier can be the ID, serial number, IP address, etc. of the device, which can accurately locate the device. The predicted fault type code specifies the specific type of fault, for example, the fault type can be electrical failure, mechanical failure, temperature anomaly, etc. Each fault type requires a different processing strategy. The fault time window marker adds a time range to each predicted fault, indicating the time window in which the fault occurs. This helps to predict the occurrence period of the fault and whether preventive intervention is needed within this period.
[0045] Further, the first-edge fault identification node sends the first-hop predicted fault to the X edge fault identification nodes of the X associated fault devices to trigger a cooperative prediction mechanism, performs a fault front prediction update, and outputs X front prediction faults. The method comprises:
[0046] The X edge fault identification nodes load X associated operating conditions according to the fault time window marker. After the X edge fault identification nodes activate the local fault identification model of the X fault identification units according to the predicted fault type code, the X edge fault identification nodes input the X associated operating conditions to perform local fault front prediction, and output the X front prediction faults.
[0047] X edge fault identification nodes load X associated operating conditions from historical operation data according to the time window mark of each fault. These operating condition data refer to equipment operation data related to fault occurrence, such as temperature, vibration, current, etc. For each associated fault, the corresponding operating condition data is aligned according to the fault time window mark, that is, the operating condition in the corresponding period is extracted from the historical data according to the predicted period of the fault, and the data is synchronized.
[0048] Each edge fault identification node is equipped with a fault identification unit. According to different fault types, the local fault identification model of the fault identification unit is activated, and the input operating condition data is analyzed. The local fault identification model is a machine learning-based model that uses the loaded operating condition data to identify whether the equipment has a precursor of a fault. After receiving the associated operating condition, local fault precursor prediction is performed, that is, whether the equipment is about to fail is identified. Each edge fault identification node outputs a precursor prediction fault. These prediction results include the type of equipment fault, the possible time of occurrence, and the possible impact range. Finally, X precursor prediction faults are obtained. These predictions provide key early warning information for subsequent fault response and repair.
[0049] Further, the X edge fault identification nodes execute local fault precursor prediction by inputting the X associated operating conditions after activating the local fault identification model of the X fault identification units according to the predicted fault type code, and output the X precursor prediction faults. The method comprises:
[0050] Interactively obtain a plurality of reference fault transmission paths, wherein each path node in the reference fault transmission path is attached with a sample equipment identifier and a sample fault type code; call a first associated equipment identifier of a first associated fault equipment; predefine a transmission dependency relationship between the first associated equipment identifier and a fault equipment identifier; take the first associated equipment identifier, the fault equipment identifier, and the predicted fault type code as a transmission path query key, take the transmission dependency relationship as a path matching criterion, traverse the plurality of reference fault transmission paths, and compare to output N sample fault types; a first edge fault identification node executes local fault precursor prediction by inputting a first associated operating condition after activating a local fault identification model of a first fault identification unit according to the N sample fault types, and outputs a first precursor prediction fault.
[0051] The benchmark fault conduction path refers to a typical path or fault link of device fault propagation in history, which describes how the fault is propagated through other devices in the network from the occurrence of a device fault, and each path node in the benchmark fault conduction path is attached with a sample device identifier and a sample fault type code, wherein the sample device identifier is a unique identifier of each device for distinguishing different devices; and the sample fault type code represents the fault type experienced by the device.
[0052] The first associated fault device refers to any device in the fault propagation chain that has a direct or indirect connection with the initial fault device, and the first associated device identifier corresponding to the first associated fault device is invoked.
[0053] The conduction dependency describes how devices depend on each other and affect each other, and the fault propagation of a device is often carried out through the connection relationship between devices, and the conduction dependency determines whether the fault of a certain device will affect other devices. Here, the conduction dependency between the first associated device identifier and the fault device identifier refers to how the fault of the first associated device will be conducted and affect the fault device, and this relationship is usually defined based on the physical connection of the device, the network structure or other hierarchical dependencies.
[0054] The first associated device identifier, the fault device identifier and the predicted fault type code are used together as a query key to find the relevant fault propagation path from the benchmark fault conduction path, and these identifiers and codes help to accurately locate the fault-related devices and fault types, so as to query the relevant fault conduction path. The conduction dependency refers to the fault dependency between devices, which determines whether the fault of a device will propagate to other devices, and according to this dependency as the path matching criterion. After obtaining the query key and the path matching criterion, a plurality of benchmark fault conduction paths are traversed, and a plurality of path nodes are compared in the traversal process to find N sample fault types, which describe different modes of device fault in historical conduction paths, including different fault types such as electrical fault, sensor fault, mechanical fault, etc.
[0055] The first edge fault identification node selects an appropriate local fault identification model to activate according to the N sample fault types, and different analysis models are required for each fault type, so this step ensures that the correct model is selected for analysis according to the fault type. The first associated operating condition is input into the local fault identification model, and the local fault identification model analyzes the input first associated operating condition, compares it with historical fault patterns, and performs local pre-fault prediction. This process involves judging whether there is a potential fault risk of the device, identifying the precursor of the device fault, and finally outputting the first pre-fault prediction, which includes the predicted type of the fault, the time window of the occurrence, etc.
[0056] Further, the pre-constructed coupling network topology, the method comprises:
[0057] According to the physical connection relationship of the hardware device, a physical topology layer is constructed; according to the logical interconnection topology relationship of the virtualized network entity, a logical topology layer is constructed; based on a real-time configuration perception engine, the physical topology layer and the logical topology layer are dynamically mapped to construct an output coupling network topology.
[0058] The physical connection relationship of the hardware device refers to how physical devices are connected to each other through physical media (such as cables, switches, routers, etc.), for example, in a data center, servers, switches and storage devices are connected together through a network, forming a physical topology, the core of the physical topology layer is to describe the actual connection mode between devices, and how these connections affect data flow, fault propagation and the dependency relationship between devices. Based on the physical connection relationship between devices, a physical topology layer is constructed to represent how devices are interconnected through physical connections.
[0059] Virtualized network entities refer to virtual devices, virtual switches, virtual machines, virtual networks and other virtual devices created in a virtualized environment, which do not depend on the specific hardware implementation of physical devices, but are organized and managed at the virtualization level. Based on the logical interconnection topology relationship of the virtualized network entity, a logical topology layer is constructed, which represents the logical connection mode between virtual devices, including the communication path between virtual machines, the configuration of virtual switches, etc.
[0060] The real-time configuration perception engine is a dynamic perception and analysis tool that can obtain and update the configuration state of physical and virtual resources in the network in real time. This engine can usually handle changes in devices in the network, such as device failures, configuration changes, resource adjustments, etc., and dynamically adjust the topology mapping according to these changes. Between the physical and logical topology layers, dynamic mapping is performed to reflect the association between the two, dynamic mapping refers to automatically adjusting the mapping relationship between the physical topology and the logical topology according to the real-time state and configuration updates of the devices. For example, when a physical device fails or a virtual machine migrates, the mapping relationship between the topology structure of the virtual network entity and the physical network is adjusted according to the feedback of the real-time configuration perception engine. By dynamically mapping the physical topology layer and the logical topology layer, a coupling network topology is finally constructed, which not only includes the connection relationship of the physical devices, but also includes the association of the virtualized network entities, and can describe the overall structure of the entire network.
[0061] Further, according to the first hop prediction fault and X front prediction faults, the accompanying entity mapping is performed to output a global logical influence domain, the method comprises:
[0062] Based on the interlayer mapping rule, the first-hop accompanying logical entity corresponding to the first-hop predicted fault and the X accompanying logical entities corresponding to the X preceding predicted faults are retrieved respectively; and the first-hop accompanying logical entity and the X accompanying logical entities are aggregated by graph theory to output the global logical impact domain.
[0063] The interlayer mapping rule is a rule describing how the physical layer and the logical layer are mapped. Through these rules, the fault and impact information of the physical layer can be mapped to the logical topology layer, ensuring that the impact of the fault on virtual network entities can also be captured. The first-hop predicted fault is the first node in the fault propagation chain, and its fault will have an impact on virtual entities in the logical layer. In this step, the first-hop accompanying logical entity related to the first-hop predicted fault is retrieved according to the interlayer mapping rule. The first-hop accompanying logical entity is usually a virtual resource or service directly or indirectly related to the faulty device. Similarly, the X accompanying logical entities related to the X preceding predicted faults are retrieved according to the interlayer mapping rule. These faults are further predicted device faults after the first-hop fault, and they will also have an impact on virtual resources.
[0064] Graph theory is a mathematical model for analyzing the relationship between nodes and edges. In this step, graph theory is used to aggregate the first-hop accompanying logical entity and the X accompanying logical entities to form a network structure. In the graph theory model, each logical entity is a node, and the connection relationship between nodes (determined by fault propagation or device dependency) is an edge in the graph. By aggregating these nodes and edges, the impact range of fault propagation and the dependency between virtual resources can be identified. The goal of the aggregation process is to combine multiple accompanying logical entities and their mutual relationships to build a comprehensive impact model. Specifically, the first-hop accompanying logical entity and the X accompanying logical entities are combined according to their fault propagation relationship to obtain a complete fault impact network, i.e., the global logical impact domain. The global logical impact domain represents the set of all virtual resources or services affected during the fault propagation process.
[0065] Further, the method includes the following steps:
[0066] Based on the physical-logical layer dynamic mapping rule, the global logical impact domain is converted into a logical entity impact strength vector. After projecting the logical entity impact strength vector to the logical topology layer, the impact propagation fitting is performed according to the out-degree connection relationship of the logical topology layer to output a logical node vulnerability value matrix. Based on a preset vulnerability value threshold, the logical node vulnerability value matrix is traversed to identify multiple high-risk logical connection paths. The multiple high-risk logical connection paths are aggregated to output the network topology change distribution.
[0067] The physical-logical layer dynamic mapping rule is a rule describing how the physical layer and the logical layer correspond and convert. Through these rules, the impact of physical device failure can be mapped to the logical topology layer, ensuring that the virtual devices in the logical layer can reflect the impact of the physical layer failure. The global logical impact domain represents the impact range of device failure on the entire logical topology layer. In this step, the impact domain is converted into a logical entity impact intensity vector, i.e., the impact degree of each logical entity in the failure propagation. The impact intensity vector represents the impact degree of the failure propagation on each logical entity, which can be in numerical form, representing the intensity of the impact.
[0068] First, the logical entity impact intensity vector is projected onto the logical topology layer, which means that according to the impact intensity vector, the impact value is distributed in the logical topology structure. The logical topology layer represents the dependency relationship of virtualized resources, services, and networks, and through projection, the impact intensity value will be propagated to these virtual resources. The result of the projection is to assign an impact intensity to each logical entity, representing the affected degree of the virtual device or service in the failure propagation. Based on the out-degree connection relationship of the logical topology layer, the impact is propagated and fitted. The out-degree connection relationship refers to the relationship from a certain logical node to other nodes. Each logical node is connected to multiple other nodes, and these connections determine how the failure affects other nodes. Through the fitting result, a logical node vulnerability value matrix is generated, which represents the vulnerability value of each logical node (virtual device or service) in the network, i.e., their degree of being affected by the failure. The higher the vulnerability value, the more likely the node is to be affected by the failure.
[0069] The preset vulnerability threshold is a set standard that represents a logical node's vulnerability reaching a certain level, which is considered a high-risk node. This threshold is set according to the system's fault tolerance capability, business requirements, or failure handling capability, etc. Traverse the entire logical node vulnerability value matrix and check the vulnerability value of each node. If the vulnerability value of a node exceeds the preset threshold, mark the node as a high-risk node. High-risk logical connection paths refer to those paths connecting high-vulnerability nodes, which are the most likely paths to cause serious impact in the failure propagation process.
[0070] All identified high-risk logical connection paths are aggregated. The aggregation of these paths helps identify which parts of the network are vulnerable to failure propagation and how these impacts will propagate to other parts of the network. Ultimately, the network topology change distribution is output, which represents the propagation pattern and impact range of the failure in the network. This distribution provides specific directions for subsequent failure management, repair, and optimization.
[0071] Embodiment two, based on the same inventive concept as the network topology change prediction method combined with big data analysis in the preceding embodiments, such as Figure 2As shown, the embodiments of the present application provide a network topology change prediction device combined with big data analysis, which comprises:
[0072] a network topology construction module 10 for pre-construction of a coupling network topology, wherein the coupling network topology comprises a physical topology layer and a logical topology layer dynamically mapping connections in real time; a prediction node deployment module 20 for locally deploying a plurality of edge fault prediction nodes in a plurality of physical devices in the physical topology layer; a fault device positioning module 30 for, after a first edge fault identification node analyzes and outputs a first hop prediction fault according to a first real-time operating condition, performing a topology cascade propagation analysis to locate X associated fault devices according to the out-degree connection relationship of the physical topology layer; a front prediction update module 40 for the first edge fault identification node sending the first hop prediction fault to X edge fault identification nodes of the X associated fault devices to trigger a cooperative prediction mechanism, performing a fault front prediction update, and outputting X front prediction faults; a concomitant entity mapping module 50 for performing concomitant entity mapping according to the first hop prediction fault and the X front prediction faults, and outputting a global logical influence domain; a change distribution positioning module 60 for projecting the global logical influence domain to the logical topology layer in the coupling network topology, and positioning a network topology change distribution; and a risk output module 70 for performing a business service dependency chain stall mapping according to the network topology change distribution, and outputting a business service interruption risk.
[0073] Further, the fault device positioning module 30 is configured to perform the following operation steps:
[0074] With a fault intervention response delay as a constraint, a first physical device is called to obtain a plurality of sample fault precursor feature data sets in a plurality of sample device fault scenarios; after a plurality of heterogeneous fault identification models are constructed based on the plurality of sample fault precursor feature data sets, a first fault identification unit is constructed by connecting the plurality of heterogeneous fault identification models in parallel; the first real-time operating condition is divided and sliced with the fault intervention response delay as a division window and 1 / 8 of the fault intervention response delay as an overlap scale to obtain a first operating condition time sequence slice; the first operating condition time sequence slice is sequentially input into the first fault identification unit for fault precursor perception until the first hop prediction fault is output; and the first physical device is taken as a cascade propagation source starting point, and a topology cascade propagation analysis is performed according to the out-degree connection relationship of the physical topology layer to locate the X associated fault devices.
[0075] Further, the fault device positioning module 30 is configured to perform the following operation steps:
[0076] interactively obtaining a plurality of historical fault logs of the plurality of physical devices; performing fault linkage evaluation based on the plurality of historical fault logs to construct a device fault linkage matrix; according to the first physical device, performing row vector correlation degree analysis on the device fault linkage matrix to output Y candidate fault devices; taking the first physical device as a cascading propagation source starting point, performing topological cascading propagation analysis according to the out-degree connection relationship of the physical topology layer, locating C initial fault devices, and then verifying the fault conduction path of the C initial fault devices by using the Y candidate fault devices, and outputting the X associated fault devices.
[0077] Further, the first-hop prediction fault is accompanied by a fault device identifier, a predicted fault type code and a fault time window marker.
[0078] Further, the front-end prediction update module 40 is configured to perform the following operation steps:
[0079] The X edge fault identification nodes load X associated operating conditions according to the fault time window marker; after the X edge fault identification nodes activate the local fault identification model of the X fault identification units according to the predicted fault type code, the X associated operating conditions are input to perform local fault front-end prediction, and the X front-end prediction faults are output.
[0080] Further, the front-end prediction update module 40 is configured to perform the following operation steps:
[0081] interactively obtaining a plurality of reference fault conduction paths, wherein each path node in the reference fault conduction path is accompanied by a sample device identifier and a sample fault type code; calling a first associated device identifier of a first associated fault device; predefining a conduction dependency relationship between the first associated device identifier and the fault device identifier; taking the first associated device identifier, the fault device identifier and the predicted fault type code as a conduction path query key, taking the conduction dependency relationship as a path matching criterion, traversing the plurality of reference fault conduction paths, and comparing to output N sample fault types; after the first edge fault identification node activates the local fault identification model of the first fault identification unit according to the N sample fault types, the first associated operating conditions are input to perform local fault front-end prediction, and the first front-end prediction fault is output.
[0082] Further, the network topology construction module 10 is configured to perform the following operation steps:
[0083] According to the physical connection relationship of the hardware device, a physical topology layer is constructed; according to the logical interconnection topology relationship of the virtualized network entity, a logical topology layer is constructed; based on the real-time configuration perception engine, the physical topology layer and the logical topology layer are dynamically mapped to construct the coupled network topology.
[0084] Further, the accompanying entity mapping module 50 is configured to perform the following steps:
[0085] Based on the inter-layer mapping rule, the first-hop accompanying logical entity of the first-hop predicted failure and the X accompanying logical entities of the X preceding predicted failures are retrieved respectively; the first-hop accompanying logical entity and the X accompanying logical entities are aggregated by graph theory, and the global logical influence domain is output.
[0086] Further, the change distribution positioning module 60 is configured to perform the following steps:
[0087] Based on the physical-logical layer dynamic mapping rule, the global logical influence domain is converted into a logical entity influence strength vector; after projecting the logical entity influence strength vector to the logical topology layer, influence propagation fitting is performed according to the out-degree connection relationship of the logical topology layer, and a logical node vulnerability value matrix is output; based on a preset vulnerability value threshold, the logical node vulnerability value matrix is traversed to identify a plurality of high-risk logical connection paths; the plurality of high-risk logical connection paths are aggregated, and the network topology change distribution is output.
[0088] Through the foregoing detailed description of the network topology change prediction method combined with big data analysis, those skilled in the art can clearly understand the network topology change prediction device combined with big data analysis in the embodiment. Since it corresponds to the method disclosed in the embodiment, it is described relatively simply, and the related part can be referred to the method part description.
[0089] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A network topology change prediction method combined with big data analysis, characterized by, The method comprises: pre-constructing a coupling network topology, wherein the coupling network topology comprises a physical topology layer and a logical topology layer connected in real time dynamically; locally deploying a plurality of edge fault prediction nodes in a plurality of physical devices in the physical topology layer; after the first edge fault identification node analyzes and outputs the first-hop predicted fault according to the first real-time operating condition, performing topological cascade propagation analysis to locate X associated fault devices according to the out-degree connection relationship of the physical topology layer; wherein, after the first edge fault identification node analyzes and outputs the first-hop predicted fault according to the first real-time operating condition, performing topological cascade propagation analysis to locate X associated fault devices according to the out-degree connection relationship of the physical topology layer, the method comprises: constraining the fault intervention response delay, and calling the first physical device to perform networking on a plurality of sample fault precursor feature data sets of a plurality of sample device fault scenarios; after constructing a plurality of heterogeneous fault identification models based on the plurality of sample fault precursor feature data sets, constructing a first fault identification unit by parallel connection of the plurality of heterogeneous fault identification models; taking the fault intervention response delay as a segmentation window and 1 / 8 of the fault intervention response delay as an overlap scale, slidingly segmenting the first real-time operating condition to obtain a first operating condition time sequence slice; sequentially inputting the first operating condition time sequence slice into the first fault identification unit for fault precursor perception until the first-hop predicted fault is outputted; taking the first physical device as a cascade propagation source starting point, performing topological cascade propagation analysis according to the out-degree connection relationship of the physical topology layer to locate the X associated fault devices; the first edge fault identification node sends the first-hop predicted fault to X edge fault identification nodes of the X associated fault devices to trigger a cooperative prediction mechanism, performs fault precursor prediction update, and outputs X precursor predicted faults; performing concomitant entity mapping according to the first-hop predicted fault and the X precursor predicted faults to output a global logical influence domain; projecting the global logical influence domain to the logical topology layer in the coupling network topology to locate a network topology change distribution; according to the network topology change distribution, performing business service dependency chain stall mapping to output a business service interruption risk. 2.The network topology change prediction method in conjunction with big data analysis of claim 1, wherein, taking the first physical device as a cascade propagation source starting point, performing topological cascade propagation analysis according to the out-degree connection relationship of the physical topology layer to locate the X associated fault devices, the method comprises: interactively obtaining a plurality of historical fault logs of the plurality of physical devices; performing fault linkage evaluation based on the plurality of historical fault logs to construct a device fault linkage matrix; according to the row vector correlation degree analysis of the first physical device in the device fault linkage matrix, outputting Y alternative fault devices; taking the first physical device as a cascade propagation source starting point, performing topological cascade propagation analysis according to the out-degree connection relationship of the physical topology layer to locate C initial fault devices, and then using the Y alternative fault devices to verify the fault conduction path of the C initial fault devices to output the X associated fault devices. 3.The network topology change prediction method in conjunction with big data analytics of claim 1, wherein, The first-hop predicted fault is accompanied by a fault device identifier, a predicted fault type code, and a fault time window marker. 4.The network topology change prediction method in conjunction with big data analytics of claim 3, wherein, The first edge fault identification node sends the first-hop predicted fault to the X edge fault identification nodes of the X associated fault devices to trigger a cooperative prediction mechanism, performs a fault front prediction update, and outputs X front prediction faults. The X edge fault identification nodes load the X associated operating conditions according to the fault time window marker. After the X edge fault identification nodes activate the local fault identification model of the X fault identification units according to the predicted fault type code, the X edge fault identification nodes input the X associated operating conditions to perform local fault front prediction, and output the X front prediction faults. 5.The network topology change prediction method in conjunction with big data analytics of claim 4, wherein, After the X edge fault identification nodes activate the local fault identification model of the X fault identification units according to the predicted fault type code, the X edge fault identification nodes input the X associated operating conditions to perform local fault front prediction, and output the X front prediction faults. Interactively obtain a plurality of benchmark fault conduction paths, wherein each path node in the benchmark fault conduction path is accompanied by a sample device identifier and a sample fault type code; Call the first associated device identifier of the first associated fault device; Predefine the conduction dependency relationship between the first associated device identifier and the fault device identifier; Use the first associated device identifier, the fault device identifier, and the predicted fault type code as a conduction path query key, use the conduction dependency relationship as a path matching criterion, traverse the plurality of benchmark fault conduction paths, and compare to output N sample fault types; After the first edge fault identification node activates the local fault identification model of the first fault identification unit according to the N sample fault types, the first edge fault identification node inputs the first associated operating condition to perform local fault front prediction, and outputs the first front prediction fault. 6.The network topology change prediction method in conjunction with big data analytics of claim 1, wherein, Pre-construct a coupling network topology, the method comprising: Construct a physical topology layer according to the physical connection relationship of the hardware device; Construct a logical topology layer according to the logical interconnection topology relationship of the virtualized network entity; Based on a real-time configuration perception engine, dynamically map the physical topology layer and the logical topology layer to construct the output coupling network topology. 7.The network topology change prediction method in conjunction with big data analytics of claim 1, wherein, According to the first-hop predicted fault and the X front prediction faults, perform accompanying entity mapping to output a global logical impact domain, the method comprising: Based on inter-layer mapping rules, respectively retrieve the first-hop accompanying logical entity of the first-hop predicted fault and the X accompanying logical entities of the X front prediction faults; Graphically aggregate the first-hop accompanying logical entity and the X accompanying logical entities to output the global logical impact domain. 8.The network topology change prediction method in conjunction with big data analytics of claim 1, wherein, In the coupling network topology, project the global logical impact domain to the logical topology layer to locate the network topology change distribution, the method comprising: Based on physical-logical layer dynamic mapping rules, convert the global logical impact domain into a logical entity impact strength vector; After projecting the logical entity impact strength vector to the logical topology layer, perform impact propagation fitting according to the out-degree connection relationship of the logical topology layer to output a logical node vulnerability value matrix; traversing the logical node vulnerability matrix based on a preset vulnerability threshold value, identifying a plurality of high-risk logical connection paths; aggregating the plurality of high-risk logical connection paths, and outputting the network topology change distribution.
9. A network topology change prediction device combined with big data analysis, characterized by, The device for implementing the network topology change prediction method based on big data analysis according to any one of claims 1-8, comprising: a network topology construction module for pre-construction of a coupled network topology, wherein the coupled network topology comprises a physical topology layer and a logical topology layer that are dynamically mapped in real time; a prediction node deployment module for locally deploying a plurality of edge fault prediction nodes in a plurality of physical devices in the physical topology layer; a fault device positioning module for, after a first edge fault identification node analyzes and outputs a first-hop prediction fault according to a first real-time operation condition, performing topology cascade propagation analysis and positioning X associated fault devices according to the out-degree connection relationship of the physical topology layer; a front prediction update module for triggering a cooperative prediction mechanism by the first edge fault identification node sending the first-hop prediction fault to X edge fault identification nodes of the X associated fault devices, performing fault front prediction update, and outputting X front prediction faults; a concomitant entity mapping module for performing concomitant entity mapping according to the first-hop prediction fault and the X front prediction faults, and outputting a global logical influence domain; a change distribution positioning module for projecting the global logical influence domain to the logical topology layer in the coupled network topology, and positioning a network topology change distribution; a risk output module for, according to the network topology change distribution, performing service service dependency chain stall mapping, and outputting a business service interruption risk.
Citation Information
Patent Citations
Intelligent visual management method and system for enterprise big data
CN120144416A
Intelligent mechanical safety protection system based on industrial Internet of Things equipment
CN120263475A