Big data-driven network operation status monitoring and management method and system
By establishing a distributed database and a multi-level collaborative network, the problems of multimodal data integration and dynamic adaptation are solved, and accurate detection and early warning of network anomalies are achieved.
Patent Information
- Application Number
- CN202510179871.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Existing technologies are unable to effectively integrate multimodal heterogeneous data and cannot adapt to dynamically changing network environments in real time, resulting in insufficient accuracy in network anomaly detection and delayed early warning.
Establish a distributed database to store multimodal data, perform time synchronization and semantic association, build a multi-dimensional association graph of abnormal states, use a multi-level collaborative network for collaborative identification, combine anomaly location and time series prediction, and generate network status warnings.
It achieves efficient fusion and real-time adaptation of multimodal data, improves the accuracy of network anomaly detection and the foresight of early warning, and provides accurate network anomaly detection and forward-looking early warning.
Smart Images

Figure CN119675988B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital information monitoring technology, and in particular to a big data-driven network operation status monitoring and management method and system. Background Art
[0002] As networks grow in size and complexity, traditional network monitoring systems face numerous challenges. Existing technologies primarily rely on static rules and simple data analysis, failing to effectively integrate heterogeneous data from diverse devices, links, and user behaviors, resulting in inaccurate anomaly detection. Furthermore, existing systems struggle to adapt to dynamically changing network environments and lack the ability to predict potential risks. Summary of the Invention
[0003] This application provides a big data-driven network operation status monitoring and management method and system to solve the technical problems that existing technologies are unable to effectively integrate multimodal heterogeneous data and adapt to dynamically changing network environments in real time, resulting in insufficient network anomaly detection accuracy and delayed warning.
[0004] The first aspect of the present application provides a big data-driven network operation status monitoring and management method, the method comprising: establishing a distributed database, the distributed database storing multimodal data obtained from network equipment, link status, user behavior and security logs; after time synchronization of the distributed database, establishing an abnormal state multidimensional association graph for the distributed database based on semantic association; establishing a multi-level collaborative network, initializing the multi-level collaborative network based on the abnormal state multidimensional association graph, inputting the time-synchronized distributed database into the multi-level collaborative network, and outputting a collaborative identification result, the multi-level collaborative network comprising a device-level identification layer, a link-level identification layer, a system-level identification layer and a collaborative analysis layer; after locating the anomaly of the distributed database according to the collaborative identification result, executing a time series development forecast and establishing additional abnormal results; after compensating the collaborative identification result according to the additional abnormal result, generating a network status warning.
[0005] The second aspect of the present application provides a big data-driven network operation status monitoring and management system, the system comprising: a database establishment module, the database establishment module being used to establish a distributed database, the distributed database storing multimodal data obtained from network devices, link status, user behavior and security logs; an abnormal state association module, the abnormal state association module being used to time synchronize the distributed database and then establish an abnormal state multidimensional association graph for the distributed database based on semantic association; a multi-level collaborative identification module, the multi-level collaborative identification module being used to establish a multi-level collaborative network, initialize the multi-level collaborative network based on the abnormal state multidimensional association graph, input the time-synchronized distributed database into the multi-level collaborative network, and output a collaborative identification result, the multi-level collaborative network comprising a device-level identification layer, a link-level identification layer, a system-level identification layer and a collaborative analysis layer; an abnormal development prediction module, the abnormal development prediction module being used to locate the abnormality of the distributed database according to the collaborative identification result, perform time series development prediction, and establish additional abnormal results; an identification result compensation module, the identification result compensation module being used to compensate the collaborative identification result according to the additional abnormal result and then generate a network status warning.
[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0007] The big data-driven network operation status monitoring and management method and system provided in this application relate to the field of digital information monitoring technology. By establishing a distributed database to store multimodal data, performing time synchronization and semantic association, constructing a multidimensional association diagram of abnormal states, and utilizing a multi-level collaborative network for collaborative identification, anomaly positioning and time series prediction are combined to generate additional anomaly results. Finally, the identification results are compensated and a network status warning is generated. This solves the technical problem that the existing technology cannot effectively integrate multimodal heterogeneous data and adapt to the dynamically changing network environment in real time, resulting in insufficient network anomaly detection accuracy and delayed warning. This achieves the technical effect of accurate network anomaly detection and forward-looking warning through big data-driven multimodal data fusion and multi-level collaborative analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0009] Figure 1 A flowchart of a big data-driven network operation status monitoring and management method provided in an embodiment of the present application;
[0010] Figure 2 A schematic diagram of the structure of a big data-driven network operation status monitoring and management system provided in an embodiment of the present application.
[0011] Explanation of the accompanying symbols: database establishment module 11, abnormal state association module 12, multi-level collaborative identification module 13, abnormal development prediction module 14, identification result compensation module 15. DETAILED DESCRIPTION
[0012] This application provides a big data-driven network operation status monitoring and management method and system to solve the technical problems that existing technologies are unable to effectively integrate multimodal heterogeneous data and adapt to dynamically changing network environments in real time, resulting in insufficient network anomaly detection accuracy and delayed warning.
[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices.
[0015] Example 1, as Figure 1 As shown, the present application provides a big data driven network operation status monitoring and management method, the method comprising:
[0016] P10: Establish a distributed database that stores multimodal data obtained from network devices, link status, user behavior, and security logs.
[0017] Specifically, establishing a distributed database is fundamental to network health monitoring. By deploying a distributed database system, multimodal data from network devices (such as router and switch performance metrics), link status (parameters reflecting network quality, such as latency, bandwidth, and packet loss), user behavior (such as access frequency and page click paths), and security logs (such as failed logins and malicious IP access records) can be uniformly stored and managed. A distributed database is a system that distributes data across multiple nodes, providing high availability and scalability. In this step, select a distributed database system that supports multimodal data storage to meet the storage requirements for structured data (such as link metric tables), semi-structured data (such as user behavior logs in JSON format), and unstructured data (such as security log text).
[0018] First, the data acquisition module collects data from multiple sources in real time. For example, network device CPU utilization and memory usage can be obtained through the SNMP (Simple Network Management Protocol) interface. Link status can be recorded using traffic monitoring tools such as NetFlow. User behavior data can be collected using website analytics tools, while security logs can be collected from firewalls or IDS (Intrusion Detection Systems). To ensure unified storage of this multi-source data, the database must support timestamp indexing and flexible insertion of multiple data formats. Sharding should also be used to distribute storage across multiple nodes to enhance system throughput.
[0019] Furthermore, to ensure data consistency and chronological order across different nodes, data in the distributed database is timestamped (e.g., using the NTP network time protocol) and stored in chronological order, enhancing the accuracy of subsequent time series analysis. This approach integrates multimodal data from diverse sources into a unified distributed storage platform, creating an efficient and flexible data storage system. This design not only improves real-time data access but also provides a high-quality input data source for subsequent anomaly analysis and model prediction.
[0020] P20: After time synchronization is performed on the distributed database, a multi-dimensional association graph of abnormal states is established for the distributed database based on semantic association.
[0021] Furthermore, step P20 in the embodiment of the present application further includes:
[0022] P21: Execute attribute feature extraction of the data source and establish an attribute feature set; P22: After standardizing the attribute feature set, calculate the covariance matrix, which represents the linear relationship of the attributes; P23: Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors; P24: Use the eigenvalues and eigenvectors to perform principal component selection and establish a key attribute set; P25: Perform semantic analysis on the key attribute set to establish strong correlation attributes, potential correlation attributes and background attributes; P26: Perform mutual information analysis on the key attribute set to establish attribute correlation; P27: Use the strong correlation attributes, the potential correlation attributes, the background attributes and attribute correlation to establish a multidimensional correlation diagram of abnormal states.
[0023] It should be understood that after completing the time synchronization of distributed database data, building a multidimensional association graph of abnormal states based on semantic associations requires a series of steps: extracting data features, analyzing attribute relationships, and ultimately constructing the association graph. This process begins with attribute feature extraction, gradually constructing a logical relationship network for the data, and finally visually presenting abnormal associations between the data.
[0024] First, attribute feature extraction is performed on the data source. This involves extracting the core attributes of each data type from the distributed database. For example, CPU utilization and memory usage can be extracted from network device data; packet loss rate and latency can be extracted from link status; access frequency and path depth can be extracted from user behavior; and the number of abnormal login events can be extracted from security logs. The extracted attribute features are stored as an attribute feature set, which is a collection of key attributes from all source data.
[0025] Next, the attribute feature set is normalized, which means that attributes with different dimensions and ranges are unified onto comparable scales. For example, CPU utilization (0100%) and packet loss rate (01) are normalized to the same range based on their maximum and minimum values. The normalized data is used to calculate the covariance matrix, a mathematical tool used to measure the linear relationship between different attributes. For example, if an increase in a device's CPU utilization is positively correlated with an increase in its network latency, their values in the covariance matrix will be positive.
[0026] Next, the covariance matrix is subjected to eigenvalue decomposition, a mathematical simplification of complex attribute relationships into key feature representations. The eigenvalue represents the importance of each principal component, while the eigenvector represents the contribution of each attribute to that principal component. For example, if the first principal component is associated with the device's CPU utilization and link latency, the eigenvector can be used to quantify its weight.
[0027] Furthermore, principal component selection is performed based on eigenvalues and eigenvectors to filter out the most critical attributes for abnormal conditions from a large number of attributes, forming a key attribute set. For example, through analysis, it may be found that CPU utilization, packet loss rate, and abnormal login events are the most representative attributes of the current system anomaly.
[0028] Next, semantic analysis is performed on the key attribute set. By analyzing the causal relationship and business context between attributes, the attributes are divided into strongly associated attributes (such as CPU utilization is closely related to latency), potentially associated attributes (such as user access frequency may affect device load), and background attributes (such as the potential impact of weather conditions on network stability).
[0029] Subsequently, combined with mutual information analysis, we further quantify the degree of information sharing between key attributes, that is, analyze the nonlinear correlation between attributes. For example, mutual information analysis may reveal a hidden nonlinear correlation between an increase in packet loss rate and a specific user's access behavior.
[0030] Finally, a multidimensional association graph of abnormal conditions is constructed using strongly correlated attributes, potentially correlated attributes, background attributes, and the attribute correlations obtained through analysis. This model is represented as a graph structure, where nodes represent key attributes and edges indicate the strength of the associations between them. For example, one node represents CPU utilization, another node represents link packet loss rate, and the edge weight between the two indicates the strength of their association. This multidimensional association graph not only reveals the propagation path of abnormalities in the current system but also reveals potential hidden associations, providing deeper insights for anomaly analysis.
[0031] Through the above steps, the constructed multi-dimensional correlation diagram of abnormal states realizes the intuitive visualization of complex multimodal data relationships, reveals the propagation mechanism of abnormal states, and provides scientific support for subsequent abnormality positioning and risk prediction.
[0032] Furthermore, step P27 of the embodiment of the present application further includes:
[0033] P27-1: Use the Kafka stream processing framework to capture the incoming data in the distributed database in real time; P27-2: Based on the incoming data, perform association calculations on the nodes and edges in the abnormal state multidimensional relationship graph, and dynamically update the graph structure according to the association calculation results.
[0034] Optionally, in order to achieve real-time dynamic update of the multi-dimensional association graph of abnormal status, this application further introduces the Kafka stream processing framework and dynamic graph structure update mechanism to cope with the continuous influx of new data in the distributed database, ensuring that the association graph always reflects the latest network operation status and potential anomalies.
[0035] First, we leverage the Kafka stream processing framework to capture incoming data from distributed databases in real time. Kafka is a high-throughput distributed message queuing system specifically designed to process real-time data streams. It continuously receives incoming data from distributed databases through a producer-consumer model. For example, when a network device uploads new performance metrics (such as increased CPU utilization) or a user action triggers a security event log (such as multiple failed logins), Kafka immediately writes this data as a message to its topic and forwards it through the stream processing pipeline to the graph update module for analysis.
[0036] Next, the association calculation is performed on the nodes and edges in the association graph based on the incoming data. In the multidimensional association graph, nodes represent key attributes (such as CPU utilization, packet loss rate, etc.), and edges represent the strength of the association between attributes. The core of the association calculation is to use an incremental update algorithm to adjust only the nodes and edges affected by the new data without rebuilding the entire graph structure. For example, when new security log data indicates that a certain IP address is associated with multiple abnormal login events, the system will update the "IP address" node and the edge weight between it and the "number of abnormal logins" node. If the new data also reveals the potential impact of a user's access behavior on device performance indicators, new nodes and edges will be dynamically added to incorporate these relationships into the graph in real time.
[0037] To achieve this dynamic update, a streaming computing strategy based on sliding windows can be employed. By defining a time range (e.g., the last 5 minutes), sliding windows ensure that computations are limited to recent data streams, improving processing efficiency and ensuring the real-time nature of the graph structure. Furthermore, to enhance computational accuracy, the system uses weighted mutual information analysis to quantify the impact of new data on existing node and edge relationships. For example, new packet loss rate data may increase the edge weight between a link delay node and a packet loss rate node, reflecting a closer relationship.
[0038] Ultimately, the graph structure is dynamically updated based on the correlation calculation results, ensuring that the multidimensional correlation graph reflects the propagation paths of anomalies in network operations in real time. For example, if a sudden delay on a link causes a degradation in user experience for multiple users, the system dynamically adds relevant nodes and edges, marking the key points of the anomaly and the scope of impact in real time. This update mechanism enables the multidimensional correlation graph to quickly adapt to environmental changes, providing more accurate anomaly location and risk warning capabilities.
[0039] By introducing the Kafka stream processing framework and dynamic update algorithms, this step significantly improves the system's ability to process real-time data, making the multidimensional correlation graph efficient and flexible in its adaptive nature. This allows for global monitoring of network operation status while also capturing local emergencies, providing technical support for timely response measures.
[0040] P30: Establish a multi-level collaborative network. After initializing the multi-level collaborative network based on the abnormal state multi-dimensional association diagram, input the time-synchronized distributed database into the multi-level collaborative network and output the collaborative identification results. The multi-level collaborative network includes a device-level identification layer, a link-level identification layer, a system-level identification layer and a collaborative analysis layer.
[0041] Furthermore, the distributed database after time synchronization is input into the multi-level collaborative network to output the collaborative recognition result. In the embodiment of the present application, step P30 further includes:
[0042] P31: Use the device-level identification layer to perform device anomaly analysis on the distributed database as follows:
[0043] ;in, Characterization equipment In time The abnormal status score is used to quantify the health status or abnormality of the device. is the total number of device features, which indicates the number of features used to describe the device status, such as temperature, pressure, current, etc. Characterize any device feature, Indicates the importance of each feature in abnormal state analysis. The weight is determined by historical data or expert knowledge. Characterization The characteristic value of each device characteristic, Characterizing device characteristics In the time window The trend function within the time window is used to describe the characteristic value Changes within, is the weight of the historical anomaly pattern matching factor, which determines the impact of historical data on the current anomaly score. Characterization equipment The matching score of the historical abnormal pattern library can measure the similarity between the current state and the historical abnormal pattern. is the weight of the equipment synergy impact factor, is the global synergistic impact factor.
[0044] It should be understood that the establishment of a multi-level collaborative network is the core step in analyzing system anomaly states. Based on a multi-dimensional association graph of anomaly states, it enables comprehensive anomaly identification from the device level to the system level by initializing the multi-level collaborative network and inputting it into a time-synchronized distributed database. The multi-level collaborative network consists of four main layers: device-level identification layer, link-level identification layer, system-level identification layer, and collaborative analysis layer. Each layer works together to output collaborative identification results.
[0045] For example, the core of equipment abnormality analysis is to use the formula to quantify the abnormal state of the equipment, extract the multi-dimensional feature data of the equipment from the distributed database, and calculate the characteristics of each feature in the time window in turn. The changing trend within For example, if the voltage of a device fluctuates abnormally in the past 10 minutes, the trend function will capture this dynamic characteristic. Then, the current state of the device is matched with the historical abnormal pattern library, and the matching score is calculated through the pattern matching algorithm. For example, if the vibration frequency of a device matches the previously recorded mechanical failure mode, the matching score is high. Next, the multi-dimensional correlation diagram of abnormal states is used to analyze the collaborative relationship between devices and calculate the global collaborative impact factor of the devices. , assess the potential impact of the device anomaly on other devices or links. Finally, combining all calculation results, the device anomaly status score is calculated comprehensively using feature weights, matching factor weights, and synergy factor weights, and a device-level anomaly analysis result is generated.
[0046] Through analysis at the device-level recognition layer, we can quickly identify potential anomalies in individual devices and, combined with collaborative relationships, expand the coverage and depth of anomaly detection. This analysis provides precise input to the link-level recognition layer, ensuring that the anomaly detection logic of the multi-layer collaborative network is gradually optimized from the lowest level to the highest level.
[0047] Device-level identification results are synchronized to the collaborative analysis layer, where contextual analysis and anomaly propagation mechanisms are used to comprehensively analyze device-level, link-level, and system-level results, ultimately outputting a global collaborative identification result. This process not only improves the accuracy of anomaly detection but also reduces the impact of single points of failure on the system through multi-layered collaborative mechanisms, ensuring efficient operation of the entire network and real-time anomaly response capabilities.
[0048] Furthermore, step P32 of the embodiment of the present application further includes:
[0049] P32-1: Use the link-level identification layer to identify link anomalies in the distributed database as follows: ;in, Characterization Link In time The interaction anomaly score of Characterization and Links The set of adjacent links, Characterizing any link Adjacent links, Characterization Link and Link The interaction weight, Characterization Link and Link The interaction function of Characterizes the abnormal score of the link itself, For Link The anomaly propagation factor characterizes the diffusion effect of anomalies on adjacent links. and They are the weights of their own characteristic anomalies and anomaly propagation respectively; P32-2: Establish link anomaly identification results, perform system anomaly identification based on a distributed database through the system-level identification layer, establish system anomaly identification results, synchronize the link anomaly identification results, the system anomaly identification results, and the device anomaly identification results to the collaborative analysis layer to establish collaborative identification results.
[0050] In one possible embodiment of the present application, at the link-level identification layer, the above-mentioned anomaly score calculation formula is used to perform a detailed analysis of the link state information in the distributed database and perform an interactive anomaly score. The interactive anomaly score for each link is calculated and compared with an anomaly score threshold to identify potential link anomalies and establish a link anomaly identification result. The anomaly score threshold can be set based on empirical data.
[0051] Then, after completing the link-level anomaly analysis, the entire network system can be further identified for anomalies through the system-level identification layer. The system-level identification layer considers not only the status of a single link, but also the interaction between links, the status of the device, and the overall health of the system. Afterwards, the link-level anomaly identification results, the device-level anomaly identification results, and the system-level anomaly identification results are synchronized to the collaborative analysis layer. At the collaborative analysis layer, the identification results of each layer are comprehensively analyzed to form a comprehensive collaborative identification result, reflecting the anomalies of each element in the network and their mutual influence. Through this hierarchical analysis, this solution can accurately identify network anomalies from multiple levels and multiple angles, and can promptly discover potential chain reactions within the system, providing comprehensive network security protection.
[0052] P40: After locating the anomaly of the distributed database based on the collaborative identification results, perform time series development prediction and establish additional anomaly results.
[0053] Furthermore, step P40 in this embodiment of the present application further includes:
[0054] P41: Create a time zero point based on the collaborative identification result; P42: Perform data backtracking extraction based on the time zero point and anomaly positioning, perform time series development prediction based on the data backtracking extraction result, and establish the additional anomaly result.
[0055] Optionally, based on the collaborative identification results from the previous step, the distributed database anomaly is first located, followed by time series development forecasting, and finally, additional anomaly results are generated. The core of this step is to accurately locate anomalies and combine time series analysis of historical data to predict future anomaly trends, ensuring that the network management system can identify potential risks in advance.
[0056] First, the collaborative identification results are used to locate anomalies in the network. This pinpoints the specific source of the anomaly, which may involve a device, a link, or the system itself. This process is crucial for reducing false positives because it helps the system focus its resources on analyzing the most urgent and critical anomalies, rather than wasting energy on irrelevant ones.
[0057] Next, based on the positioning results, the time series development forecasting phase begins. This process analyzes historical data trends to predict possible future anomalies. The system uses a data backtracking approach, starting from a specific time zero and working backwards to extract relevant time series data. This data provides the foundation for subsequent forecasting and analysis, ensuring accurate and timely forecasts. Using forecasting models (such as ARIMA or LSTM), the system not only monitors current anomaly changes but also analyzes their potential expansion and duration, providing managers with more comprehensive network status warnings.
[0058] After executing a time series forecast, additional anomaly results are generated. These results can reveal potential future risk points and changing trends. The introduction of additional anomaly results enables the system to identify potential issues before anomalies occur, providing personnel with ample time to take preventative measures. This not only enhances the system's early warning capabilities but also optimizes resource allocation and risk management strategies.
[0059] In summary, the combination of anomaly location and time series prediction provides a highly forward-looking and accurate early warning mechanism for network management systems. It can quickly identify and predict abnormal events that could impact network stability, providing decision support for operations and maintenance personnel, enabling them to proactively implement necessary interventions and ensure stable network operation.
[0060] P50: After compensating the collaborative identification result according to the additional abnormal result, a network status warning is generated.
[0061] Specifically, the previously obtained collaborative identification results are compensated based on the generated additional anomaly results, and a network status warning is ultimately generated. The purpose of this process is to make the final warning more accurate and timely by correcting the predicted anomaly trends, thereby providing more reliable decision support for network operations and maintenance.
[0062] First, the collaborative identification results are compensated with additional anomaly results. These additional anomaly results, derived from time series forecasts, provide information about the future state of the network, helping the system detect potential anomalies that might have been missed by the initial collaborative identification results. In this step, the system combines these additional anomaly results with the previous collaborative identification results to correct and adjust them, ensuring that the collaborative identification results more accurately reflect the current health of the network while taking into account future trends.
[0063] By compensating for identification results, errors caused by prediction lags or data fluctuations can be eliminated, thereby reducing the probability of false positives and improving the accuracy of anomaly detection. For example, in some cases, initial collaborative identification may fail to capture potential abnormal trends in certain devices or links. Additional anomaly results can promptly identify and correct these omissions, thus avoiding missing important early warning signals.
[0064] Ultimately, the compensated identification results are used to generate network status alerts. These alerts, based on more precise and complete anomaly analysis, accurately reflect the current network's operational status and potential risks. Network status alerts not only provide immediate anomaly warnings but also prioritize alerts based on severity, providing managers with appropriate response strategies. For example, based on the compensated results, alerts can be issued for device failures, link congestion, or system bottlenecks, enabling operations personnel to intervene before issues escalate, improving the accuracy and timeliness of alerts.
[0065] Furthermore, the embodiment of the present application further includes step P60, which further includes:
[0066] P61: Establish a network status monitoring database; P62: Perform an early warning superposition analysis based on the network status monitoring database and the network status early warning, and issue an early warning according to the early warning superposition analysis result.
[0067] In a possible embodiment of the present application, a network status monitoring database can be further established, and an early warning superposition analysis can be performed based on the database and the network status early warning, and finally an early warning alarm can be generated according to the analysis results. That is, through multi-dimensional information integration and analysis, a more comprehensive and detailed early warning report can be provided for network management, thereby improving the response capability and accuracy to potential problems.
[0068] First, a network status monitoring database was built to store network operation information from various data sources. This database aggregates real-time operational data from various network devices, link status, user behavior, historical warning records, and more. This data provides strong support for subsequent analysis. The core function of the monitoring database is to continuously update and store dynamic network status information, providing a unified data platform for the entire warning system.
[0069] Next, based on the various data stored in the monitoring database and previously generated network status alerts, an overlay analysis is performed. This analysis aims to integrate alert information from different levels and sources, identifying more complex or potential risks through multi-dimensional cross-validation. For example, an abnormal alert for a device might be related to factors such as excessive link load or insufficient system processing capacity. By synthesizing this information, overlay analysis can uncover potential chain reactions or more complex system failure modes.
[0070] After the early warning overlay analysis is complete, pre-alarm alerts are generated based on the analysis results. These alerts provide detailed warning information based on the severity and impact of the anomaly, along with appropriate response recommendations. This information helps network managers understand the current health of the network, potential issues, and possible solutions, enabling timely intervention to prevent further deterioration.
[0071] In summary, the embodiments of the present application have at least the following technical effects:
[0072] This application establishes a distributed database to store multimodal data from network devices, link status, user behavior, and security logs, synchronizes time and semantically correlates these data, and constructs a multidimensional association graph of abnormal states. Subsequently, a multi-level collaborative network (including device-level, link-level, and system-level identification layers, as well as a collaborative analysis layer) is used for collaborative identification. This combines anomaly location and time series prediction to generate additional anomaly results. Ultimately, the identification results are compensated to generate network status warnings, enabling accurate detection and prediction of network anomalies.
[0073] The technical effect of achieving accurate network anomaly detection and forward-looking early warning has been achieved through big data-driven multimodal data fusion and multi-level collaborative analysis.
[0074] Example 2, based on the same inventive concept as the big data driven network operation status monitoring and management method in the above embodiment, Figure 2 As shown, the present application provides a big data driven network operation status monitoring and management system. The system and method embodiments in the present application are based on the same inventive concept. The system includes:
[0075] The database establishment module 11 is used to establish a distributed database, which stores multimodal data obtained from network devices, link status, user behavior and security logs.
[0076] The abnormal state association module 12 is used to establish an abnormal state multi-dimensional association graph for the distributed database based on semantic association after time synchronization of the distributed database.
[0077] A multi-level collaborative identification module 13 is used to establish a multi-level collaborative network. After initializing the multi-level collaborative network based on the abnormal state multi-dimensional association diagram, the time-synchronized distributed database is input into the multi-level collaborative network, and the collaborative identification result is output. The multi-level collaborative network includes a device-level identification layer, a link-level identification layer, a system-level identification layer and a collaborative analysis layer.
[0078] The abnormal development prediction module 14 is used to perform time series development prediction after locating the abnormality of the distributed database according to the collaborative identification result, and establish additional abnormality results.
[0079] The recognition result compensation module 15 is used to generate a network status warning after compensating the collaborative recognition result according to the additional abnormal result.
[0080] Furthermore, the abnormal state association module 12 is further configured to perform the following steps:
[0081] Execute attribute feature extraction of the data source and establish an attribute feature set; after standardizing the attribute feature set, calculate the covariance matrix, the covariance matrix represents the linear relationship of the attributes; perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors; use the eigenvalues and eigenvectors to perform principal component selection and establish a key attribute set; perform semantic analysis on the key attribute set to establish strong correlation attributes, potential correlation attributes and background attributes; perform mutual information analysis on the key attribute set to establish attribute correlation; use the strong correlation attributes, the potential correlation attributes, the background attributes and attribute correlation to establish a multidimensional correlation diagram of abnormal conditions.
[0082] Furthermore, the abnormal state association module 12 is further configured to perform the following steps:
[0083] Use the Kafka stream processing framework to capture incoming data from distributed databases in real time;
[0084] Based on the inflow data, association calculation of nodes and edges in the abnormal state multidimensional relationship graph is performed, and the graph structure is dynamically updated according to the association calculation result.
[0085] Furthermore, the multi-level collaborative identification module 13 is further configured to perform the following steps:
[0086] The device-level identification layer is used to perform device anomaly analysis on the distributed database as follows: ;in, Characterization equipment In time Abnormal status score, is the total number of device features, Characterize any device feature, Characterizing device characteristics The weight of Characterization The characteristic value of each device characteristic, Characterizing device characteristics In the time window The changing trend function within is the weight of the historical abnormal pattern matching factor, Characterization equipment The matching score of the historical abnormal pattern library, is the weight of the equipment synergy impact factor, is a global collaborative influencing factor; establishes a device anomaly identification result, synchronizes the device anomaly identification result to the collaborative analysis layer, and establishes a collaborative identification result.
[0087] Furthermore, the multi-level collaborative identification module 13 is further configured to perform the following steps:
[0088] The link-level identification layer is used to identify link anomalies in a distributed database as follows:
[0089] ;in, Characterization Link In time The interaction anomaly score of Characterization and Links The set of adjacent links, Characterizing any link Adjacent links, Characterization Link and Link The interaction weight, Characterization Link and Link The interaction function of Characterizes the abnormal score of the link itself, For Link The anomaly propagation factor characterizes the diffusion effect of anomalies on adjacent links. and They are the weights of their own characteristic anomalies and anomaly propagation respectively; establish a link anomaly identification result, perform system anomaly identification based on a distributed database through the system-level identification layer, establish a system anomaly identification result, synchronize the link anomaly identification result, the system anomaly identification result, and the device anomaly identification result to the collaborative analysis layer to establish a collaborative identification result.
[0090] Furthermore, the abnormal development prediction module 14 is further configured to perform the following steps:
[0091] A time zero point is created based on the collaborative identification result; data backtracking extraction is performed based on the time zero point and anomaly location, and a time series development forecast is performed according to the data backtracking extraction result to establish the additional anomaly result.
[0092] Furthermore, the system further comprises:
[0093] A network status monitoring database establishment module is used to establish a network status monitoring database; an early warning superposition analysis module is used to perform early warning superposition analysis based on the network status monitoring database and the network status early warning, and issue early warnings according to the early warning superposition analysis results.
[0094] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0095] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
[0096] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.
Claims
1. A big data driven network operation status monitoring and management method, characterized in that: The method comprises: Establishing a distributed database that stores multimodal data obtained from network devices, link status, user behavior, and security logs; After time synchronization is performed on the distributed database, a multi-dimensional association graph of abnormal states is established for the distributed database based on semantic association; Establishing a multi-level collaborative network, initializing the multi-level collaborative network based on the abnormal state multi-dimensional association graph, inputting the time-synchronized distributed database into the multi-level collaborative network, and outputting collaborative identification results, wherein the multi-level collaborative network includes a device-level identification layer, a link-level identification layer, a system-level identification layer, and a collaborative analysis layer; After locating anomalies in the distributed database based on the collaborative identification results, a time series development forecast is performed to establish additional anomaly results; After compensating the collaborative identification result according to the additional abnormal result, generating a network status warning; The method of establishing a multi-dimensional association graph of abnormal states for a distributed database based on semantic association includes: Execute attribute feature extraction of data source and establish attribute feature set; After normalizing the attribute feature set, a covariance matrix is calculated, where the covariance matrix represents the linear relationship of the attributes; Performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors; Performing principal component selection using the eigenvalues and the eigenvectors to establish a key attribute set; Performing semantic analysis on the key attribute set to establish strong correlation attributes, potential correlation attributes and background attributes; Performing mutual information analysis on the key attribute set to establish attribute correlation; Establishing a multi-dimensional association graph of abnormal states using the strong association attributes, the potential association attributes, the background attributes, and attribute correlations; The step of establishing a multi-dimensional association graph of abnormal states further includes: Use the Kafka stream processing framework to capture incoming data from distributed databases in real time; Performing association calculations on nodes and edges in a multidimensional relationship graph of abnormal states based on the inflow data, and dynamically updating the graph structure according to the association calculation results; After locating the anomaly of the distributed database according to the collaborative identification results, executing time series development prediction and establishing additional anomaly results include: Creating a time zero point based on the collaborative recognition result; Performing data backtracking extraction based on the time zero point and anomaly location, performing time series development prediction based on the data backtracking extraction results, and establishing the additional anomaly results; The method further comprises: Establish a network status monitoring database; Performing an early warning superposition analysis based on the network status monitoring database and the network status early warning, and issuing an early warning according to the early warning superposition analysis result; The step of inputting the time-synchronized distributed database into the multi-level collaborative network and outputting collaborative recognition results includes: The device-level identification layer is used to perform device anomaly analysis on the distributed database as follows: Among them, S i (t) represents the abnormal status score of device i at time t, n is the total number of device features, j represents any device feature, w j The weight representing device feature j, X j (t) The characteristic value representing the jth device characteristic, f j (X j (t),Δt) represents the changing trend function of device feature j within the time window Δt, α is the weight of the historical abnormal pattern matching factor, P i Characterizes the matching score of the historical abnormal pattern library of device i, β is the weight of the device collaborative impact factor, and (X(t)) is the global collaborative impact factor; Establishing a device anomaly identification result, synchronizing the device anomaly identification result to the collaborative analysis layer, and establishing a collaborative identification result; The step of synchronizing the device anomaly identification result to the collaborative analysis layer to establish the collaborative identification result includes: The link-level identification layer is used to identify link anomalies in a distributed database as follows: L score (p,t)=∑ q∈N(p) φ pq ·F pq (t)+δ·Self(R p (t))+ε·Prop(p,t); Among them, L score (p,t) represents the interaction anomaly score of link p at time t, N(p) represents the set of links adjacent to link p, q represents any link adjacent to link p, φ pq Characterizes the interaction weight between link p and link q, F pq (t) represents the interaction function between link p and link q, Self(R p (t)) represents the link's own characteristic anomaly score, Prop(p,t) is the anomaly propagation factor of link p, which represents the diffusion effect of the anomaly on adjacent links, and δ and ε are the weights of the own characteristic anomaly and anomaly propagation, respectively; Establish link anomaly identification results, perform system anomaly identification based on a distributed database through the system-level identification layer, establish system anomaly identification results, synchronize the link anomaly identification results, the system anomaly identification results, and the device anomaly identification results to the collaborative analysis layer to establish collaborative identification results.
2. Big data driven network operation status monitoring and management system, characterized by: The system is used to perform the method according to claim 1, and the system includes: A database establishment module, the database establishment module is used to establish a distributed database, the distributed database stores multimodal data obtained from network devices, link status, user behavior and security logs; An abnormal state association module, the abnormal state association module is used to establish an abnormal state multi-dimensional association graph for the distributed database based on semantic association after time synchronization of the distributed database; A multi-level collaborative identification module, which is used to establish a multi-level collaborative network, initialize the multi-level collaborative network based on the abnormal state multi-dimensional association graph, input the time-synchronized distributed database into the multi-level collaborative network, and output a collaborative identification result. The multi-level collaborative network includes a device-level identification layer, a link-level identification layer, a system-level identification layer, and a collaborative analysis layer; An abnormal development prediction module, which is used to perform time series development prediction and establish additional abnormal results after locating anomalies in the distributed database based on the collaborative identification results; The recognition result compensation module is used to generate a network status warning after compensating the collaborative recognition result according to the additional abnormal result.
Citation Information
Patent Citations
Configuration method and system for overall electrical scheme of machine room
CN119443720A