Network detection method and device, computer device, readable storage medium and program product
By acquiring network topology resources and traffic information, and using a hybrid anomaly detection model and a large language model for network alarm analysis, the problems of alarm interference and insufficient human experience in traditional methods are solved. This enables accurate identification and efficient screening of network alarms, improving the accuracy and efficiency of network detection.
Patent Information
- Application Number
- CN202411966324.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2044-12-30
AI Technical Summary
When faced with large-scale cloud computing and IoT platforms, traditional network operation and maintenance methods suffer from the problem that the increased data volume leads to a large number of minor or invalid alarms interfering with core alarms. Furthermore, the analysis relying on human experience is not very accurate and it is difficult to accurately identify the root cause of network alarms.
By acquiring network topology resource association information and traffic information, the target anomaly score of alarm information is calculated using an anomaly detection hybrid model. Root cause alarm identification is performed by combining probabilistic graphical models and large language models, extracting target alarm information and its dependencies to achieve accurate screening and intelligent analysis.
It effectively reduces false alarms and false alarms, significantly improves the accuracy of identifying the root cause of alarms in network security management, and enhances the accuracy and efficiency of network detection results.
Smart Images

Figure CN119696995B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a network detection method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] With the development of cloud computing, Internet of Things, big data and other technologies, network operation and maintenance also faces new challenges. In terms of cloud resource management, data security protection, optimization of large-scale network architecture, network security operation and maintenance management needs to be further strengthened.
[0003] In the traditional method, when the operation and maintenance personnel analyze the multi-source heterogeneous data by relying on artificial experience, they will first integrate diversified data from different systems, devices and platforms, including logs, monitoring indicators, alarm information, etc. By comparing historical data and current state, combining business background and artificial experience, abnormal indicators or alarms are identified, and the abnormal data is analyzed in depth to locate the possible link of the problem, such as server, network, database or application, etc. Based on the understanding of system architecture and operation mechanism, the operation and maintenance personnel comprehensively judge the root cause of the abnormality and analyze the root cause to determine the hidden trouble troubleshooting scheme.
[0004] However, in the current traditional method, with the increase of data volume of cloud computing, Internet of Things, big data and other platforms, the alarm information also increases, a large number of secondary or invalid alarms interfere with the core alarms, cover up the real alarm causes, and the subjectivity of the operation and maintenance personnel relying on artificial experience for analysis leads to poor accuracy of network alarm root cause detection and hidden trouble troubleshooting scheme. SUMMARY
[0005] Therefore, it is necessary to provide a network detection method, device, computer equipment, computer readable storage medium and computer program product to solve the above technical problems.
[0006] In a first aspect, the present application provides a network detection method, comprising:
[0007] obtaining preset topological resource association information, traffic information and first alarm information of a plurality of alarm types of a target network;
[0008] calculating a target anomaly score corresponding to each of the first alarm information, and preliminarily screening the first alarm information based on the target anomaly score to obtain second alarm information;
[0009] performing root cause alarm identification on the second alarm information based on the preset topological resource association information and the target anomaly score corresponding to the second alarm information, and extracting target alarm information and a dependency relationship corresponding to the target alarm information;
[0010] According to the preset large language model, the traffic information, the target alarm information and the dependency relationship are analyzed and processed to obtain a network detection result.
[0011] In one of the embodiments, the computing the target anomaly score corresponding to each of the first alarm information comprises:
[0012] According to the second correlation relationship between the alarm sample data in the preset transaction library, the first correlation relationship between the alarm types corresponding to each of the first alarm information is determined;
[0013] According to the first correlation relationship, the first alarm information is divided into a plurality of alarm groups;
[0014] Based on the anomaly detection hybrid model and the first correlation relationship corresponding to the alarm group, the target anomaly score corresponding to each of the first alarm information is calculated.
[0015] In one of the embodiments, before the first correlation relationship between the alarm types corresponding to each of the first alarm information is determined according to the first correlation relationship between the alarm sample data in the preset transaction library, the method further comprises:
[0016] Obtain a sample data set, and divide the sample data set into a plurality of sample data subsets; the sample data set comprises a plurality of alarm sample data of a plurality of preset types;
[0017] According to the alarm frequency and alarm traffic of the alarm sample data, the time window and the sliding step of the correlation analysis model are determined, and the first weight corresponding to each of the alarm sample data and the second weight of the transaction corresponding to each of the time windows are determined;
[0018] Based on the time window, the sliding step, the first weight, the second weight and the correlation analysis model, the alarm sample data in each of the sample data subsets are analyzed in parallel for correlation to obtain a second correlation relationship of a preset type;
[0019] Based on the second correlation relationship, a preset transaction library is constructed.
[0020] In one of the embodiments, the anomaly detection hybrid model comprises an independent forest sub-model and a density clustering sub-model;
[0021] The anomaly detection hybrid model and the first correlation relationship corresponding to the alarm group are used to calculate the target anomaly score corresponding to each of the first alarm information, comprising:
[0022] Feature extraction is performed on the first alarm information to obtain a first feature matrix composed of the first alarm information, and the feature vectors in the first feature matrix are re-divided according to the alarm group to obtain a second feature matrix corresponding to each alarm group.
[0023] For each alarm group corresponding to the second feature matrix, based on the first association relationship in the second feature matrix and the independent forest sub-model, an anomaly score is calculated on the feature vector in the second feature matrix to obtain the first anomaly score corresponding to each first alarm information;
[0024] Based on the density clustering sub-model, anomaly scores are calculated on the feature vectors in the first feature matrix to obtain the second anomaly score corresponding to each alarm message.
[0025] The target anomaly score is determined based on the first anomaly score and the second anomaly score.
[0026] In one embodiment, before determining the target anomaly score based on the first anomaly score and the second anomaly score, the method further includes:
[0027] Obtain a preset transaction library; the preset transaction library contains alarm sample data and sample alarm groups formed by the second association between the alarm sample data;
[0028] Feature extraction is performed on the alarm sample data to obtain a first sample feature matrix composed of the alarm sample data. The sample feature vectors in the first sample feature matrix are then re-divided according to the sample alarm groups to obtain a second sample feature matrix corresponding to each sample alarm group.
[0029] Based on the first chi-square value calculated by the independent forest sub-model for the sample feature vectors and the second association relationship in the second sample feature matrix, and the second chi-square value calculated by the density clustering sub-model for each sample feature vector in the first sample feature matrix, the first weight corresponding to the independent forest sub-model and the second weight corresponding to the density clustering sub-model are determined respectively.
[0030] The step of determining the target anomaly score based on the first anomaly score and the second anomaly score includes:
[0031] The first abnormal score and the second abnormal score are weighted and summed according to the first weight and the second weight to obtain the target abnormal score.
[0032] In one embodiment, the step of performing root cause alarm identification on the second alarm information based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, and extracting the dependency relationship between the target alarm information and the target alarm information, includes:
[0033] An alarm graph is constructed based on the preset topology resource association information, the target anomaly score corresponding to the second alarm information, and the feature matrix corresponding to the second alarm information;
[0034] The alarm graph is analyzed and processed according to the probabilistic graphical model to obtain the target probability value corresponding to each second alarm information, and the target alarm information is determined based on the target probability value and the preset probability threshold.
[0035] The dependency relationships corresponding to the target alarm information are determined based on the target alarm information and the alarm graph.
[0036] Secondly, this application also provides a network detection device, comprising:
[0037] The first acquisition module is used to acquire preset topology resource association information, traffic information, and first alarm information of multiple alarm types of the target network;
[0038] The calculation module is used to calculate the target anomaly score corresponding to each of the first alarm messages, and to perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages;
[0039] The identification module is used to identify the root cause of the alarm based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, and to extract the target alarm information and the dependency relationship corresponding to the target alarm information.
[0040] The analysis module is used to perform root cause alarm analysis processing on the traffic information, the target alarm information and the dependency relationship according to the preset large language model, and obtain network detection results.
[0041] In one embodiment, the calculation module is specifically used to determine the first association relationship between the alarm types corresponding to each of the first alarm information based on the second association relationship between the alarm sample data in the preset transaction database;
[0042] The first alarm information is divided into multiple alarm groups based on the first association relationship;
[0043] The target anomaly score corresponding to each of the first alarm messages is calculated based on the anomaly detection hybrid model and the first association relationship corresponding to the alarm group.
[0044] In one embodiment, the device further includes:
[0045] The second acquisition module is used to acquire a sample dataset and divide the sample dataset into multiple sample data subsets; the sample dataset includes alarm sample data of multiple preset types;
[0046] The first determining module is used to determine the time window and sliding step of the correlation analysis model based on the alarm frequency and alarm traffic of the alarm sample data, and to determine the first weight corresponding to each alarm sample data and the second weight of the transaction corresponding to each time window.
[0047] The correlation analysis module is used to perform parallel correlation analysis on the alarm sample data in each of the sample data subsets based on the time window, the sliding step size, the first weight, the second weight and the correlation analysis model, to obtain a second correlation relationship of a preset type.
[0048] The construction module is used to construct a preset transaction library based on the second association relationship.
[0049] In one embodiment, the anomaly detection hybrid model includes an independent forest sub-model and a density clustering sub-model;
[0050] The calculation module is specifically used to extract features from the first alarm information to obtain a first feature matrix composed of the first alarm information, and to re-divide the feature vectors in the first feature matrix according to the alarm group to obtain a second feature matrix corresponding to each alarm group.
[0051] For each alarm group corresponding to the second feature matrix, based on the first association relationship in the second feature matrix and the independent forest sub-model, an anomaly score is calculated on the feature vector in the second feature matrix to obtain the first anomaly score corresponding to each first alarm information;
[0052] Based on the density clustering sub-model, anomaly scores are calculated on the feature vectors in the first feature matrix to obtain the second anomaly score corresponding to each alarm message.
[0053] The target anomaly score is determined based on the first anomaly score and the second anomaly score.
[0054] In one embodiment, the device further includes:
[0055] The third acquisition module is used to acquire a preset transaction library; the preset transaction library includes alarm sample data and sample alarm groups formed by the second association relationship between the alarm sample data;
[0056] The feature extraction module is used to extract features from the alarm sample data to obtain a first sample feature matrix composed of the alarm sample data, and to re-divide the sample feature vectors in the first sample feature matrix according to the sample alarm groups to obtain a second sample feature matrix corresponding to each sample alarm group.
[0057] The second determining module is used to determine the first weight corresponding to the independent forest sub-model and the second association relationship based on the first chi-square value calculated by the independent forest sub-model on the sample feature vectors and the second association relationship in the second sample feature matrix and the second chi-square value calculated by the density clustering sub-model on each of the sample feature vectors in the first sample feature matrix, respectively.
[0058] The calculation module is specifically used to perform a weighted summation of the first abnormal score and the second abnormal score according to the first weight and the second weight to obtain the target abnormal score.
[0059] In one embodiment, the identification module is specifically used to construct an alarm graph based on the preset topology resource association information, the target anomaly score corresponding to the second alarm information, and the feature matrix corresponding to the second alarm information;
[0060] The alarm graph is analyzed and processed according to the probabilistic graphical model to obtain the target probability value corresponding to each second alarm information, and the target alarm information is determined based on the target probability value and the preset probability threshold.
[0061] The dependency relationships corresponding to the target alarm information are determined based on the target alarm information and the alarm graph.
[0062] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0063] Obtain the target network's preset topology resource association information, traffic information, and the first alarm information for multiple alarm types;
[0064] Calculate the target anomaly score corresponding to each of the first alarm messages, and perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages;
[0065] Based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, the root cause alarm is identified for the second alarm information, and the target alarm information and the dependency relationship corresponding to the target alarm information are extracted.
[0066] Based on a preset large language model, the traffic information, the target alarm information, and the dependency relationship are subjected to root cause alarm analysis to obtain network detection results.
[0067] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0068] Obtain the target network's preset topology resource association information, traffic information, and the first alarm information for multiple alarm types;
[0069] Calculate the target anomaly score corresponding to each of the first alarm messages, and perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages;
[0070] Based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, the root cause alarm is identified for the second alarm information, and the target alarm information and the dependency relationship corresponding to the target alarm information are extracted.
[0071] Based on a preset large language model, the traffic information, the target alarm information, and the dependency relationship are subjected to root cause alarm analysis to obtain network detection results.
[0072] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0073] Obtain the target network's preset topology resource association information, traffic information, and the first alarm information for multiple alarm types;
[0074] Calculate the target anomaly score corresponding to each of the first alarm messages, and perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages;
[0075] Based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, the root cause alarm is identified for the second alarm information, and the target alarm information and the dependency relationship corresponding to the target alarm information are extracted.
[0076] Based on a preset large language model, the traffic information, the target alarm information, and the dependency relationship are subjected to root cause alarm analysis to obtain network detection results.
[0077] The aforementioned network detection method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire preset topology resource association information, and first alarm information according to preset periodic traffic information and multiple alarm types; calculate the target anomaly score corresponding to each first alarm information based on an anomaly detection hybrid model, and perform preliminary screening of the first alarm information based on the target anomaly scores to obtain second alarm information and the target anomaly scores corresponding to the second alarm information; determine the target alarm information and the corresponding dependencies in the second alarm information based on the preset topology resource association information, the target anomaly scores corresponding to the second alarm information, the feature matrix corresponding to the second alarm information, and the probabilistic graphical model; and analyze and process the traffic information, target alarm information, and dependencies according to a preset large language model to obtain object detection results. This method acquires preset topology resource association information, periodic traffic information, and first alarm information of various alarm types. It uses an anomaly detection hybrid model to calculate target anomaly scores and perform preliminary screening. Combined with probabilistic graphical model analysis, it obtains target alarm information and alarm dependencies. Then, through comprehensive analysis and processing using a large language model, it achieves accurate identification, efficient screening, and intelligent analysis of network alarms. This avoids human intervention, effectively reduces false alarms and false negatives, significantly improves the accuracy of identifying target alarm information as the root cause of alarms in network security management, and enhances the accuracy of network detection results. Attached Figure Description
[0078] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0079] Figure 1 This is a flowchart illustrating a network detection method in one embodiment;
[0080] Figure 2 This is a schematic diagram of a network detection architecture in one embodiment;
[0081] Figure 3 This is a flowchart illustrating the calculation of a target anomaly score in one embodiment;
[0082] Figure 4 This is a schematic diagram illustrating the process of constructing a preset transaction library in one embodiment;
[0083] Figure 5 This is a schematic diagram illustrating the process of obtaining the target anomaly score based on the independent forest sub-model and the density clustering sub-model in one embodiment;
[0084] Figure 6This is a flowchart illustrating the process of determining the first weight and the second weight in one embodiment;
[0085] Figure 7 This is a schematic diagram illustrating the process of constructing an alarm graph and identifying target alarm information based on the alarm graph in one embodiment;
[0086] Figure 8 This is a structural block diagram of a network detection device in one embodiment;
[0087] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0088] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0089] In one embodiment, such as Figure 1 As shown, a network detection method is provided. This embodiment illustrates the method applied to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0090] Step 102: Obtain the preset topology resource association information, traffic information, and first alarm information of multiple alarm types of the target network.
[0091] In this embodiment, for large-scale enterprise networks, data centers, or cloud platforms, the terminal can collect traffic data, Syslog logs, and preset topology resource association information according to a preset period or in real time. The Syslog logs of network devices include alarm information and device status; the traffic data of network devices includes traffic rate and traffic pattern; and the preset topology resource association information can be the connection relationships between devices and resources in the network, including the connection status of devices such as routers, switches, and servers, as well as the configuration of each subnet and VLAN. Then, as... Figure 2As shown, after the terminal completes data collection, it preprocesses the traffic data and Syslog logs. Specifically, the terminal extracts key information from the collected data to obtain traffic information of high value for network root cause detection and first alarm information for multiple alarm types, such as traffic rate, alarm type, alarm time, and IP (Internet Protocol) address. The first alarm information is the complete set of alarm information from the Syslog logs. It includes corresponding log information; for example, log 1 includes: Device: Router A (IP: 192.168.xx); Time: 2024-12-25 08:00:00; Alarm type: Interface traffic anomaly; Alarm description: Interface GigabitEthernet0 / 1 traffic rate exceeds the threshold (100 Mbps).
[0092] Step 104: Calculate the target anomaly score corresponding to each first alarm message, and perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages.
[0093] In this embodiment, the terminal assesses the anomaly level of the collected first alarm information and calculates a target anomaly score for each alarm information to quantify the anomaly level of each first alarm information. Specifically, the terminal can calculate the target anomaly score corresponding to each first alarm information according to a preset strategy. For example, it can perform a weighted calculation based on preset weights corresponding to the alarm type and alarm time of the first alarm information. For instance, different alarm types correspond to different preset scores, different alarm time periods have interval scores, and alarm type and alarm time have preset weights. Then, based on the scores corresponding to the alarm type and alarm time, and combined with the preset weights, the target anomaly score corresponding to each first alarm information is calculated. Alternatively, the terminal can also input the first alarm information into a neural network model, analyze and map each first alarm information through the neural network model, and output the target anomaly score for each first alarm information through the neural network model.
[0094] The terminal performs preliminary screening of first alarm information based on the calculated target anomaly score. The terminal may include a pre-set threshold; only first alarm information with a target anomaly score exceeding the threshold will be retained as second alarm information. This reduces the number of alarms processed subsequently, allowing the terminal to focus on alarms that are most likely to represent the real problem when performing root cause alarm analysis later.
[0095] Step 106: Based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, perform root cause alarm identification on the second alarm information, and extract the target alarm information and the dependency relationship corresponding to the target alarm information.
[0096] In this embodiment, the terminal combines preset topology resource association information with the target anomaly score of the second alarm information to identify the root cause alarm in the second alarm information, which is then used as the target alarm information; that is, the target alarm information is the root cause of other alarms. The terminal analyzes the connection relationships between network devices through the preset topology resource association information and the target anomaly score of the second alarm information, and further infers which second alarm information is the source of other alarms, thereby extracting the target alarm information and its dependencies. Specifically, the terminal can construct a graph structure based on the preset topology resource association information, the second alarm information, and their corresponding target anomaly scores, and analyze the graph structure using a graph neural network model to identify the target alarm information and its dependencies.
[0097] Step 108: Perform root cause alarm analysis on traffic information, target alarm information and dependency relationships according to the preset large language model to obtain network detection results.
[0098] In the embodiments of this application, such as Figure 2 As shown, the preset large language model can be a fine-tuned large model. The terminal integrates traffic information, target alarm information, and the dependencies between target alarm information. Through comprehensive analysis of the preset large language model, network detection results are obtained, which constitute specific hazard investigation solutions. Utilizing the language understanding and generation capabilities of the large model, combined with the analysis results of each module, the system can automatically generate specific hazard investigation solutions and optimization suggestions, and provide intelligent decision support to assist operation and maintenance personnel in quickly formulating effective solutions, improving the accuracy and efficiency of problem solving. Specifically, the hazard investigation solution is obtained through the analysis and calculation of the preset large language model. As shown in the formula below:
[0099]
[0100] Where F represents the characteristics of network traffic, logs, and preset topology resource association information, and R represents the target alarm information (root alarm) and its dependencies.
[0101] In the aforementioned network detection method, by acquiring preset topology resource association information, periodic traffic information, and first alarm information of various alarm types, the target anomaly score is calculated using an anomaly detection hybrid model and preliminary screening is performed. The target alarm information and alarm dependency relationship are obtained by combining probabilistic graphical model analysis, and then comprehensive analysis and processing are carried out through large language model. This achieves accurate identification, efficient screening, and intelligent analysis of network alarms, avoids human intervention, effectively reduces false alarms and false negatives, significantly improves the accuracy of identifying target alarm information as the root cause of alarms in network security management, and improves the accuracy of network detection results.
[0102] In one exemplary embodiment, such as Figure 3As shown, step 104 includes steps 302 to 306. Wherein:
[0103] Step 302: Based on the second correlation between alarm sample data in the preset transaction database, determine the first correlation between alarm types corresponding to each first alarm information.
[0104] In this embodiment, the preset transaction library is a pre-built database that serves as a knowledge graph in network detection. It includes the association relationships (second association relationships) between various alarm sample data. The transaction library, which is pre-built based on the alarm sample data and contains the second association relationships between various alarm types, aims to capture common combinations and dependencies between different alarm types. The specific construction process of the preset transaction library is described in detail in the following embodiments.
[0105] The terminal matches the alarm types of each first alarm information in a preset transaction database to obtain the first association relationship between the first alarm information. For example, if the second association relationship between alarm sample data in the preset transaction database is {A, B, C, D} and {C, E, F}, and the alarm types of the first alarm information include B, C, and F, then the first association relationship is obtained by matching the alarm types and the second association relationship as {B, C} and {C, F}.
[0106] Step 304: Divide the first alarm information into multiple alarm groups according to the first association relationship.
[0107] In this embodiment, the terminal groups related alarm information to facilitate more effective anomaly detection and analysis. By grouping related first alarm information together, their overall impact and potential root causes can be assessed more accurately. Specifically, based on the first correlation of the first alarm information, the terminal divides the first alarm information into different alarm groups, each containing mutually related alarm information. For example, if the first correlation of the first alarm information is {B, C} and {C, F}, the first alarm information of type B includes alarm 1 and alarm 2, the first alarm information of type C includes alarm 3, and the first alarm information of type F includes alarm 4 and alarm 5, then the divided alarm groups are alarm group 1 and alarm group 2. Alarm group 1 includes alarm 1, alarm 2, and alarm 3, and alarm group 2 includes alarm 3, alarm 4, and alarm 5.
[0108] Step 306: Calculate the target anomaly score corresponding to each first alarm information based on the anomaly detection hybrid model and the first correlation relationship corresponding to the alarm group.
[0109] In this embodiment, the anomaly detection hybrid model is a comprehensive model that can combine multiple anomaly detection algorithms to perform first correlation calculation on alarm groups, analyze the correlation between alarm types within each alarm group, and then, for each first alarm information, the terminal calculates a target anomaly score based on its alarm group and the correlation between alarm types within the group through the anomaly detection hybrid model.
[0110] In this embodiment, the first association between first alarm information is determined by the second association between alarm sample data in the preset transaction database, and the relevant alarm information is divided into multiple alarm groups. The target anomaly score is calculated by combining the anomaly detection hybrid model, thereby realizing accurate grouping and anomaly detection of alarm information. It can effectively capture common combinations and dependencies between alarm types, improve the accuracy and efficiency of alarm analysis, and at the same time, by dividing alarm groups and calculating anomaly scores, it can assist in the identification and location of potential root causes (target alarm information) and improve the accuracy of target alarm information identification, thereby improving the accuracy of network detection results.
[0111] In one exemplary embodiment, such as Figure 4 As shown, before step 302, the method further includes steps 402 to 408. Wherein:
[0112] Step 402: Obtain the sample dataset and divide the sample dataset into multiple sample data subsets.
[0113] The sample dataset includes alarm sample data of multiple preset types.
[0114] In this embodiment, the terminal can collect a sample dataset from historical alarm records. The sample dataset covers various types of alarm information, such as device failure, network congestion, and security threats. Then, the terminal divides the sample dataset into multiple data subsets to achieve parallel processing of multiple sample data subsets.
[0115] Step 404: Determine the time window and sliding step of the correlation analysis model based on the alarm frequency and alarm traffic of the alarm sample data, and determine the first weight corresponding to each alarm sample data and the second weight of each transaction corresponding to each time window.
[0116] In this embodiment of the application, the time window size is dynamically adjusted. and adaptive sliding step size The terminal optimizes the construction of the transaction database based on real-time network traffic and alarm frequency to ensure the timeliness and flexibility of alarm correlation analysis. The time window refers to the time range for analyzing alarm data. For example, the terminal can define the trial lecture window as 10 minutes, 30 minutes, or 1 hour. The adaptive sliding step size refers to the step size used to move between time windows. For example, if the time window is 30 minutes, the sliding step size could be 5 minutes, ensuring overlap between each time window and improving the accuracy of the analysis. Time window size. and adaptive sliding step size The calculation formula is as follows:
[0117]
[0118]
[0119] The size of the time window determines the granularity of the analysis. A larger time window can capture patterns over a longer period, but may result in a loss of detail; a smaller time window is better suited for capturing short-term, localized anomalies. The adaptive sliding step size can be dynamically adjusted based on the characteristics of the alarm data. For example, if there are few alarm events within the current time window (e.g., during periods of low traffic or low activity), the adaptive sliding step size can be appropriately reduced to capture more detail; if there are many alarm events, the adaptive sliding step size can be appropriately increased to improve analysis efficiency.
[0120] Step 406: Based on the time window, sliding step size, first weight, second weight and correlation analysis model, perform parallel correlation analysis on the alarm sample data in each sample data subset to obtain the second correlation relationship of the preset type.
[0121] In this embodiment, the correlation analysis model can be a model that analyzes alarm sample data in a subset of sample data based on the Apriori algorithm, and the first weight is the weight corresponding to the alarm type to which the alarm sample data belongs. The second weight is the weight of each alarm sample data constituting a transaction within the time window. The time window, sliding step size, first weight, and second weight can be set by operations and maintenance personnel based on prior experience, or obtained from the neural network model used for weight mapping.
[0122] The terminal uses an association analysis model to perform parallel association analysis on subsets of sample data according to pre-set time windows, sliding steps, first weights, and second weights. Specifically, the terminal calls upon the GPU (Graphics Processing Unit) to perform association analysis on each subset of sample data simultaneously, introducing weighted support calculation and multi-level association rule mining, by assigning weights to different types of alarm sample data. The importance of alarm sample data is considered in the support calculation, and this is extended to multi-level association rules to ensure that high-weight alarms have a greater influence in association rule mining, thereby obtaining alarm sample data with strong correlations, which serve as the second association between these alarm sample data. Support is... and confidence level The calculation formula is as follows:
[0123]
[0124] In an exemplary embodiment, the terminal employs parallel computing and the incremental Apriori method to improve the computational efficiency of frequent itemset mining. This is achieved by dividing the sample dataset, composed of alarm sample data, into multiple subsets for parallel processing during the transaction database construction process. Furthermore, the preset transaction database is dynamically updated when the network architecture changes. During the update process, frequent itemset updates are only performed on newly added transactions, reducing computational complexity and time overhead. The formulas for performing correlation analysis on the original sample dataset and on newly added transactions are shown below:
[0125]
[0126] in, This represents the k-th subset in the preset transaction database, where K is the number of subsets. Represents newly added transactions. In incremental transaction processing, New Frequent Itemsets are applied to new transactions using the Apriori algorithm. It is calculated based on incremental transactions. The terminal uses the Apriori algorithm to discover frequent itemsets, denoted as New Frequent Itemsets.
[0127] Meanwhile, the frequent itemsets of existing transactions are obtained by analyzing the first K transactions. The frequent itemsets are merged to obtain the following: .
[0128] When processing incremental transactions, New Frequent Itemsets are added to the frequent itemsets of existing transactions, thereby updating the overall frequent itemsets. Specific update methods may include simple merging, or using strategies to handle changes to itemsets and the deletion of outdated itemsets.
[0129] Step 408: Construct a preset transaction library based on the second association relationship.
[0130] In this embodiment, the terminal stores the second association relationship corresponding to the alarm type to which the alarm sample data belongs in the database to obtain a preset transaction database. ,in, Size of the included time window Transactions involving internal alarm sample data. and For phase distance adaptive sliding step size Two transactions. And it serves as the knowledge graph for the hybrid anomaly detection model and the pre-defined large language model in network detection.
[0131] In this embodiment, firstly, by integrating a pre-set large language model, knowledge graph, and traditional artificial intelligence model into a system construction method, the advantages of various models can be fully integrated to achieve efficient fusion and collaborative analysis of multi-source data, thereby improving the overall performance and intelligence level of the system.
[0132] Secondly, by acquiring historical alarm records and dividing them into multiple data subsets, dynamically adjusting the time window and sliding step size, and combining weight allocation and weighted support calculation, this method utilizes parallel computing and the incremental Apriori algorithm to perform efficient correlation analysis on the alarm sample data, ultimately constructing a pre-defined transaction database. This approach improves the timeliness, flexibility, and accuracy of alarm correlation analysis. Simultaneously, through dynamic updates and incremental processing, it reduces computational complexity and time overhead, ultimately providing strong knowledge support for anomaly detection in network monitoring and improving the efficiency and accuracy of network alarm root cause identification and anomaly detection.
[0133] In one exemplary embodiment, the anomaly detection hybrid model includes an independent forest sub-model and a density clustering sub-model; such as Figure 5 As shown, step 306 includes steps 502 to 508. Wherein:
[0134] Step 502: Extract features from the first alarm information to obtain a first feature matrix composed of the first alarm information, and re-divide the feature vectors in the first feature matrix according to the alarm group to obtain a second feature matrix corresponding to each alarm group.
[0135] In this embodiment, the terminal can extract features from the first alarm information using a word embedding model, converting the first alarm information into a numerical feature vector, and combining the feature vectors corresponding to all the first alarm information into a large feature matrix, which serves as the first feature matrix. Simultaneously, the terminal re-divides the feature vectors in the first feature matrix according to alarm groups defined in a preset transaction database, so that each alarm group corresponds to a second feature vector. This allows the terminal to perform independent anomaly detection and analysis for each alarm group.
[0136] Step 504: For the second feature matrix corresponding to each alarm group, calculate the anomaly score of the feature vector in the second feature matrix according to the first association relationship and independent forest sub-model in the second feature matrix, and obtain the first anomaly score corresponding to each first alarm information.
[0137] In this embodiment, the terminal processes the second feature matrix corresponding to each alarm group separately, inputting the first alarm information in the alarm group and the second association relationship to which each alarm information in the alarm group belongs to the independent forest sub-model (a sub-model based on the Isolation Forest algorithm). The independent forest sub-model uses the feature vector corresponding to the first alarm information as the node (data point) of the independent forest, and determines the independent forest structure of each node in the alarm group according to the second association relationship, and then calculates the first anomaly score of each feature vector in the second feature matrix. First abnormal score The calculation formula is as follows:
[0138]
[0139] in, For data points Average path length, To the size of the sample The relevant expected path length, To adapt to tree depth.
[0140] Step 506: Calculate the anomaly score of the feature vectors in the first feature matrix based on the density clustering sub-model to obtain the second anomaly score corresponding to each alarm message.
[0141] In this embodiment, the density clustering sub-model can be a clustering model based on the DBSCAN algorithm. The terminal calculates anomaly scores for each feature vector in the first feature matrix according to the DBSCAN algorithm. Specifically, the terminal uses each first alarm message as a cluster center and analyzes the correlation between the first alarm message as a cluster center and other alarm messages. Then, areas with higher density are designated as normal clusters, and nodes with lower density are designated as abnormal nodes. A second anomaly score is calculated for the feature vector corresponding to each alarm message. The core clustering algorithm is shown below:
[0142]
[0143] Where ϵ represents the density condition and MinPts represents the minimum number of points. That is, the second anomaly score of a core point can be the minimum value among the base scores. The anomaly score of boundary points in the clustering results can be calculated based on the proximity of their neighbor counts to the minimum number of points MinPts. The second anomaly score of noise points is set to the maximum value among the base scores, indicating that the point is a significant anomaly. Optionally, the terminal can also adjust the second anomaly score in conjunction with the first correlation. For example, if a first alarm message has a strong correlation with other alarm messages (i.e., they often appear simultaneously), its anomaly score can be reduced; if a first alarm message has a weak correlation with other alarm messages, its anomaly score can be increased, because this means it is more likely to be an independent anomaly. That is, second anomaly score = base anomaly score × (1 + correlation adjustment factor).
[0144] Step 508: Determine the target abnormal score based on the first abnormal score and the second abnormal score.
[0145] In this embodiment of the application, the terminal may ultimately use the sum of the first abnormal score and the second abnormal score as the target abnormal score, or the terminal may also perform a weighted summation of the first score and the second abnormal score to obtain the target abnormal score.
[0146] In this embodiment, by combining the target anomaly score calculation methods of the independent forest sub-model and the density clustering sub-model, multi-dimensional anomaly detection is performed on the alarm information. This not only considers isolated points in the feature space but also density differences, thereby providing more comprehensive and accurate anomaly detection results. Furthermore, by adjusting the weights, the performance of network anomaly detection can be further optimized, improving the accuracy of identifying root cause alarms (target alarm information) in complex network environments.
[0147] Furthermore, by combining independent forest sub-models and density clustering sub-models, the real-time processing capabilities for alarm correlation analysis and hazard detection are significantly improved. This enables the terminal to analyze and respond instantly upon data generation, achieving real-time hazard detection, accurate location, and rapid troubleshooting, greatly shortening problem response time and improving operational efficiency.
[0148] In an exemplary embodiment, the first weight and the second weight can be obtained in advance based on alarm sample data contained in a preset transaction database, such as... Figure 6 As shown, before step 508, the method further includes steps 602 to 606. Wherein:
[0149] Step 602: Obtain the preset transaction library.
[0150] The preset transaction library contains alarm sample data and sample alarm groups formed by the second association between alarm sample data.
[0151] In this embodiment, the preset transaction library is constructed from historical alarm data through correlation analysis. It includes various alarm sample data and their second correlation relationships, as well as different sample alarm groups composed of these correlation relationships. The terminal first loads the preset transaction library to obtain each alarm group and various alarm sample data and their second correlation relationships in the preset transaction library.
[0152] Step 604: Extract features from the alarm sample data to obtain the first sample feature matrix composed of the alarm sample data, and re-divide the sample feature vectors in the first sample feature matrix according to the sample alarm group to obtain the second sample feature matrix corresponding to each sample alarm group.
[0153] In this embodiment, the terminal constructs the first sample feature matrix and the second sample feature matrix according to the same principle as step 502. The construction process of the first sample feature matrix and the second sample feature matrix will not be described in detail in this embodiment.
[0154] Step 606: Based on the first chi-square value calculated by the independent forest sub-model for the sample feature vectors and the second association relationship in the second sample feature matrix, and the second chi-square value calculated by the density clustering sub-model for the sample feature vectors in the first sample feature matrix, determine the first weight corresponding to the independent forest sub-model and the second weight corresponding to the density clustering sub-model, respectively.
[0155] In this embodiment, the terminal, following the same principle as steps 504 and 506, evaluates the performance of the independent forest sub-model in calculating the second sample feature matrix using the chi-square test to obtain a first chi-square value. Simultaneously, it evaluates the performance of the density clustering sub-model in processing the first sample feature matrix to obtain a second chi-square value. Then, it calculates a first weight based on the first and second chi-square values. The formula for calculating the first weight is as follows:
[0156]
[0157] in, This is the first chi-square value calculated for the feature matrix of the second sample in the independent forest sub-model. This is the second chi-square value when the density clustering sub-model processes the feature matrix of the first sample. Furthermore, the second weight is... .
[0158] Step 508 includes step 5081, wherein:
[0159] Step 5081: The first abnormal score and the second abnormal score are weighted and summed according to the first weight and the second weight to obtain the target abnormal score.
[0160] In this embodiment of the application, the terminal ultimately generates a comprehensive anomaly score by weighted fusion of the first anomaly score and the second anomaly score:
[0161]
[0162] Where α is the weighting coefficient. It is the chi-square value that reflects the reliability of the hybrid anomaly detection model across multiple tasks.
[0163] In this embodiment, by combining the advantages of traditional AI models such as independent forest sub-models and density clustering sub-models, the incompatibility issues between large language models and traditional AI models in terms of data format and accuracy are resolved, enabling collaborative work between different types of models. This compatibility improves resource utilization and flexibility in network detection, allowing for dynamic adjustment of model combinations based on actual needs, and optimizing the configuration and utilization efficiency of operational resources.
[0164] Simultaneously, by dynamically adjusting weights using the chi-square test, weighted fusion analysis is performed on alarm sample data to generate a comprehensive anomaly score. This improves the accuracy and reliability of anomaly detection, effectively identifying isolated anomalies and capturing density differences. This provides support for network fault diagnosis and anomaly detection, achieving deep compatibility between a pre-set large language model and a knowledge graph, ensuring seamless data transmission and sharing between different models. Through this compatibility, the terminal can provide more accurate and intelligent network vulnerability detection and analysis, while ensuring data security and integrity, avoiding information leakage or misprocessing due to data incompatibility.
[0165] In one exemplary embodiment, such as Figure 7 As shown, step 106 includes steps 702 to 706. Wherein:
[0166] Step 702: Construct an alarm graph based on the preset topology resource association information, the target anomaly score corresponding to the second alarm information, and the feature matrix corresponding to the second alarm information.
[0167] The feature matrix corresponding to the second alarm information includes the feature vector of each second alarm information.
[0168] In this embodiment, the terminal obtains the feature vector of each second alarm information from the feature matrix corresponding to the second alarm information, and concatenates the target anomaly score corresponding to the second alarm information with the feature vector of the second alarm information to obtain a concatenated vector. The terminal uses this concatenated vector as a node of the alarm graph, and determines the relationship between each second alarm information according to the preset topology resource association information, which is used as the edge of the alarm graph, thereby obtaining the alarm graph G=(V,E). Specifically, based on the physical connection, dependency relationship or logical association between topology resource association information, such as the network connection between network devices, the calling relationship between services, etc., the terminal can identify the direct or indirect association between the second alarm information, and then abstract the association relationship as the edge in the alarm graph.
[0169] Step 704: Analyze and process the alarm graph according to the probabilistic graph model to obtain the target probability value corresponding to each second alarm information, and determine the target alarm information based on the target probability value and the preset probability threshold.
[0170] The probabilistic graphical model is a dynamic Bayesian network.
[0171] In this embodiment of the application, the terminal uses graph embedding technology to transform the alarm graph G=(V,E) into a low-dimensional vector representation Embedding. A dynamic Bayesian network (BN) is constructed to capture the causal relationships between second alarm information through a probabilistic graphical model.
[0172]
[0173] in, This represents the number of nodes V in the alarm graph. Let V be the node in the alarm graph. The target probability value is used to characterize each second alarm message as a target alarm message. Then, the terminal compares the target probability value with a preset probability threshold, and determines the second alarm message with a target probability value higher than the preset probability threshold as the target alarm message.
[0174] For the training process of dynamic Bayesian network (BN), the terminal can employ dynamic structure learning based on a second-order optimization method to update the structural parameters θ of the Bayesian network in real time, where η is the learning rate. The Hessian matrix of the loss function:
[0175]
[0176] Through causal inference using Bayesian networks, it is possible to identify and prioritize the processing of target alarm information (Root Cause) that serves as the root alarm. This allows for the optimization of network security response strategies.
[0177]
[0178] in, Characterization in given evidence In this case, the probability that the second alarm message is the target alarm message. This includes the second alarm information, the dependencies related to the second alarm information in the preset topology resource association information, and the target anomaly score.
[0179] Step 706: Determine the dependencies corresponding to the target alarm information based on the target alarm information and the alarm graph.
[0180] In this embodiment, the terminal determines the node corresponding to the target alarm information in the alarm graph. By traversing the alarm graph, it finds the edges directly connected to the node of the target alarm information. The edges represent direct dependencies. The dependencies are extracted from these edges, including the direction of the dependency, the strength of the dependency, or other attributes. Optionally, the terminal can further traverse the nodes indirectly connected to the node of the target alarm information to obtain deeper dependencies. Finally, the terminal obtains the dependencies corresponding to the target alarm information and organizes the extracted dependencies, including operations such as classification, sorting, or calculating relevance, to facilitate subsequent analysis and processing.
[0181] In this embodiment, the feature vector of the second alarm information and the target anomaly score are concatenated to form nodes of the alarm graph, which can comprehensively capture the multi-dimensional features of the alarm information and provide richer evidence for anomaly detection. By pre-setting topological resource association information, the edges in the alarm graph can be constructed, which can reflect the physical connections, dependencies, or logical associations between alarm information, and help to identify the causal relationships between alarm information. By using dynamic Bayesian networks for probabilistic inference, the dynamic causal relationships between alarm information can be captured. By updating the network structure parameters in real time, the network environment can be adapted to changes, improving the accuracy of the probabilistic graphical model in identifying target alarm information, thereby improving the accuracy of network detection results.
[0182] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0183] Based on the same inventive concept, this application also provides a network detection apparatus for implementing the network detection method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more network detection apparatus embodiments provided below can be found in the limitations of the network detection method described above, and will not be repeated here.
[0184] In one exemplary embodiment, such as Figure 8 As shown, a network detection device 800 is provided, including: a first acquisition module 801, a calculation module 802, an identification module 803, and an analysis module 804, wherein:
[0185] The first acquisition module 801 is used to acquire preset topology resource association information, traffic information, and first alarm information of multiple alarm types of the target network;
[0186] The calculation module 802 is used to calculate the target anomaly score corresponding to each first alarm information, and to perform preliminary screening of the first alarm information based on the target anomaly score to obtain the second alarm information;
[0187] The identification module 803 is used to identify the root cause alarm of the second alarm information based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, and extract the target alarm information and the dependency relationship corresponding to the target alarm information.
[0188] The analysis module 804 is used to perform root cause alarm analysis on traffic information, target alarm information and dependency relationships based on a preset large language model to obtain network detection results.
[0189] In one embodiment, the calculation module 802 is specifically used to determine the first association relationship between the alarm types corresponding to each first alarm information based on the second association relationship between the alarm sample data in the preset transaction database.
[0190] The first alarm information is divided into multiple alarm groups based on the first association relationship;
[0191] The target anomaly score is calculated based on the anomaly detection hybrid model and the first correlation relationship corresponding to the alarm group.
[0192] In one embodiment, the device 800 further includes:
[0193] The second acquisition module is used to acquire the sample dataset and divide the sample dataset into multiple sample data subsets; the sample dataset includes alarm sample data of multiple preset types;
[0194] The first determination module is used to determine the time window and sliding step of the correlation analysis model based on the alarm frequency and alarm traffic of the alarm sample data, and to determine the first weight corresponding to each alarm sample data and the second weight of each transaction corresponding to each time window.
[0195] The correlation analysis module is used to perform parallel correlation analysis on alarm sample data in each sample data subset based on time window, sliding step size, first weight, second weight and correlation analysis model, and obtain the second correlation relationship of preset type;
[0196] The building module is used to build a preset transaction library based on the second association relationship.
[0197] In one embodiment, the anomaly detection hybrid model includes an independent forest sub-model and a density clustering sub-model;
[0198] The calculation module 802 is specifically used to extract features from the first alarm information to obtain a first feature matrix composed of the first alarm information, and to re-divide the feature vectors in the first feature matrix according to the alarm group to obtain a second feature matrix corresponding to each alarm group.
[0199] For each alarm group, the anomaly score is calculated on the feature vectors in the second feature matrix based on the first correlation relationship and the independent forest sub-model, so as to obtain the first anomaly score corresponding to each first alarm information.
[0200] Based on the density clustering sub-model, anomaly scores are calculated on the feature vectors in the first feature matrix to obtain the second anomaly score corresponding to each alarm message.
[0201] The target anomaly score is determined based on the first and second anomaly scores.
[0202] In one embodiment, the device 800 further includes:
[0203] The third acquisition module is used to acquire a preset transaction library; the preset transaction library contains alarm sample data and sample alarm groups formed by the second association between alarm sample data;
[0204] The feature extraction module is used to extract features from the alarm sample data to obtain the first sample feature matrix composed of the alarm sample data, and to re-divide the sample feature vectors in the first sample feature matrix according to the sample alarm group to obtain the second sample feature matrix corresponding to each sample alarm group.
[0205] The second determining module is used to determine the first weight corresponding to the independent forest sub-model and the second weight corresponding to the density clustering sub-model based on the first chi-square value calculated by the independent forest sub-model on the sample feature vectors and the second correlation relationship in the second sample feature matrix and the second chi-square value calculated by the density clustering sub-model on the sample feature vectors in the first sample feature matrix, respectively.
[0206] The calculation module 802 is specifically used to perform a weighted summation of the first abnormal score and the second abnormal score according to the first weight and the second weight to obtain the target abnormal score.
[0207] In one embodiment, the identification module 803 is specifically used to construct an alarm map based on preset topology resource association information, the target anomaly score corresponding to the second alarm information, and the feature matrix corresponding to the second alarm information;
[0208] The alarm graph is analyzed and processed according to the probabilistic graphical model to obtain the target probability value corresponding to each second alarm information. The target alarm information is determined based on the target probability value and the preset probability threshold.
[0209] Determine the dependencies corresponding to the target alarm information based on the target alarm information and the alarm graph.
[0210] Each module in the aforementioned network detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0211] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data from a pre-defined transaction database. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a network detection method.
[0212] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0213] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0214] Obtain the target network's preset topology resource association information, traffic information, and the first alarm information for multiple alarm types;
[0215] Calculate the target anomaly score corresponding to each first alarm message, and perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages;
[0216] Based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, the root cause alarm is identified for the second alarm information, and the target alarm information and the corresponding dependency relationship are extracted.
[0217] Based on the pre-set large language model, root cause alarm analysis is performed on traffic information, target alarm information and dependency relationships to obtain network detection results.
[0218] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0219] Based on the second correlation between alarm sample data in the preset transaction database, determine the first correlation between alarm types corresponding to each first alarm information;
[0220] The first alarm information is divided into multiple alarm groups based on the first association relationship;
[0221] The target anomaly score is calculated based on the anomaly detection hybrid model and the first correlation relationship corresponding to the alarm group.
[0222] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0223] Obtain the sample dataset and divide it into multiple sample data subsets; the sample dataset includes alarm sample data of multiple preset types;
[0224] The time window and sliding step size of the correlation analysis model are determined based on the alarm frequency and alarm traffic of the alarm sample data, and the first weight corresponding to each alarm sample data and the second weight corresponding to each time window are determined.
[0225] Based on time window, sliding step size, first weight, second weight and correlation analysis model, parallel correlation analysis is performed on alarm sample data in each sample data subset to obtain the second correlation relationship of the preset type;
[0226] A pre-defined transaction library is constructed based on the second association relationship.
[0227] In one embodiment, the anomaly detection hybrid model includes an independent forest sub-model and a density clustering sub-model; the processor also performs the following steps when executing the computer program:
[0228] Feature extraction is performed on the first alarm information to obtain the first feature matrix composed of the first alarm information. The feature vectors in the first feature matrix are then re-divided according to the alarm group to obtain the second feature matrix corresponding to each alarm group.
[0229] For each alarm group, the anomaly score is calculated on the feature vectors in the second feature matrix based on the first correlation relationship and the independent forest sub-model, so as to obtain the first anomaly score corresponding to each first alarm information.
[0230] Based on the density clustering sub-model, anomaly scores are calculated on the feature vectors in the first feature matrix to obtain the second anomaly score corresponding to each alarm message.
[0231] The target anomaly score is determined based on the first and second anomaly scores.
[0232] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0233] Obtain the preset transaction library; the preset transaction library contains alarm sample data and sample alarm groups formed by the second association between alarm sample data;
[0234] Feature extraction is performed on the alarm sample data to obtain the first sample feature matrix composed of the alarm sample data. The sample feature vectors in the first sample feature matrix are then re-divided according to the sample alarm groups to obtain the second sample feature matrix corresponding to each sample alarm group.
[0235] Based on the first chi-square value calculated by the independent forest sub-model on the sample feature vectors and the second association relationship in the second sample feature matrix, and the second chi-square value calculated by the density clustering sub-model on the sample feature vectors in the first sample feature matrix, the first weight corresponding to the independent forest sub-model and the second weight corresponding to the density clustering sub-model are determined respectively.
[0236] The target anomaly score is determined based on the first anomaly score and the second anomaly score, including:
[0237] The first and second outlier scores are weighted and summed according to the first and second weights to obtain the target outlier score.
[0238] In one embodiment, the processor, when executing a computer program, also performs the following steps:
[0239] An alarm graph is constructed based on the preset topology resource association information, the target anomaly score corresponding to the second alarm information, and the feature matrix corresponding to the second alarm information.
[0240] The alarm graph is analyzed and processed according to the probabilistic graphical model to obtain the target probability value corresponding to each second alarm information. The target alarm information is determined based on the target probability value and the preset probability threshold.
[0241] Determine the dependencies corresponding to the target alarm information based on the target alarm information and the alarm graph.
[0242] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0243] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0244] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0245] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0246] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0247] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A network detection method, characterized in that, The method includes: Obtain the target network's preset topology resource association information, traffic information, preset transaction library, and first alarm information of multiple alarm types; the preset transaction library contains sample alarm groups composed of alarm sample data and second association relationships between the alarm sample data; Feature extraction is performed on the alarm sample data, and the sample feature vectors in the first sample feature matrix obtained by feature extraction are re-divided according to the sample alarm group to obtain the second sample feature matrix of each sample alarm group. Based on the first chi-square value calculated by the independent forest sub-model for the sample feature vectors in the second sample feature matrix and the second association relationship, and the second chi-square value calculated by the density clustering sub-model for each sample feature vector in the first sample feature matrix, the first weight corresponding to the independent forest sub-model and the second weight corresponding to the density clustering sub-model are determined. The first anomaly score and the second anomaly score are weighted and summed according to the first weight and the second weight to obtain the target anomaly score; the first anomaly score is calculated by the independent forest sub-model of the anomaly detection hybrid model and the first alarm information; the second anomaly score is calculated by the density clustering sub-model of the anomaly detection hybrid model and the first alarm information. The first alarm information is initially filtered based on the target anomaly score to obtain the second alarm information; Based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, the root cause alarm is identified for the second alarm information, and the target alarm information and the dependency relationship corresponding to the target alarm information are extracted. Based on a preset large language model, the traffic information, the target alarm information, and the dependency relationship are subjected to root cause alarm analysis to obtain network detection results.
2. The method according to claim 1, characterized in that, The calculation of the target anomaly score corresponding to each of the first alarm messages includes: Based on the second correlation between alarm sample data in the preset transaction database, determine the first correlation between the alarm types corresponding to each of the first alarm information; The first alarm information is divided into multiple alarm groups based on the first association relationship; The target anomaly score corresponding to each of the first alarm messages is calculated based on the anomaly detection hybrid model and the first association relationship corresponding to the alarm group.
3. The method according to claim 2, characterized in that, Before determining the first association relationship between the alarm types corresponding to each first alarm information based on the first association relationship between alarm sample data in the preset transaction database, the method further includes: Obtain a sample dataset and divide the sample dataset into multiple sample data subsets; the sample dataset includes alarm sample data of multiple preset types; The time window and sliding step size of the correlation analysis model are determined based on the alarm frequency and alarm traffic of the alarm sample data, and the first weight corresponding to each alarm sample data and the second weight corresponding to each time window are determined. Based on the time window, the sliding step size, the first weight, the second weight, and the correlation analysis model, parallel correlation analysis is performed on the alarm sample data in each of the sample data subsets to obtain a second correlation relationship of a preset type. A preset transaction database is constructed based on the second association relationship.
4. The method according to claim 2, characterized in that, The anomaly detection hybrid model includes an independent forest sub-model and a density clustering sub-model; The calculation of the target anomaly score corresponding to each of the first alarm messages based on the anomaly detection hybrid model and the first association relationship corresponding to the alarm group includes: Feature extraction is performed on the first alarm information to obtain a first feature matrix composed of the first alarm information, and the feature vectors in the first feature matrix are re-divided according to the alarm group to obtain a second feature matrix corresponding to each alarm group. For each alarm group corresponding to the second feature matrix, based on the first association relationship in the second feature matrix and the independent forest sub-model, an anomaly score is calculated on the feature vector in the second feature matrix to obtain the first anomaly score corresponding to each first alarm information; Based on the density clustering sub-model, anomaly scores are calculated on the feature vectors in the first feature matrix to obtain the second anomaly score corresponding to each alarm message. The target anomaly score is determined based on the first anomaly score and the second anomaly score.
5. The method according to claim 1, characterized in that, The step of identifying the root cause of the second alarm information based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, and extracting the dependency relationship between the target alarm information and the target alarm information, includes: An alarm graph is constructed based on the preset topology resource association information, the target anomaly score corresponding to the second alarm information, and the feature matrix corresponding to the second alarm information; The alarm graph is analyzed and processed according to the probabilistic graphical model to obtain the target probability value corresponding to each second alarm information, and the target alarm information is determined based on the target probability value and the preset probability threshold. The dependency relationships corresponding to the target alarm information are determined based on the target alarm information and the alarm graph.
6. A network detection device, characterized in that, The device includes: The first acquisition module is used to acquire preset topology resource association information, traffic information, and first alarm information of multiple alarm types of the target network; The calculation module is used to calculate the target anomaly score corresponding to each of the first alarm messages, and to perform preliminary screening of the first alarm messages based on the target anomaly scores to obtain the second alarm messages; The identification module is used to identify the root cause of the alarm based on the preset topology resource association information and the target anomaly score corresponding to the second alarm information, and to extract the target alarm information and the dependency relationship corresponding to the target alarm information. The analysis module is used to perform root cause alarm analysis on the traffic information, the target alarm information and the dependency relationship according to the preset large language model, and obtain network detection results; The device further includes: The third acquisition module is used to acquire a preset transaction library; the preset transaction library includes alarm sample data and sample alarm groups formed by the second association relationship between the alarm sample data; The feature extraction module is used to extract features from the alarm sample data and re-divide the sample feature vectors in the first sample feature matrix obtained by feature extraction according to the sample alarm group to obtain the second sample feature matrix corresponding to each sample alarm group. The second determining module is used to determine the first weight corresponding to the independent forest sub-model and the second association relationship based on the first chi-square value calculated by the independent forest sub-model on the sample feature vectors and the second association relationship in the second sample feature matrix and the second chi-square value calculated by the density clustering sub-model on each of the sample feature vectors in the first sample feature matrix, respectively. The calculation module is specifically used to perform a weighted summation of the first anomaly score and the second anomaly score based on the first weight and the second weight to obtain the target anomaly score; wherein, the first anomaly score is calculated based on the independent forest sub-model of the anomaly detection hybrid model and the first alarm information; the second anomaly score is calculated based on the density clustering sub-model of the anomaly detection hybrid model and the first alarm information.
7. The apparatus according to claim 6, characterized in that, The device further includes: The second acquisition module is used to acquire a sample dataset and divide the sample dataset into multiple sample data subsets; the sample dataset includes alarm sample data of multiple preset types; The first determining module is used to determine the time window and sliding step of the correlation analysis model based on the alarm frequency and alarm traffic of the alarm sample data, and to determine the first weight corresponding to each alarm sample data and the second weight of the transaction corresponding to each time window. The correlation analysis module is used to perform parallel correlation analysis on the alarm sample data in each of the sample data subsets based on the time window, the sliding step size, the first weight, the second weight and the correlation analysis model, to obtain a second correlation relationship of a preset type. The construction module is used to construct a preset transaction library based on the second association relationship.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Alarm root cause analysis method, electronic equipment and storage medium
CN112087334A
Root cause locating method, electronic device, and storage medium
WO2022237088A1