Intelligent Identification Method and Device for System Hotspot Problems
By unifying the timestamp format and causal inference model of multi-source data, a knowledge graph is built to identify hot issues in the system, solving the problem of traditional methods relying on manual intervention and low data processing efficiency, and realizing intelligent, rapid identification and early warning of system problems.
Patent Information
- Application Number
- CN202510560353.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional system problem identification and analysis methods rely on manual intervention, which is time-consuming and labor-intensive, and it is difficult to quickly and effectively process massive data. The lack of intelligent support makes it difficult for operation and maintenance personnel to understand the system status and problems in a timely manner, affecting the timely resolution of problems.
By unifying the timestamp format of multi-source data in the historical system, comparing it on the same time axis, using the association rule algorithm and causal inference model to mine events related to extremely strong system performance in the system multi-source data, building a knowledge graph and performing causal inference, forming a system hot spot recognition model, and receiving real-time data for abnormal pattern recognition and dynamic adjustment.
It improves the efficiency and accuracy of system hot issues identification, can promptly detect and warn of potential problems, reduce human intervention, and improve system operation efficiency.
Smart Images

Figure CN120069045B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and specifically to an intelligent identification method and device for system hot issues. Background Art
[0002] With the rapid development of information technology, modern complex systems (such as large software systems, cloud computing platforms, Internet of Things systems, etc.) generate a vast amount of data during daily operations. This data covers multiple aspects such as system logs, performance monitoring metrics, network traffic information, sensor data, etc., which is crucial for the stable operation and performance optimization of the system. However, traditional system problem identification and analysis methods usually have the following limitations:
[0003] Relying on manual intervention: Traditional system monitoring and analysis often require manual participation, and potential problems are identified by viewing logs, analyzing performance metrics, etc. This method not only consumes time and effort but is also easily affected by human factors, making it difficult to guarantee the accuracy and timeliness of the analysis results.
[0004] Low data processing efficiency: Facing a vast amount of system data, traditional processing methods often have difficulty in quickly and effectively processing and analyzing the data. Especially in scenarios with high real-time requirements, traditional analysis methods may not be able to respond in a timely manner to changes in the system state, thus missing the best opportunity to discover problems.
[0005] Lack of intelligent support: Traditional system monitoring and analysis methods often lack intelligent support and are difficult to automatically mine potential problems and trends from complex data. This results in the system often being able to only respond passively when problems occur, rather than being able to give early warnings and interventions through intelligent means.
[0006] The above limitations make it difficult for operation and maintenance personnel to quickly understand the system state and the location of problems, thus affecting the timely resolution of problems and making it difficult to meet the efficiency requirements of system operation. Summary of the Invention
[0007] In view of the problems in the prior art, this application provides an intelligent identification method and device for system hot issues, which can improve the efficiency and accuracy of identifying system hot issues.
[0008] To solve at least one of the above problems, this application provides the following technical solutions:
[0009] In a first aspect, this application provides an intelligent identification method for system hot issues, including:
[0010] Obtain multi-source historical system data, perform timestamp format standardization operations on the multi-source historical system data, and perform comparison operations on the multi-source historical system data after the timestamp format standardization operations to determine corresponding multi-source historical system data on the same time axis, where the multi-source historical system data is used to characterize system performance;
[0011] Perform abnormal event correlation relationship mining operations on the multi-source historical system data on the same time axis according to a preset abnormal correlation algorithm to determine strongly correlated events corresponding to abnormal events, perform causal relationship reasoning operations on the strongly correlated events according to a set causal inference model to determine corresponding strongly correlated event causal chains, and perform loss calculation and iterative optimization operations on the root cause events obtained according to the strongly correlated event causal chains and preset actual root cause event labels to determine corresponding system hotspot recognition models;
[0012] Receive multi-source real-time system data, perform abnormal identification operations on the multi-source real-time system data according to preset system hotspot problem detection rules to determine corresponding abnormal patterns, perform dynamic adjustment operations on the abnormal patterns according to the system hotspot recognition model, and perform system hotspot problem recognition according to the abnormal patterns obtained after the dynamic adjustment operations to determine corresponding system hotspot problems.
[0013] Further, the performing timestamp format standardization operations on the multi-source historical system data, and performing comparison operations on the multi-source historical system data after the timestamp format standardization operations to determine corresponding multi-source historical system data on the same time axis includes;
[0014] Perform format unification operations on the timestamps in the multi-source historical system data according to a preset timestamp format;
[0015] And perform comparison operations on the multi-source historical system data after the format unification operations to determine corresponding multi-source historical system data on the same time axis.
[0016] Further, the performing comparison operations on the multi-source historical system data after the format unification operations to determine corresponding multi-source historical system data on the same time axis includes:
[0017] Perform deviation adjustment operations and timestamp sorting operations on the multi-source historical system data after the format unification operations to determine corresponding sequential multi-source historical system data;
[0018] Perform time zone division operations on the sequential multi-source historical system data according to a preset time window algorithm to determine multi-source historical system data within corresponding multiple time regions;
[0019] Perform synchronous alignment operations on the historical system multi-source data in the multiple time regions respectively, and perform missing value filling operations and correlation fusion operations on the historical system multi-source data after the synchronous alignment operations to determine the corresponding historical system multi-source data on the same time axis.
[0020] Further, the operation of mining the correlation relationship of abnormal events on the historical system multi-source data on the same time axis according to the preset abnormal correlation algorithm to determine the strongly correlated events corresponding to the abnormal events includes:
[0021] Perform a recursive traversal operation on the historical system multi-source data on the same time axis to determine the frequent item sets corresponding to the abnormal events;
[0022] Perform a confidence calculation operation on the frequent item sets to determine the corresponding frequent item confidence;
[0023] Perform a screening operation on the frequent item confidence according to the preset confidence threshold to determine the strongly correlated events corresponding to the abnormal events.
[0024] Further, before performing the causal relationship reasoning operation on the strongly correlated events according to the set causal inference model to determine the corresponding causal chain of the strongly correlated events, it includes:
[0025] Perform variable definition operations on the historical system multi-source data on the same time axis according to the preset system performance influencing factors to determine the corresponding key variables, where the key variables include explicit variables and latent variables;
[0026] Perform causal relationship definition operations on the key variables, and perform causal model construction operations according to the causal paths obtained after the causal relationship definition operations to determine the corresponding initial causal model;
[0027] Perform model tuning operations on the initial causal model according to the preset goodness-of-fit criteria to determine the corresponding causal inference model.
[0028] Further, the operation of performing causal relationship definition operations on the key variables and performing causal model construction operations according to the causal paths obtained after the causal relationship definition operations to determine the corresponding initial causal model includes:
[0029] Perform index construction operations on the latent variables in the key variables according to the preset causal relationship hypothesis graph to determine the measurement indexes corresponding to each latent variable;
[0030] Perform causal relationship definition operations according to the measurement indexes to determine the corresponding causal paths, and perform causal model construction operations according to the causal paths to determine the corresponding initial causal model, where the causal paths include causal paths between latent variables and causal paths between latent variables and explicit variables.
[0031] Further, the causal relationship reasoning operation is performed on the strongly associated events according to the set causal inference model to determine the corresponding causal chain of the strongly associated events, including:
[0032] Performing a causal relationship reasoning operation on the strongly associated events according to the set causal inference model to determine the causal order of the strongly associated events and the root cause event corresponding to the causal order;
[0033] Determining the corresponding causal chain of the strongly associated events according to the causal order of the strongly associated events and the root cause event corresponding to the causal order.
[0034] In a second aspect, the present application provides an intelligent recognition device for system hot issues, including:
[0035] A multi-source data preprocessing module, configured to obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to reflect system performance;
[0036] A hot spot recognition model construction module, configured to perform an abnormal event correlation relationship mining operation on the historical system multi-source data on the same time axis according to a preset abnormal correlation algorithm to determine strongly associated events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strongly associated events according to a set causal inference model to determine the corresponding causal chain of the strongly associated events, and perform a loss calculation and iterative tuning operation on the root cause event obtained according to the causal chain of the strongly associated events and a preset actual root cause event label to determine the corresponding system hot spot recognition model;
[0037] A hot spot recognition module, configured to receive real-time system multi-source data, perform an abnormal recognition operation on the real-time system multi-source data according to a preset system hot issue detection rule to determine the corresponding abnormal pattern, perform a dynamic adjustment operation on the abnormal pattern according to the system hot spot recognition model, and perform system hot issue recognition according to the abnormal pattern obtained after the dynamic adjustment operation.
[0038] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the intelligent recognition method for system hot issues are implemented.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the intelligent recognition method for system hot issues are implemented.
[0040] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions, which when executed by a processor implement the steps of the intelligent recognition method for system hot issues described above.
[0041] As can be seen from the above technical solutions, the present application provides an intelligent recognition method and device for system hot issues. By unifying the timestamp formats of multi-source data of historical systems, comparing them on the same time axis, obtaining multi-source data of historical systems on the same time axis, and inputting the data into a preset initial model, events strongly correlated with system performance anomalies in the multi-source data of the system are mined through an association rule algorithm and mapped into a knowledge graph. Causal relationship reasoning operations are performed on the knowledge graph according to a causal inference model to obtain the root cause events leading to system performance anomalies. Loss calculations are performed on the root cause events and preset root cause event tags to obtain a system hot issue recognition model; real-time multi-source data of the system is received, abnormal patterns are identified according to system hot issue detection rules, and the abnormal patterns are dynamically adjusted according to the system hot issue recognition model to determine system hot issues, thereby improving the efficiency and accuracy of system hot issue recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is a schematic flowchart of the intelligent recognition method for system hot issues in an embodiment of the present application;
[0044] Figure 2 It is a schematic flowchart of the intelligent recognition method for system hot issues in an embodiment of the present application;
[0045] Figure 3 It is a schematic flowchart of the intelligent recognition method for system hot issues in an embodiment of the present application;
[0046] Figure 4 It is a schematic flowchart of the intelligent recognition method for system hot issues in an embodiment of the present application;
[0047] Figure 5 It is a schematic flowchart of the intelligent recognition method for system hot issues in an embodiment of the present application;
[0048] Figure 6 It is a schematic flowchart of the intelligent recognition method for system hot issues in an embodiment of the present application;
[0049] Figure 7 This is the seventh flowchart of the intelligent recognition method for system hot issues in the embodiments of the present application;
[0050] Figure 8 This is the structural diagram of the intelligent recognition device for system hot issues in the embodiments of the present application;
[0051] Figure 9 This is the structural diagram of the electronic device in the embodiments of the present application.
[0052] Reference numerals:
[0053] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0055] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0056] Considering that the existing system problem handling method makes it difficult for operation and maintenance personnel to quickly understand the system status and the location of problems, thus affecting the timely resolution of problems and making it difficult to meet the efficiency requirements of system operation. This application provides an intelligent identification method and device for system hot issues. By unifying the timestamp formats of multi-source data of historical systems, comparing them on the same time axis, obtaining multi-source data of historical systems on the same time axis, and inputting the data into a preset initial model, events strongly related to system performance anomalies in the multi-source data of the system are mined through an association rule algorithm and mapped into a knowledge graph. Causal relationship reasoning operations are performed on the knowledge graph according to a causal inference model to obtain the root cause events leading to system performance anomalies. The root cause events and preset root cause event tags are used for loss calculation to obtain a system hot issue identification model; real-time multi-source data of the system is received, abnormal patterns are identified according to system hot issue detection rules, and the abnormal patterns are dynamically adjusted according to the system hot issue identification model to determine system hot issues, thereby improving the efficiency and accuracy of system hot issue identification.
[0057] To improve the efficiency and accuracy of system hot issue identification, this application provides an embodiment of an intelligent identification method for system hot issues. Refer to Figure 1 The intelligent identification method for system hot issues specifically includes the following content:
[0058] Step S101: Obtain multi-source data of historical systems, perform timestamp format standardization operations on the multi-source data of historical systems, and perform comparison operations on the multi-source data of historical systems after the timestamp format standardization operations to determine the corresponding multi-source data of historical systems on the same time axis, where the multi-source data of historical systems is used to characterize system performance;
[0059] Optionally, in this embodiment, the purpose of this step is to first preprocess the multi-source data to ensure that data from different sources can be compared and analyzed on the same time axis.
[0060] Optionally, in this embodiment, the multi-source data of historical systems is to collect historical system data from different sources to ensure that multiple aspects of the system status (performance, network, logs, etc.) are covered, including but not limited to:
[0061] System logs: Log information of system operation.
[0062] Performance monitoring tool data: CPU usage, memory occupancy, disk I / O, etc.
[0063] Network traffic data: Used to analyze the network behavior of the system.
[0064] Application sensor data: Outputs of, for example, temperature sensors, pressure sensors, etc.
[0065] Optionally, in this embodiment, for the operation of standardizing the timestamp format, different data sources may adopt different time zones or time formats, resulting in inconsistent timestamps. To ensure that data can be aligned in chronological order, the timestamps of all data need to be in a unified format. Preferably, in this embodiment, it is unified into the UTC timestamp format. UTC time is not affected by time zone and daylight saving time changes and is the standard time globally.
[0066] Optionally, in this embodiment, the multi-source data of the same time-axis history system is determined after synchronizing and comparing the data from different sources with standardized timestamps. The specific method is as follows: First, the Kafka stream processing platform and time window technology are used to synchronize the data to ensure that the data within the same time window can be reasonably compared and integrated. It can be understood that the Kafka stream processing platform realizes the real-time transmission and processing of data, ensuring that the system can stably receive and process data regardless of the data scale. The time window technology divides the data into time blocks of a fixed size, which can ensure the integrity and alignment of the data within each time window. Then, after the timestamp standardization and synchronization of the data are completed, the cross-data-source comparison operation is performed. The data aligned by timestamps can be associated to determine the multi-source data at the same time point, which is the comparison of different data sources (such as log data, performance monitoring data, etc.) within the same time window. In this way, the historical data from different systems or sensors can be integrated under the same time framework, facilitating analysis and decision-making.
[0067] For example, if the CPU usage rate is abnormally high during a certain period, data such as network traffic and memory occupancy can be viewed simultaneously to form a comprehensive view to determine the root cause of the problem.
[0068] Optionally, in this embodiment, subsequently, the data after the synchronization and comparison operations will be stored in the Spark big data platform for subsequent query and analysis. The big data platform Spark supports the execution of complex query and analysis models. To improve the query efficiency, preferably, an index is created using the timestamp to accelerate the query of data within a specific time range.
[0069] After the above step S101, the timestamp standardization and synchronization comparison of the multi-source data are successfully completed, laying a solid data foundation for the subsequent machine learning data patterns.
[0070] Step S102: Perform an abnormal event correlation relationship mining operation on the multi-source historical system data on the same time axis according to a preset abnormal correlation algorithm, determine the strongly correlated events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strongly correlated events according to a set causal inference model, determine the corresponding causal chain of the strongly correlated events, and perform a loss calculation and iterative optimization operation on the root cause events obtained according to the causal chain of the strongly correlated events and the preset actual root cause event labels to determine the corresponding system hotspot recognition model;
[0071] Optionally, in this embodiment, this step trains a system hotspot recognition model, which is composed of an association analysis module, a graph database module, and a causal inference module.
[0072] Optionally, in this embodiment, the association analysis module aims to discover strong correlations between events from historical data and automatically mine potential patterns in the data. First, scan the data set once to find frequent 1-itemsets and sort them in descending order of frequency to obtain a list L. Then, based on list L, scan the data set again and process each original transaction: delete the items not in L and sort them in the order of L to obtain the modified transaction set T'. Next, construct an FP-tree, sort and link the data in T' according to the frequent items to form a tree with NULL as the root node. Record the support degree of each node at each node. Then, starting from the bottom (leaf node) of the tree and going up, perform recursive mining of the conditional pattern basis for each node to find all frequent itemsets. Specifically, for each node, first find all its successor nodes (directly connected nodes), and then perform recursive mining on each successor node. During the recursive process, it is necessary to continuously update the conditional pattern basis and conditional FP-tree of each node until no more frequent itemsets can be found. Finally, after identifying the frequent itemsets, discover strongly correlated events by calculating confidence. Specifically, for each frequent itemset, calculate the confidence between all the items it contains. Confidence represents the probability that the result item appears under the condition that the premise item appears. According to the set confidence threshold, filter out the association rules with confidence higher than the threshold, and these rules are the strongly correlated events.
[0073] Optionally, in this embodiment, the graph database module constructs an event association graph through the Neo4j graph database to show the strong correlations between different events. The graph shows the influence chain between each event, helps identify which events are strongly correlated with each other, and helps further locate system hotspot problems.
[0074] Optionally, in this embodiment, the causal inference module uses a structural equation model (SEM) for causal reasoning to analyze the potential causal relationships between events. Specifically, the construction of the structural equation model (SEM):
[0075] First, identify the hotspots, abnormal patterns, and potential causal relationships in the system performance that we need to address. Then, within this framework, establish the theoretical relationships between latent variables and manifest variables.
[0076] Latent variables are variables that cannot be directly observed, such as system performance, abnormal patterns, and causal relationships. For example, "system performance" includes multiple latent variables (such as load, response time, etc.).
[0077] Manifest variables are data items that can be observed and directly measured, such as CPU usage rate, network latency, etc. collected through a monitoring system.
[0078] Define how manifest variables reflect latent variables through data modeling, and establish measurement indicators for each latent variable. For example, measure the latent variable of system performance through multiple manifest variables (such as system load, response time, CPU usage rate).
[0079] Then, define the causal relationships between latent variables and set the causal paths. Among them, the causal paths between latent variables and between latent variables and manifest variables need to be set according to the theoretical framework. For example, "event correlation" can be set as a latent variable, and abnormal events identified through machine learning or rule engines can be used as manifest variables. For example, "system load" affects "response time", and "response time" in turn affects the latent variable of "user experience".
[0080] Next, for the defined initial causal model, use maximum likelihood estimation (ML) for model estimation, use the lavaan package (structural equation analysis package) in AMOS for model fitting, evaluate whether the fitting indicators meet the standards, and check the path coefficients and significance between latent variables and between latent variables and manifest variables.
[0081] Preferably, use the TLI>0.90 standard to evaluate the goodness of fit of the model, ensure that the coefficients of each path are significant, and verify the causal relationships between latent variables. According to the test results of the goodness of fit and path coefficients, adjust the structure of the model, add or delete paths until the model goodness of fit reaches the best state, and obtain a causal inference model. Based on the path coefficients of the causal inference model, conduct causal reasoning, analyze potential causal relationships, and explore how to improve system performance by adjusting certain variables.
[0082] For example, the system discovers strong correlations between events from multi-source data through the association rule algorithm. By mining a large amount of historical data, the association rule algorithm can discover events that frequently occur simultaneously and their mutual relationships. Such as the frequent association between "high load" and "database connection failure". Then, the Neo4j graph database is used to construct an event association graph, visually representing the relationships between various events. In the graph, nodes represent events, and edges represent the correlations between events.
[0083] Event A: The database response time exceeds the preset threshold.
[0084] Event B: The system CPU usage exceeds 90%.
[0085] Event C: The server reboots.
[0086] In the graph, assume that a strong correlation is found between Event A and Event B (i.e., when the database response time is too long, the system CPU usage also increases sharply), and there is also a causal relationship between Event B and Event C (too high CPU causes the system to reboot)
[0087] After the knowledge graph is constructed, next, a causal inference model is applied to infer causal relationships. The causal inference model analyzes the potential causal chains between events, revealing which events are the root causes leading to the occurrence of other events. Based on the association relationships already identified in the graph, the causal inference model further analyzes whether there are causal relationships between various events: when Event A (the database response time is too long) occurs, it directly causes Event B (the CPU usage increases), and Event B then triggers Event C (the server reboots). This causal chain describes the process of a system performance problem occurring.
[0088] It should be noted that when constructing the causal chain, the causal inference model not only considers the causal relationships between single events but also combines other influencing factors (such as the external environment, system load changes, etc.) to ensure the accuracy and reliability of the inference results.
[0089] Causal chain: Event A → Event B → Event C
[0090] Event A: High concurrent requests cause an increase in the database response time.
[0091] Event B: The increase in the database response time causes an increase in the CPU load.
[0092] Event C: The CPU load is too high, ultimately causing the server to reboot.
[0093] Through the reasoning of this causal chain, the system found the root cause leading to the server restart, which is the "high concurrent requests" event. Subsequently, the system can give early warnings when "high concurrent requests" occur and predict subsequent problems based on the causal chain, thus providing improvement suggestions for the operation and maintenance personnel.
[0094] Optionally, in this embodiment, after obtaining the root cause event output by the causal inference module, the initial model is trained according to the real root cause label. Specifically, the root cause of each event can be assigned a label according to its impact on the system state. For example, if an event causes the system to crash, its label can be marked as "abnormal", otherwise it is marked as "normal". The gap between the root cause event output by the model and the label is calculated through the cross-entropy loss function, and the backpropagation algorithm and the gradient descent method are used to adjust the parameters. Through multiple iterations, the model is continuously optimized, gradually reducing the difference between the predicted label and the actual label.
[0095] It can be understood that in each iteration, the model will refine and adjust the causal chain of strongly associated events to identify more potential hot issues, optimize the system identification performance, and finally form an efficient and accurate system hot issue identification model through continuous iteration and optimization for real-time identification of hot issues in the system.
[0096] Step S103: Receive multi-source data of the real-time system, perform abnormal identification operations on the multi-source data of the real-time system according to the preset system hot issue detection rules, determine the corresponding abnormal pattern, perform dynamic adjustment operations on the abnormal pattern according to the system hot issue identification model, and identify system hot issues according to the abnormal pattern obtained after the dynamic adjustment operation, and determine the corresponding system hot issues.
[0097] Optionally, in this step, a unified data format is formulated through the API (data interface), and the rule engine and the machine learning model obtain the same data input through the interface.
[0098] After receiving the real-time data, the system will perform preliminary abnormal identification according to the preset rules. The rule engine (Drools) will perform rule matching by checking whether the data conforms to a specific abnormal pattern. In this stage, the system detects known abnormal patterns through the rule library. For example, it is stipulated that "if the server CPU usage rate exceeds 90% and lasts for more than 5 minutes, an abnormal alarm will be triggered". These patterns are defined according to historical data, system experience rules, and manually set standards. When the server CPU usage rate reaches 90%, the rule engine will immediately trigger an alarm.
[0099] Meanwhile, for exceptions not recorded in the rule engine, they are identified by the hot issue recognition model we trained in step S102. For example, within a specific time period, although the CPU usage rate of each server does not exceed 90%, the mutual dependence between different servers causes the overall system performance to decline, ultimately affecting the user experience. The hot issue recognition model can discover some potential association patterns by training on historical data. For example, certain specific events (such as promotional activities, high-concurrency requests) can lead to resource contention among multiple service nodes, resulting in system crashes and performance anomalies.
[0100] It is worth noting that the hot issue recognition model can identify not only immediate abnormal patterns but also gradually changing abnormal trends. For example, although the CPU usage rate for a single time does not exceed 90%, in the past few months, the CPU usage rate has been continuously rising, leading to a gradual deterioration of performance. At this time, the system needs to monitor and identify this abnormal trend in order to make adjustments in a timely manner. Through time series analysis and causal inference models, the system can discover the gradual upward trend of performance indicators. For example, the continuous increase in CPU usage rate is due to overloaded system load, low efficiency of certain code, or uneven resource allocation. This trend is identified in the real-time monitoring of the system, providing a basis for the warning system.
[0101] Next, based on the results of the second abnormal pattern (complex pattern recognition) and abnormal trend (trend analysis), the system will dynamically adjust the first abnormal pattern (known rule abnormal pattern). For example, if the hot issue recognition model discovers that the impact of a specific event or trend is significant, it can adjust the thresholds or policies in the rule engine so that the alarm system can respond to system problems more timely and accurately.
[0102] For example, the original rule was that an alarm is triggered when the CPU usage rate exceeds 90% and lasts for 5 minutes. However, through the analysis of trends and complex patterns, it is adjusted to: when the CPU usage rate exceeds 85% and lasts for 3 minutes, or when certain specific events (such as high-concurrency access) occur, an alarm is triggered immediately.
[0103] After dynamic adjustment, the system can more accurately identify potential hot issues. For example, through the dynamically adjusted rule engine, the system can detect the increase in CPU usage rate in a timely manner and associate it with other abnormal events (such as database response latency), and finally identify the hot issue of the system - the performance bottleneck of the database leads to overuse of server resources.
[0104] Finally, through the integrated visualization tool in this solution, the system visually displays the hot issues to the operation and maintenance personnel in the form of charts, heat maps, etc., to help them quickly locate and handle problems.
[0105] This example demonstrates how the present embodiment collaborates with a rule engine and a hot issue recognition model to perform real-time analysis of the influencing factors of system anomalies, quickly and accurately identify hot issues, and find the root factors affecting system performance, thereby improving the efficiency and accuracy of system hot issue recognition.
[0106] As can be seen from the above description, the intelligent recognition method for system hot issues provided by the embodiments of the present application can unify the timestamp formats of multi-source data of historical systems, compare them on the same time axis to obtain multi-source data of historical systems on the same time axis, input the data into a preset initial model, mine events strongly correlated with system performance anomalies in the multi-source data of the system through an association rule algorithm, map them into a knowledge graph, perform causal relationship reasoning operations on the knowledge graph according to a causal inference model to obtain root cause events leading to system performance anomalies, calculate the loss between the root cause events and preset root cause event tags to obtain a system hot issue recognition model; receive real-time multi-source data of the system, identify abnormal patterns according to system hot issue detection rules, and dynamically adjust the abnormal patterns according to the system hot issue recognition model to determine system hot issues, thereby improving the efficiency and accuracy of system hot issue recognition.
[0107] In an embodiment of the intelligent recognition method for system hot issues of the present application, referring to Figure 2 , it may further specifically include the following content:
[0108] Step S201: Perform a format unification operation on the timestamps in the multi-source data of the historical system according to a preset timestamp format;
[0109] Step S202: Perform a comparison operation on the multi-source data of the historical system after the format unification operation to determine the corresponding multi-source data of the historical system on the same time axis.
[0110] Optionally, to ensure that the data can be aligned in chronological order, the timestamps of all data need to be in a unified format. Preferably, the present embodiment unifies to the UTC timestamp format. The UTC time is not affected by time zones and daylight saving time changes and is the standard time globally.
[0111] Specifically, different data sources may adopt different time formats. For example:
[0112] UNIX timestamp (the number of seconds / milliseconds represents the number of seconds since January 1, 1970)
[0113] ISO 8601 format (2024-12-23T12:34:56Z)
[0114] Custom time format (yyyy-mm-dd hh:mm:ss)
[0115] To ensure the consistency of the timestamp formats for all data sources, we perform a standardized conversion of the timestamp formats in the data access platform or the data processing pipeline, converting all timestamps to a unified timestamp format.
[0116] Next, for multi-source data comparison, we perform deviation adjustment and sorting on the multi-source data with unified timestamps. Specifically, since there will be time deviations in data from different sources (such as due to network latency, clock drift, etc.), we adopt a synchronization strategy. At the data access stage, according to the transmission time of the data packet, we adjust the received timestamp. For example, when receiving a data packet, we can record the current reception time and calculate the transmission delay to adjust the timestamp. At the same time, there are clock drifts in different devices, resulting in inconsistent timestamps. We use the NTP (Network Time Protocol) protocol to synchronize the system times of each node regularly to ensure the consistency of timestamps. After deviation adjustment, the multi-source data is sorted and stored in timestamp order.
[0117] To align the sorted multi-source data on the time axis for convenient multi-source data comparison, we group events into time windows, for example, classifying events by hour, by minute, or by day, so as to unify the data within the same time period onto the same time axis. For cases where there are gaps between timestamps, we use a linear interpolation algorithm to fill in the missing time periods.
[0118] If there are certain delays in the timestamps of some data sources, we ensure data consistency by setting a reasonable delay tolerance threshold. For example, allowing a certain delay time, preferably 0.3 seconds, as long as the data is still within the acceptable timeliness range, synchronization can be performed.
[0119] Through step S202, this embodiment obtains multi-source data on the same time axis, which is convenient for subsequent comparison between multi-source data and lays a solid data foundation for the machine learning algorithm to discover the associations between multi-source data.
[0120] In an embodiment of the intelligent identification method for system hot issues in this application, refer to Figure 3 , and it may specifically include the following content:
[0121] Step S301: Perform deviation adjustment operations and timestamp sorting operations on the historical system multi-source data after the format unification operation to determine the corresponding sequential historical system multi-source data;
[0122] Step S302: Perform time zone division operations on the sequential historical system multi-source data according to a preset time window algorithm to determine the historical system multi-source data within the corresponding multiple time regions;
[0123] Step S303: Perform synchronous alignment operations on the historical system multi-source data in the multiple time regions respectively, and perform missing value filling operations and correlation fusion operations on the historical system multi-source data after the synchronous alignment operations to determine the corresponding historical system multi-source data on the same time axis.
[0124] Optionally, in this embodiment, this step is a refinement of the step of comparing the historical system multi-source data after timestamp synchronization. Specifically speaking, after the timestamps are unified, our multi-source data already has a unified timestamp format that can be used for comparison. At this time, when receiving data, perform deviation adjustment and sorting according to the unified timestamp to obtain sequential historical system multi-source data.
[0125] Specifically, in the data access stage, the deviation adjustment is to adjust the received timestamp according to the transmission time of the data packet. For example, when receiving a data packet, the current reception time can be recorded and the transmission delay can be calculated to adjust the timestamp. At the same time, there are drifts in the clocks of different devices, resulting in inconsistent timestamps. We use the NTP (Network Time Protocol) protocol to synchronize the system time of each node regularly to ensure the consistency of timestamps. When the data flows into the system, use a buffer to temporarily store the data for batch processing and analysis. Sort the data in the buffer according to the standardized timestamp to ensure that all events are arranged in chronological order in the subsequent processing steps.
[0126] Next, perform data processing on the received sequential historical system multi-source data, including alignment between multi-source data. Specifically speaking, use the time window technology to divide the data into individual time intervals (such as every minute, every hour). Within each time window, align the data using the timestamp. Use the interpolation method to identify and fill in the missing data points found during the synchronization process to ensure the integrity of the data in each time window. On the basis of alignment, correlate and fuse the data from different sources according to the time window for subsequent unified analysis.
[0127] Through step S303, this embodiment obtains multi-source data on the same time axis, which is convenient for subsequent comparative analysis between multi-source data and lays a solid data foundation for machine learning algorithms to discover the correlations between multi-source data.
[0128] In an embodiment of the intelligent identification method for system hot issues in this application, refer to Figure 4 , and it may also specifically include the following content:
[0129] Step S401: Perform a recursive traversal operation on the historical system multi-source data on the same time axis to determine the frequent item sets corresponding to abnormal events;
[0130] Step S402: Perform a confidence calculation operation on the frequent item sets to determine the corresponding frequent item confidence levels.
[0131] Step S403: Perform a screening operation on the frequent item confidence levels according to a preset confidence threshold to determine strongly associated events corresponding to abnormal events.
[0132] Optionally, in this embodiment, the purpose of step S401 is to find highly correlated frequent item sets. First, scan the data set once to find frequent 1-item sets, and sort them in descending order of frequency to obtain a list L. Then, based on list L, scan the data set again and process each original transaction: delete the items not in L and arrange them in the order of L to obtain the modified transaction set T'. Next, construct an FP-tree, sort and link the data in T' according to the frequent items to form a tree with NULL as the root node. Record the support degree of the appearance of each node at each node. Then, starting from the bottom (leaf nodes) of the tree and going up, perform recursive mining of the conditional pattern bases for each node to find all frequent item sets. Specifically, for each node, first find all its successor nodes (directly connected nodes), and then perform recursive mining on each successor node. During the recursive process, it is necessary to continuously update the conditional pattern bases and conditional FP-trees of each node until no more frequent item sets can be found.
[0133] Optionally, in this embodiment, the purposes of steps S402 and S403 are to find events strongly associated with system abnormal events based on the frequent item sets. Specifically, for each frequent item set, calculate the confidence levels between all the items it contains. Confidence level represents the probability of the result item appearing under the condition that the premise item appears. According to the set confidence threshold, filter out the association rules with confidence levels higher than the threshold, and these rules are the strongly associated events.
[0134] Illustrate with an example. Suppose the events collected by the system include:
[0135] Event 1: CPU usage exceeds 90%.
[0136] Event 2: Memory usage exceeds 85%.
[0137] Event 3: Network traffic exceeds the predetermined threshold.
[0138] Using the association algorithm, the following associations will be found:
[0139] There is a strong association between Event 1 and Event 2, that is, high CPU usage is usually accompanied by an increase in memory occupancy.
[0140] There is also a strong association between Event 2 and Event 3. When memory occupancy is high, network traffic is usually large.
[0141] Through step S403, in this embodiment, events strongly related to system exception events and the relationships between events are successfully discovered from the multi-source data of the historical system on the same time axis, laying a foundation for finding the root cause events through causal inference subsequently.
[0142] In an embodiment of the intelligent identification method for system hot issues of the present application, refer to Figure 5 , and it may specifically include the following content:
[0143] Step S501: Perform variable definition operations on the multi-source data of the historical system on the same time axis according to preset system performance influencing factors to determine corresponding key variables, where the key variables include manifest variables and latent variables;
[0144] Step S502: Perform causal relationship definition operations on the key variables, and perform causal model construction operations according to the causal paths obtained after the causal relationship definition operations to determine corresponding initial causal models;
[0145] Step S503: Perform model tuning operations on the initial causal model according to preset goodness-of-fit criteria to determine corresponding causal inference models.
[0146] Optionally, this step is the construction process of the causal inference model. First, it is determined that the problem to be solved is to identify hot issues, abnormal patterns, and potential causal relationships of system performance. Then, within this framework, key variables affecting system performance are defined, including latent variables and manifest variables.
[0147] Latent variables are variables that cannot be directly observed, such as system performance, abnormal patterns, and causal relationships. For example, "system performance" includes multiple latent variables (such as load, response time, etc.).
[0148] Manifest variables are data items that can be observed and directly measured. For example, CPU usage rate, network latency, etc. collected through the monitoring system.
[0149] An initial causal model is constructed through data modeling to define how manifest variables reflect latent variables and establish measurement indicators for each latent variable. For example, a latent variable such as system performance is measured through multiple manifest variables (such as system load, response time, CPU usage rate).
[0150] Then, the causal relationships of latent variables are defined, and causal paths are set. Among them, the causal paths between latent variables and between latent variables and manifest variables need to be set according to the theoretical framework. For example, "event association" can be set as a latent variable, and abnormal events identified through machine learning or rule engines can be used as manifest variables. For example, "system load" affects "response time", and "response time" in turn affects a latent variable such as "user experience".
[0151] Next, for the defined initial causal model, maximum likelihood estimation (ML) is used for model estimation, and the lavaan package (structural equation analysis package) in AMOS is used for model fitting. The goodness-of-fit index is evaluated to check whether it meets the standard, and the path coefficients and significance between latent variables and between latent variables and manifest variables are examined.
[0152] Preferably, the goodness of fit uses the TLI>0.90 standard to evaluate the goodness of fit of the model, ensuring that the coefficient of each path is significant and verifying the causal relationship between latent variables. According to the test results of the goodness of fit and path coefficients, the structure of the model is adjusted, paths are added or deleted until the model goodness of fit reaches the best state, and a causal inference model is obtained.
[0153] Through step S503, this embodiment successfully obtains a causal inference model, laying a solid foundation for finding the root cause of system performance anomaly events.
[0154] In an embodiment of the intelligent identification method for system hot issues of the present application, referring to Figure 6 it may further specifically include the following content:
[0155] Step S601: Perform an index construction operation on the latent variables in the key variables according to a preset causal relationship hypothesis diagram to determine measurement indicators corresponding to each latent variable;
[0156] Step S602: Perform a causal relationship definition operation according to the measurement indicators to determine corresponding causal paths, and perform a causal model construction operation according to the causal paths to determine corresponding initial causal models, where the causal paths include causal paths between latent variables and between latent variables and manifest variables.
[0157] Optionally, in this embodiment, the causal relationship hypothesis diagram is a graphical representation method for showing the causal relationships between different variables. This diagram consists of latent variables and manifest variables. By defining the relationships between latent variables and manifest variables, it helps to infer the causal paths between these variables and guides the subsequent construction of causal models. In the causal relationship hypothesis diagram, latent variables need to be quantified through some measurement indicators. Measurement indicators are manifest variables related to latent variables and can reflect the characteristics of latent variables. Based on prior domain knowledge and data analysis, variables that can accurately measure latent variables are selected. For example, if the latent variable is "system load", then the corresponding measurement indicators are "CPU usage rate", "memory occupancy rate", etc.
[0158] Optionally, in this embodiment, the causal path is a causal chain between latent variables and manifest variables or between latent variables.
[0159] Causal paths between latent variables: There are mutual influence relationships between latent variables. For example, one latent variable (such as "network latency") affects another latent variable (such as "system response time").
[0160] Causal paths between latent variables and manifest variables: Manifest variables, as measurement indicators of latent variables, affect latent variables. For example, CPU usage rate (manifest variable) can be used as a measurement indicator of system load (latent variable) and in turn affects system performance (latent variable).
[0161] By transforming the above causal relationship paths into a structural equation model to describe the relationships between variables, after the preliminary construction of the model, testing and adjustment can be carried out until the model can effectively describe the causal relationships in the data.
[0162] Through step S602, this embodiment successfully obtains the initial causal model, laying a solid foundation for optimizing the initial model to obtain a causal inference model for finding the root causes affecting system performance.
[0163] In an embodiment of the intelligent identification method for system hot issues in this application, refer to Figure 7 , and it may specifically include the following content:
[0164] Step S701: Perform a causal relationship reasoning operation on the strongly correlated events according to the set causal inference model to determine the causal order of the strongly correlated events and the root cause events corresponding to the causal order;
[0165] Step S702: Determine the corresponding causal chain of strongly correlated events according to the causal order of the strongly correlated events and the root cause events corresponding to the causal order.
[0166] Optionally, in this embodiment, the events and their association relationships in the graph database are used as inputs, and a causal inference model is used for causal reasoning to obtain a causal chain. The goal of causal inference is to identify which events are the root causes leading to other events. For example, whether excessive memory usage directly causes excessive CPU usage rate, or is affected by other factors.
[0167] Illustrate with an example. Assume that the knowledge graph records events A, B, C and their relationships:
[0168] Event A: The database response time exceeds the preset threshold.
[0169] Event B: The system CPU usage rate exceeds 90%.
[0170] Event C: The server restarts.
[0171] Relationship: A strong association between event A and event B, and there is also a causal relationship between event B and event C.
[0172] Based on the association relationships already identified in the graph, the causal inference model further analyzes whether there is a causal order among events: When event A (the database response time is too long) occurs, it directly causes event B (the CPU usage rate increases), and event B in turn triggers event C (the server restarts). This causal chain describes the process of a system performance problem occurring.
[0173] It should be noted that when constructing the causal chain, the causal inference model not only considers the causal relationships between single events, but also combines other influencing factors (such as the external environment, system load changes, etc.) to ensure the accuracy and reliability of the inference results.
[0174] Causal chain: Event A → Event B → Event C
[0175] Event A: High concurrent requests lead to an increase in the database response time.
[0176] Event B: The increase in the database response time causes the CPU load to increase.
[0177] Event C: The CPU load is too high, which ultimately leads to the server restart.
[0178] Through the inference of this causal chain, the system found the root cause of the server restart, that is, the "high concurrent request" event. Subsequently, the system can give early warnings when "high concurrent requests" occur and predict subsequent problems according to the causal chain, so as to provide improvement suggestions for operation and maintenance personnel.
[0179] Through step S702, this embodiment successfully obtained the system performance causal chain. Through the causal chain, the root factors affecting the system performance can be effectively found, increasing the efficiency of identifying system hot issues.
[0180] To improve the efficiency and accuracy of identifying system hot issues, this application provides an embodiment of an intelligent identification device for system hot issues that implements all or part of the content of the intelligent identification method for system hot issues. See Figure 8 , the intelligent identification device for system hot issues specifically includes the following content:
[0181] The multi-source data preprocessing module 10 is used to obtain historical system multi-source data, perform timestamp format standardization operations on the historical system multi-source data, and perform comparison operations on the historical system multi-source data after the timestamp format standardization operations to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to characterize system performance;
[0182] A hotspot recognition model construction module 20 is configured to perform an abnormal event correlation relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm, determine strongly correlated events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strongly correlated events according to a set causal inference model, determine corresponding strongly correlated event causal chains, perform a loss calculation and iterative tuning operation on the root cause events obtained according to the strongly correlated event causal chains and preset actual root cause event labels, and determine a corresponding system hotspot recognition model;
[0183] A hotspot recognition module 30 is configured to receive multi-source data of a real-time system, perform an abnormal recognition operation on the multi-source data of the real-time system according to a preset system hotspot problem detection rule, determine a corresponding abnormal pattern, perform a dynamic adjustment operation on the abnormal pattern according to the system hotspot recognition model, and identify a system hotspot problem according to the abnormal pattern obtained after the dynamic adjustment operation.
[0184] As can be seen from the above description, the intelligent recognition device for system hotspot problems provided by the embodiments of the present application can unify the timestamp formats of multi-source data of the historical system, compare them on the same time axis to obtain multi-source data of the historical system on the same time axis, input the data into a preset initial model, mine events strongly related to system performance anomalies in the multi-source data of the system through an association rule algorithm, map them into a knowledge graph, perform a causal relationship reasoning operation on the knowledge graph according to a causal inference model to obtain root cause events leading to system performance anomalies, perform a loss calculation on the root cause events and preset root cause event labels to obtain a system hotspot recognition model; receive multi-source data of the real-time system, identify an abnormal pattern according to a system hotspot problem detection rule, perform a dynamic adjustment on the abnormal pattern according to the system hotspot recognition model, and determine a system hotspot problem, thereby improving the efficiency and accuracy of system hotspot problem recognition.
[0185] From a hardware level, in order to improve the efficiency and accuracy of system hotspot problem recognition, the present application provides an embodiment of an electronic device for implementing all or part of the content of the intelligent recognition method for system hotspot problems, and the electronic device specifically includes the following content:
[0186] A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete communication with each other through the bus; the communications interface is used to implement information transmission between the intelligent identification method for system hot issues and related devices such as a core business system, a user terminal, and a related database; this logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, this logic controller can be implemented with reference to the embodiments of the intelligent identification method for system hot issues in the embodiments, as well as the embodiments of the intelligent identification method for system hot issues, and the content is incorporated herein, and the repeated parts will not be elaborated again.
[0187] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0188] In practical applications, part of the intelligent identification method for system hot issues can be executed on the electronic device side as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor.
[0189] The above-mentioned client device may have a communication module (i.e., a communication unit), and can be communicatively connected to a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform communicatively linked to the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.
[0190] Figure 9 This is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of the present application. As Figure 9 shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 9 is exemplary; other types of structures can also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0191] In one embodiment, the function of the intelligent identification method for system hot issues can be integrated into the central processing unit 9100. Among them, the central processing unit 9100 can be configured to perform the following controls:
[0192] Step S101: Obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis. Among them, the historical system multi-source data is used to characterize system performance;
[0193] Step S102: Perform an abnormal event correlation relationship mining operation on the historical system multi-source data on the same time axis according to a preset abnormal correlation algorithm to determine strong correlation events corresponding to abnormal events. Perform a causal relationship reasoning operation on the strong correlation events according to a set causal inference model to determine the corresponding causal chain of the strong correlation events. Perform a loss calculation and iterative tuning operation on the root cause event obtained according to the causal chain of the strong correlation events and a preset actual root cause event label to determine the corresponding system hot spot identification model;
[0194] Step S103: Receive real-time system multi-source data, perform an abnormal identification operation on the real-time system multi-source data according to a preset system hot issue detection rule to determine the corresponding abnormal pattern. Perform a dynamic adjustment operation on the abnormal pattern according to the system hot spot identification model, and perform system hot issue identification according to the abnormal pattern obtained after the dynamic adjustment operation to determine the corresponding system hot issue.
[0195] As can be seen from the above description, the electronic device provided in the embodiment of the present application unifies the timestamp format of the historical system multi-source data, performs a comparison on the same time axis to obtain the historical system multi-source data on the same time axis, and inputs the data into a preset initial model. By using the association rule algorithm to mine the events strongly related to the system performance abnormality in the system multi-source data and map them into the knowledge graph, perform a causal relationship reasoning operation on the knowledge graph according to the causal inference model to obtain the root cause event leading to the system performance abnormality, perform a loss calculation on the root cause event and a preset root cause event label to obtain the system hot spot identification model; receive real-time system multi-source data, identify the abnormal pattern according to the system hot issue detection rule, and perform dynamic adjustment on the abnormal pattern according to the system hot spot identification model to determine the system hot issue, thereby improving the efficiency and accuracy of system hot issue identification.
[0196] In another embodiment, the intelligent identification method for system hot issues can be separately configured from the central processing unit 9100. For example, the intelligent identification method for system hot issues can be configured as a chip connected to the central processing unit 9100, and the function of the intelligent identification method for system hot issues is realized through the control of the central processing unit.
[0197] As Figure 9 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 9 all the components shown in; in addition, the electronic device 9600 may further include Figure 9 components not shown in, and reference may be made to the prior art.
[0198] As Figure 9 shown, the central processing unit 9100, sometimes also referred to as a controller or operation control, may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.
[0199] Among them, the memory 9140, for example, may be one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processing unit 9100 can execute the program stored in the memory 9140 to implement information storage or processing, etc.
[0200] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.
[0201] The memory 9140 may be a solid-state memory, for example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that stores information even when powered off, can be selectively erased and has more data. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, and the application / function storage unit 9142 is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.
[0202] The memory 9140 may further include a data storage unit 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0203] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.
[0204] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 may include any suitable buffers, decoders, amplifiers, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that it is possible to record on the local machine through the microphone 9132 and play the sound stored on the local machine through the speaker 9131.
[0205] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the intelligent identification method for system hot issues in the above embodiments where the execution subject is a server or a client. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements all steps of the intelligent identification method for system hot issues in the above embodiments where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:
[0206] Step S101: Obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to characterize system performance;
[0207] Step S102: Perform an abnormal event correlation relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm, determine strongly correlated events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strongly correlated events according to a set causal inference model, determine the corresponding causal chain of the strongly correlated events, and perform a loss calculation and iterative tuning operation on the root cause event obtained according to the causal chain of the strongly correlated events and a preset actual root cause event label to determine the corresponding system hotspot identification model;
[0208] Step S103: Receive multi-source data of the real-time system, perform an abnormal identification operation on the multi-source data of the real-time system according to a preset system hotspot problem detection rule, determine the corresponding abnormal pattern, perform a dynamic adjustment operation on the abnormal pattern according to the system hotspot identification model, and identify the system hotspot problem according to the abnormal pattern obtained after the dynamic adjustment operation.
[0209] As can be seen from the above description, the computer-readable storage medium provided by the embodiment of the present application unifies the timestamp format of the multi-source data of the historical system, compares them on the same time axis to obtain the multi-source data of the historical system on the same time axis, inputs the data into a preset initial model, mines the events strongly related to the system performance abnormality in the multi-source data of the system through an association rule algorithm, maps them into a knowledge graph, performs a causal relationship reasoning operation on the knowledge graph according to a causal inference model to obtain the root cause event leading to the system performance abnormality, performs a loss calculation on the root cause event and a preset root cause event label to obtain a system hotspot identification model; receives multi-source data of the real-time system, identifies an abnormal pattern according to a system hotspot problem detection rule, dynamically adjusts the abnormal pattern according to the system hotspot identification model, and determines the system hotspot problem, thereby being able to improve the efficiency and accuracy of identifying the system hotspot problem.
[0210] The embodiment of the present application further provides a computer program product capable of implementing all the steps in the intelligent identification method of the system hotspot problem with the execution subject being a server or a client in the above embodiment. When the computer program / instructions are executed by a processor, the steps of the intelligent identification method of the system hotspot problem are implemented. For example, the computer program / instructions implement the following steps:
[0211] Step S101: Obtain multi-source data of the historical system, perform a timestamp format standardization operation on the multi-source data of the historical system, and perform a comparison operation on the multi-source data of the historical system after the timestamp format standardization operation to determine the corresponding multi-source data of the historical system on the same time axis, where the multi-source data of the historical system is used to characterize the system performance;
[0212] Step S102: Perform an abnormal event correlation relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm to determine strongly correlated events corresponding to the abnormal events. Perform a causal relationship reasoning operation on the strongly correlated events according to a set causal inference model to determine the corresponding causal chain of the strongly correlated events. Perform a loss calculation and iterative tuning operation on the root cause events obtained according to the causal chain of the strongly correlated events and preset actual root cause event labels to determine the corresponding system hot spot recognition model;
[0213] Step S103: Receive multi-source data of the real-time system, perform an abnormal recognition operation on the multi-source data of the real-time system according to a preset system hot spot problem detection rule to determine the corresponding abnormal pattern. Perform a dynamic adjustment operation on the abnormal pattern according to the system hot spot recognition model, and identify system hot spot problems according to the abnormal pattern obtained after the dynamic adjustment operation.
[0214] As can be seen from the above description, the computer program product provided by the embodiment of the present application unifies the time stamp format of the multi-source data of the historical system, compares them on the same time axis to obtain the multi-source data of the historical system on the same time axis, and inputs the data into a preset initial model. By using an association rule algorithm to mine events strongly related to system performance anomalies in the multi-source data of the system and map them into a knowledge graph, perform a causal relationship reasoning operation on the knowledge graph according to a causal inference model to obtain the root cause events leading to system performance anomalies, and perform a loss calculation on the root cause events and preset root cause event labels to obtain a system hot spot recognition model; receive multi-source data of the real-time system, identify abnormal patterns according to the system hot spot problem detection rule, perform dynamic adjustment on the abnormal patterns according to the system hot spot recognition model, and determine system hot spot problems, thereby being able to improve the efficiency and accuracy of identifying system hot spot problems.
[0215] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0216] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0217] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0218] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0219] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An intelligent recognition method for system hotspot problems, characterized in that, The method includes: Obtaining multi-source historical system data, performing a timestamp format standardization operation on the multi-source historical system data, and performing a comparison operation on the multi-source historical system data after the timestamp format standardization operation to determine the corresponding multi-source historical system data on the same time axis. Among them, the multi-source historical system data includes system log data, performance monitoring tool data, network traffic data, and application sensor data. The performance monitoring tool data includes CPU usage rate and memory occupancy. The application sensor data includes the output of a temperature sensor and the output of a pressure sensor; Performing an abnormal event correlation relationship mining operation on the multi-source historical system data on the same time axis according to a preset abnormal correlation algorithm to determine strongly correlated events corresponding to abnormal events. Performing a variable definition operation on the multi-source historical system data on the same time axis according to preset system performance influencing factors to determine corresponding key variables. Among them, the key variables include explicit variables and latent variables; performing a causal relationship definition operation on the key variables, and constructing a causal model according to the causal path obtained after the causal relationship definition operation to determine a corresponding initial causal model; performing a model tuning operation on the initial causal model according to a preset goodness-of-fit criterion to determine a corresponding causal inference model, performing a causal relationship reasoning operation on the strongly correlated events according to the causal inference model to determine a corresponding strongly correlated event causal chain, and performing a loss calculation and iterative tuning operation on the root cause event obtained according to the strongly correlated event causal chain and a preset actual root cause event label to determine a corresponding system hot spot identification model; Receiving multi-source real-time system data, performing an abnormal identification operation on the multi-source real-time system data according to a preset system hot spot problem detection rule to determine a corresponding abnormal pattern, performing a dynamic adjustment operation on the abnormal pattern according to the system hot spot identification model, and identifying a system hot spot problem according to the abnormal pattern obtained after the dynamic adjustment operation to determine a corresponding system hot spot problem.
2. The intelligent recognition method for system hot issues according to claim 1, wherein The performing a timestamp format standardization operation on the multi-source historical system data, and performing a comparison operation on the multi-source historical system data after the timestamp format standardization operation to determine the corresponding multi-source historical system data on the same time axis includes: Performing a format unification operation on the timestamps in the multi-source historical system data according to a preset timestamp format; And performing a comparison operation on the multi-source historical system data after the format unification operation to determine the corresponding multi-source historical system data on the same time axis.
3. The intelligent recognition method for system hot issues according to claim 2, characterized in that The and performing a comparison operation on the multi-source historical system data after the format unification operation to determine the corresponding multi-source historical system data on the same time axis includes: Performing a deviation adjustment operation and a timestamp sorting operation on the multi-source historical system data after the format unification operation to determine corresponding sequential multi-source historical system data; Performing a time zone division operation on the sequential multi-source historical system data according to a preset time window algorithm to determine the multi-source historical system data within corresponding multiple time regions; Perform synchronous alignment operations on the historical system multi-source data in the multiple time regions respectively, and perform missing value filling operations and correlation fusion operations on the historical system multi-source data after the synchronous alignment operations to determine the corresponding historical system multi-source data on the same time axis.
4. The intelligent recognition method for system hot issues according to claim 1, characterized in that The operation of mining the correlation relationship of abnormal events on the historical system multi-source data on the same time axis according to the preset abnormal correlation algorithm to determine the strongly correlated events corresponding to the abnormal events includes: Perform a recursive traversal operation on the historical system multi-source data on the same time axis to determine the frequent item sets corresponding to the abnormal events; Perform a confidence calculation operation on the frequent item sets to determine the corresponding frequent item confidence; Perform a screening operation on the frequent item confidence according to the preset confidence threshold to determine the strongly correlated events corresponding to the abnormal events.
5. The intelligent identification method for system hot issues according to claim 4, characterized in that, The operation of defining the causal relationship for the key variables and constructing a causal model according to the causal path obtained after the causal relationship definition operation to determine the corresponding initial causal model includes: Perform an index construction operation on the latent variables in the key variables according to the preset causal relationship hypothesis graph to determine the measurement indexes corresponding to each latent variable; Perform a causal relationship definition operation according to the measurement indexes to determine the corresponding causal path, and perform a causal model construction operation according to the causal path to determine the corresponding initial causal model, where the causal path includes the causal path between latent variables and the causal path between latent variables and manifest variables.
6. The intelligent identification method for system hot issues according to claim 1, wherein The operation of inferring the causal relationship on the strongly correlated events according to the set causal inference model to determine the corresponding causal chain of the strongly correlated events includes: Perform a causal relationship inference operation on the strongly correlated events according to the set causal inference model to determine the causal order of the strongly correlated events and the root cause events corresponding to the causal order; Determine the corresponding causal chain of the strongly correlated events according to the causal order of the strongly correlated events and the root cause events corresponding to the causal order.
7. An intelligent recognition device for system hot issues, characterized in that The device includes: A multi-source data preprocessing module, configured to obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data includes system log data, performance monitoring tool data, network traffic data, and application sensor data, the performance monitoring tool data includes CPU usage rate and memory occupancy, and the application sensor data includes the output of a temperature sensor and the output of a pressure sensor; A hotspot recognition model construction module is used to perform anomaly event correlation relationship mining operations on the multi-source data of the historical system on the same time axis according to a preset anomaly correlation algorithm, determine strongly correlated events corresponding to the anomaly events, perform variable definition operations on the multi-source data of the historical system on the same time axis according to preset system performance impact factors, and determine corresponding key variables, where the key variables include explicit variables and latent variables; perform causal relationship definition operations on the key variables, perform causal model construction operations according to the causal paths obtained after the causal relationship definition operations, and determine corresponding initial causal models; perform model tuning operations on the initial causal models according to a preset goodness-of-fit criterion, determine corresponding causal inference models, perform causal relationship reasoning operations on the strongly correlated events according to the causal inference models, determine corresponding strongly correlated event causal chains, perform loss calculation and iterative tuning operations on the root cause events obtained according to the strongly correlated event causal chains and preset actual root cause event labels, and determine corresponding system hotspot recognition models; A hotspot recognition module is used to receive multi-source data of a real-time system, perform anomaly recognition operations on the multi-source data of the real-time system according to preset system hotspot problem detection rules, determine corresponding anomaly patterns, perform dynamic adjustment operations on the anomaly patterns according to the system hotspot recognition model, and identify system hotspot problems according to the anomaly patterns obtained after the dynamic adjustment operations.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the intelligent recognition method for system hotspot problems according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the intelligent recognition method for system hotspot problems according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Root cause recognition model training method, root cause recognition method, device and equipment
CN118827326A
Interventional information brokering medical tracking interface
US20150123769A1