Intelligent identification method and device for system hotspot problem
Through intelligent methods, using historical and real-time system multi-source data, a system hotspot recognition model is built, which solves the problems of manual dependence, low data processing efficiency and lack of intelligent support in traditional methods, and achieves efficient and accurate identification of system hotspot problems.
Patent Information
- Application Number
- CN202510560353.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional system problem identification and analysis methods rely on manual intervention, low data processing efficiency and lack of intelligent support, which makes it difficult to guarantee the efficiency and accuracy of system hot issues identification.
It provides an intelligent identification method for system hotspot problems. By obtaining and standardizing historical system multi-source data, using preset exception association algorithms and causal inference models, mining the relationship between abnormal events and inferring causal chains, building a system hotspot recognition model, and identifying system hotspot problems in real time.
It improves the efficiency and accuracy of system hot issues identification, reduces manual intervention, and achieves rapid response and intelligent early warning for system performance abnormalities.
Smart Images

Figure CN120069045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and particularly to an intelligent recognition method and device for system hot-spot problems. Background Art
[0002] With the rapid development of information technology, modern complex systems (such as large software systems, cloud computing platforms, Internet of Things systems, etc.) generate a vast amount of data during daily operations. This data covers multiple aspects such as system logs, performance monitoring metrics, network traffic information, sensor data, etc., and is crucial for the stable operation and performance optimization of the system. However, traditional system problem identification and analysis methods usually have the following limitations: Dependence on manual intervention: Traditional system monitoring and analysis often require manual participation, identifying potential problems by viewing logs, analyzing performance metrics, etc. This method is not only time-consuming and laborious but also vulnerable to human factors, making it difficult to ensure the accuracy and timeliness of the analysis results.
[0003] Low data processing efficiency: Facing the vast amount of system data, traditional processing methods are often difficult to process and analyze the data quickly and effectively. Especially in scenarios with high real-time requirements, traditional analysis methods may not be able to respond to changes in the system state in a timely manner, thus missing the best opportunity to discover problems.
[0004] Lack of intelligent support: Traditional system monitoring and analysis methods often lack intelligent support and are difficult to automatically extract potential problems and trends from complex data. This results in the system often being able to only respond passively when problems occur, rather than being able to issue early warnings and interventions through intelligent means.
[0005] The above limitations make it difficult for operation and maintenance personnel to quickly understand the system state and the location of problems, thus affecting the timely resolution of problems and making it difficult to meet the efficiency requirements of system operation. Summary of the Invention
[0006] Aiming at the problems in the prior art, this application provides an intelligent recognition method and device for system hot-spot problems, which can improve the efficiency and accuracy of system hot-spot problem recognition.
[0007] To solve at least one of the above problems, this application provides the following technical solutions: In a first aspect, this application provides an intelligent recognition method for system hot-spot problems, including: Obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to characterize system performance; Perform anomaly event correlation relationship mining operations on the historical system multi-source data on the same time axis according to a preset anomaly correlation algorithm to determine strongly correlated events corresponding to the anomaly events. Perform causal relationship reasoning operations on the strongly correlated events according to a set causal inference model to determine the corresponding causal chain of the strongly correlated events. Calculate the loss and perform iterative optimization operations on the root cause events obtained according to the causal chain of the strongly correlated events and the preset actual root cause event tags to determine the corresponding system hot spot identification model; Receive real-time system multi-source data, perform anomaly identification operations on the real-time system multi-source data according to preset system hot spot problem detection rules to determine the corresponding anomaly patterns, perform dynamic adjustment operations on the anomaly patterns according to the system hot spot identification model, and identify system hot spot problems according to the anomaly patterns obtained after the dynamic adjustment operations.
[0008] Further, the operation of standardizing the time stamp format of the historical system multi-source data and performing a comparison operation on the historical system multi-source data after the time stamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis includes: Perform format unification operations on the time stamps in the historical system multi-source data according to a preset time stamp format; And perform a comparison operation on the historical system multi-source data after the format unification operation to determine the corresponding historical system multi-source data on the same time axis.
[0009] Further, the operation of performing a comparison operation on the historical system multi-source data after the format unification operation to determine the corresponding historical system multi-source data on the same time axis includes: Perform deviation adjustment operations and time stamp sorting operations on the historical system multi-source data after the format unification operation to determine the corresponding sequential historical system multi-source data; Perform time zone division operations on the sequential historical system multi-source data according to a preset time window algorithm to determine the historical system multi-source data within corresponding multiple time regions; Perform synchronization alignment operations on the historical system multi-source data within the multiple time regions respectively, and perform missing value filling operations and association fusion operations on the historical system multi-source data after the synchronization alignment operation to determine the corresponding historical system multi-source data on the same time axis.
[0010] Further, the operation of performing anomaly event correlation relationship mining operations on the historical system multi-source data on the same time axis according to a preset anomaly correlation algorithm to determine strongly correlated events corresponding to the anomaly events includes: Perform recursive traversal operations on the historical system multi-source data on the same time axis to determine frequent item sets corresponding to the anomaly events; Perform a confidence calculation operation on the frequent item sets to determine the confidence of the corresponding frequent items; Perform a screening operation on the confidence of the frequent items according to a preset confidence threshold to determine the strong association events corresponding to the abnormal events.
[0011] Further, before performing a causal relationship reasoning operation on the strong association events according to the set causal inference model to determine the corresponding causal chain of the strong association events, it includes: Perform a variable definition operation on the multi-source data of the historical system on the same time axis according to the preset system performance influencing factors to determine the corresponding key variables, where the key variables include explicit variables and latent variables; Perform a causal relationship definition operation on the key variables, and perform a causal model construction operation according to the causal path obtained after the causal relationship definition operation to determine the corresponding initial causal model; Perform a model tuning operation on the initial causal model according to the preset goodness-of-fit criterion to determine the corresponding causal inference model.
[0012] Further, the performing a causal relationship definition operation on the key variables, and performing a causal model construction operation according to the causal path obtained after the causal relationship definition operation to determine the corresponding initial causal model includes: Perform an index construction operation on the latent variables in the key variables according to a preset causal relationship hypothesis graph to determine the measurement indexes corresponding to each latent variable; Perform a causal relationship definition operation according to the measurement indexes to determine the corresponding causal path, and perform a causal model construction operation according to the causal path to determine the corresponding initial causal model, where the causal path includes the causal path between latent variables and the causal path between latent variables and explicit variables.
[0013] Further, the performing a causal relationship reasoning operation on the strong association events according to the set causal inference model to determine the corresponding causal chain of the strong association events includes: Perform a causal relationship reasoning operation on the strong association events according to the set causal inference model to determine the causal order of the strong association events and the root cause events corresponding to the causal order; Determine the corresponding causal chain of the strong association events according to the causal order of the strong association events and the root cause events corresponding to the causal order.
[0014] In a second aspect, the present application provides an intelligent identification device for system hot issues, including: A multi-source data preprocessing module, which is used to obtain multi-source data of a historical system, perform timestamp format standardization operations on the multi-source data of the historical system, and perform comparison operations on the multi-source data of the historical system after the timestamp format standardization operations to determine corresponding multi-source data of the historical system on the same time axis, wherein the multi-source data of the historical system is used to reflect system performance; A hot-spot recognition model construction module, which is used to perform abnormal event correlation relationship mining operations on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm, determine strongly correlated events corresponding to abnormal events, perform causal relationship reasoning operations on the strongly correlated events according to a set causal inference model, determine corresponding strongly correlated event causal chains, perform loss calculation and iterative tuning operations on the root cause events obtained according to the strongly correlated event causal chains and preset actual root cause event tags, and determine a corresponding system hot-spot recognition model; A hot-spot recognition module, which is used to receive multi-source data of a real-time system, perform abnormal recognition operations on the multi-source data of the real-time system according to preset system hot-spot problem detection rules, determine corresponding abnormal patterns, perform dynamic adjustment operations on the abnormal patterns according to the system hot-spot recognition model, and identify system hot-spot problems according to the abnormal patterns obtained after the dynamic adjustment operations.
[0015] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the intelligent recognition method for system hot-spot problems are implemented.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the intelligent recognition method for system hot-spot problems are implemented.
[0017] In a fifth aspect, the present application provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the intelligent recognition method for system hot-spot problems are implemented.
[0018] As can be seen from the above technical solutions, the present application provides an intelligent recognition method and device for system hot issues. By unifying the timestamp formats of multi-source data of historical systems, comparing them on the same timeline, obtaining multi-source data of historical systems on the same timeline, inputting the data into a preset initial model, mining events strongly related to system performance anomalies in the multi-source data of the system through an association rule algorithm, and mapping them into a knowledge graph, performing causal relationship reasoning operations on the knowledge graph according to a causal inference model, obtaining the root cause events leading to system performance anomalies, calculating the loss between the root cause events and preset root cause event tags, obtaining a system hot issue recognition model; receiving multi-source data of a real-time system, identifying abnormal patterns according to system hot issue detection rules, and dynamically adjusting the abnormal patterns according to the system hot issue recognition model to determine system hot issues, thereby being able to improve the efficiency and accuracy of system hot issue recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is one of the flow diagrams of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 2 It is the second flow diagram of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 3 It is the third flow diagram of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 4 It is the fourth flow diagram of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 5 It is the fifth flow diagram of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 6 It is the sixth flow diagram of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 7 It is the seventh flow diagram of the intelligent recognition method for system hot issues in the embodiments of the present application; Figure 8 It is the structural diagram of the intelligent recognition device for system hot issues in the embodiments of the present application; Figure 9 It is the structural diagram of an electronic device in the embodiments of the present application.
[0021] Reference Signs: An electronic device 9600, a central processing unit 9100, a memory 9140, a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, a power supply 9170, a buffer memory 9141, an application / function storage unit 9142, a data storage unit 9143, a driver program storage unit 9144, an antenna 9111, a speaker 9131, and a microphone 9132. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0023] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0024] Considering that the existing system problem handling method makes it difficult for operation and maintenance personnel to quickly understand the system status and the location of problems, thus affecting the timely resolution of problems and making it difficult to meet the efficiency requirements of system operation. The present application provides an intelligent identification method and device for system hot issues. By unifying the timestamp formats of multi-source data of historical systems, comparing them on the same time axis, obtaining multi-source data of historical systems on the same time axis, and inputting the data into a preset initial model, events strongly related to system performance anomalies in the multi-source data of the system are mined through an association rule algorithm and mapped into a knowledge graph. Causal relationship reasoning operations are performed on the knowledge graph according to a causal inference model to obtain root cause events leading to system performance anomalies. Loss calculations are performed on the root cause events and preset root cause event tags to obtain a system hot issue identification model. Real-time multi-source data of the system is received, abnormal patterns are identified according to system hot issue detection rules, and the abnormal patterns are dynamically adjusted according to the system hot issue identification model to determine system hot issues, thereby improving the efficiency and accuracy of system hot issue identification.
[0025] To improve the efficiency and accuracy of system hot issue identification, an embodiment of an intelligent identification method for system hot issues is provided in the present application. Refer to Figure 1 The intelligent identification method for system hot issues specifically includes the following contents: Step S101: Obtain historical system multi-source data, perform timestamp format standardization operations on the historical system multi-source data, and perform comparison operations on the historical system multi-source data after the timestamp format standardization operations to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to characterize system performance; Optionally, in this embodiment, the purpose of this step is to first preprocess the multi-source data to ensure that data from different sources can be compared and analyzed on the same time axis.
[0026] Optionally, in this embodiment, the historical system multi-source data is to collect historical system data from different sources to ensure that multiple aspects of the system state (performance, network, logs, etc.) are covered, including but not limited to: System logs: Log information of system operation.
[0027] Performance monitoring tool data: CPU usage, memory occupancy, disk I / O, etc.
[0028] Network traffic data: Used to analyze the network behavior of the system.
[0029] Application sensor data: Outputs of, for example, temperature sensors, pressure sensors, etc.
[0030] Optionally, in this embodiment, for the timestamp format standardization operation, different data sources may use different time zones or time formats, resulting in inconsistent timestamps. To ensure that the data can be aligned in chronological order, the timestamps of all data need to be in a unified format. Preferably, in this embodiment, it is unified into the UTC timestamp format. UTC time is not affected by time zone and daylight saving time changes and is the standard time globally.
[0031] Optionally, in this embodiment, the historical system multi-source data on the same time axis is determined after synchronizing and comparing the data from different sources with standardized timestamps. The specific method is as follows: First, use the Kafka stream processing platform and time window technology to synchronize the data to ensure that the data within the same time window can be reasonably compared and integrated. It can be understood that the Kafka stream processing platform realizes the real-time transmission and processing of data, ensuring that the system can stably receive and process data regardless of the data scale. The time window technology divides the data into fixed-size time blocks, which can ensure the integrity and alignment of the data within each time window. Then, after completing the timestamp standardization and synchronization of the data, perform cross-data source comparison operations. The data with aligned timestamps can be associated to determine the multi-source data at the same time point, that is, the comparison of different data sources (such as log data, performance monitoring data, etc.) within the same time window. In this way, historical data from different systems or sensors can be integrated under the same time framework, facilitating analysis and decision-making.
[0032] For example, if the CPU usage rate is abnormally high during a certain period, data such as network traffic and memory occupancy can be viewed simultaneously to form a comprehensive view to determine the root cause of the problem.
[0033] Optionally, in this embodiment, the data after synchronization and comparison operations will be stored in the Spark big data platform for subsequent query and analysis. The big data platform Spark supports the execution of complex query and analysis models. To improve query efficiency, preferably, timestamps are used to create indexes to accelerate the query of data within a specific time range.
[0034] After the above step S101, the timestamp standardization and synchronization comparison of multi-source data are successfully performed, laying a solid data foundation for subsequent machine learning data patterns.
[0035] Step S102: Perform an abnormal event correlation relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm to determine strong correlation events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strong correlation events according to a set causal inference model to determine the corresponding causal chain of the strong correlation events, and perform a loss calculation and iterative optimization operation on the root cause events obtained according to the causal chain of the strong correlation events and a preset actual root cause event label to determine the corresponding system hotspot recognition model; Optionally, in this embodiment, this step trains a system hotspot recognition model, which is composed of an association analysis module, a graph database module, and a causal inference module.
[0036] Optionally, in this embodiment, the association analysis module aims to discover strong associations between events from historical data and automatically mine potential patterns in the data. First, scan the data set once to find frequent 1-itemsets and sort them in descending order of frequency to obtain a list L. Then, based on the list L, scan the data set again and process each original transaction: delete the items not in L and sort them in the order of L to obtain the modified transaction set T'. Next, construct an FP-tree, sort and link the data in T' according to the frequent items to form a tree with NULL as the root node. Record the support degree of each node at each node. Then, starting from the bottom (leaf node) of the tree and going up, recursively mine the conditional pattern bases for each node to find all frequent itemsets. Specifically, for each node, first find all its successor nodes (directly connected nodes), and then recursively mine each successor node. During the recursive process, it is necessary to continuously update the conditional pattern bases and conditional FP-trees of each node until no more frequent itemsets can be found. Finally, after identifying the frequent itemsets, discover strong association events by calculating the confidence. Specifically, for each frequent itemset, calculate the confidence between all the items it contains. The confidence represents the probability of the result item appearing under the condition that the premise item appears. According to the set confidence threshold, filter out the association rules with a confidence higher than the threshold, and these rules are the strong association events.
[0037] Optionally, in this embodiment, the graph database module constructs an event association graph through the Neo4j graph database to display the strong associations between different events. The graph shows the influence chains between events, helps identify which events are strongly associated with each other, and helps further locate the hot issues of the system.
[0038] Optionally, in this embodiment, the causal inference module uses the structural equation model (SEM) for causal reasoning to analyze the potential causal relationships between events. Specifically, the construction of the structural equation model (SEM): First, determine that the problem we want to solve is to identify the hot issues, abnormal patterns, and potential causal relationships of system performance. Then, under this framework, establish the theoretical relationships between latent variables and manifest variables.
[0039] Latent variables are variables that cannot be directly observed, such as system performance, abnormal patterns, and causal relationships. For example, "system performance" includes multiple latent variables (such as load, response time, etc.).
[0040] Manifest variables are data items that can be observed and directly measured, such as CPU usage rate, network latency, etc. collected through the monitoring system.
[0041] Define how observable variables reflect latent variables through data modeling, and establish measurement indicators for each latent variable. For example, use multiple observable variables (such as system load, response time, CPU usage) to measure the latent variable of system performance.
[0042] Then, define the causal relationships of the latent variables and set the causal paths. Among them, the causal paths between latent variables and between latent variables and observable variables need to be set according to the theoretical framework. For example, "event correlation" can be set as a latent variable, and abnormal events identified by machine learning or rule engines can be used as observable variables. For example, "system load" affects "response time", and "response time" in turn affects the latent variable of "user experience".
[0043] Next, for the defined initial causal model, use maximum likelihood estimation (ML) for model estimation, use the lavaan package (structural equation analysis package) in AMOS for model fitting, evaluate whether the fitting indicators meet the standards, and check the path coefficients and significance between latent variables and between latent variables and observable variables.
[0044] Preferably, use TLI>0.90 as the standard to evaluate the goodness of fit of the model, ensure that the coefficient of each path is significant, and verify the causal relationship between latent variables. According to the test results of the goodness of fit and path coefficients, adjust the structure of the model, add or delete paths until the model goodness of fit reaches the best state, and obtain a causal inference model. Based on the path coefficients of the causal inference model, conduct causal reasoning, analyze potential causal relationships, and explore how to improve system performance by adjusting certain variables.
[0045] For example, the system discovers strong correlations between events from multi-source data through an association rule algorithm. Through the mining of a large amount of historical data, the association rule algorithm can discover frequently co-occurring events and their relationships. Such as the frequent association between "high load" and "database connection failure". Then, use the Neo4j graph database to construct an event association graph, visually representing the relationships between various events. In the graph, nodes represent events, and edges represent the associations between events.
[0046] Event A: The database response time exceeds the preset threshold.
[0047] Event B: The system CPU usage exceeds 90%.
[0048] Event C: The server restarts.
[0049] In the graph, assume that a strong association is found between Event A and Event B (i.e., when the database response time is too long, the system CPU usage will also increase sharply), and there is also a causal relationship between Event B and Event C (too high CPU causes the system to restart). After the knowledge graph is constructed, the causal inference model is then applied to infer causal relationships. The causal inference model analyzes the potential causal chains between events and reveals which events are the root causes leading to the occurrence of other events. Based on the association relationships already identified in the graph, the causal inference model further analyzes whether there are causal relationships between events: when event A (long database response time) occurs, it directly causes event B (increased CPU usage), and event B in turn triggers event C (server restart). This causal chain describes the process of a system performance problem occurring.
[0050] It should be noted that when constructing the causal chain, the causal inference model not only considers the causal relationships between single events, but also combines other influencing factors (such as external environment, system load changes, etc.) to ensure the accuracy and reliability of the inference results.
[0051] Causal chain: Event A → Event B → Event C Event A: High concurrent requests lead to an increase in database response time.
[0052] Event B: The increase in database response time leads to an increase in CPU load.
[0053] Event C: Excessive CPU load ultimately leads to server restart.
[0054] Through the inference of this causal chain, the system finds the root cause of the server restart, that is, the "high concurrent requests" event. Subsequently, the system can give early warnings when "high concurrent requests" occur and predict subsequent problems based on the causal chain, thus providing improvement suggestions for operation and maintenance personnel.
[0055] Optionally, in this embodiment, after obtaining the root cause event output by the causal inference module, the initial model is trained according to the real root cause label. Specifically, the root cause of each event can be assigned a label according to its impact on the system state. For example, if an event causes the system to crash, its label can be marked as "abnormal", otherwise it is marked as "normal". The gap between the root cause event output by the model and the label is calculated through the cross-entropy loss function, and the backpropagation algorithm and gradient descent method are used to adjust the parameters. Through multiple iterations, the model is continuously optimized, gradually reducing the difference between the predicted label and the actual label.
[0056] It can be understood that in each iteration, the model will refine and adjust the causal chain of strongly associated events to identify more potential hot issues, optimize the system identification performance, and finally form an efficient and accurate system hot issue identification model for real-time identification of hot issues in the system through continuous iteration and tuning.
[0057] Step S103: Receive multi-source data of the real-time system, perform an anomaly recognition operation on the multi-source data of the real-time system according to the preset system hot issue detection rules, determine the corresponding anomaly pattern, perform a dynamic adjustment operation on the anomaly pattern according to the system hot issue recognition model, and identify the system hot issue according to the anomaly pattern obtained after the dynamic adjustment operation.
[0058] Optionally, in this step, a unified data format is formulated through the API (data interface), and the rule engine and the machine learning model obtain the same data input through the interface.
[0059] After receiving the real-time data, the system will perform a preliminary anomaly recognition according to the preset rules. The rule engine (Drools) will perform rule matching by checking whether the data conforms to a specific anomaly pattern. At this stage, the system detects the known anomaly patterns through the rule library. For example, it is stipulated that "if the server CPU usage rate exceeds 90% and lasts for more than 5 minutes, an anomaly alarm will be triggered". These patterns are defined according to historical data, system experience rules, and manually set standards. When the server CPU usage rate reaches 90%, the rule engine will immediately trigger an alarm.
[0060] At the same time, for the anomalies not recorded in the rule engine, they are identified by the hot issue recognition model trained in step S102 of ours. For example, within a specific time period, although the CPU usage rate of each server does not exceed 90%, the mutual dependence between different servers leads to a decline in the overall system performance, ultimately affecting the user experience. The hot issue recognition model can discover some potential correlation patterns by training historical data. For example, certain specific events (such as promotional activities, high-concurrency requests) will cause resource contention among multiple service nodes, resulting in system crashes and performance anomalies.
[0061] It should be noted that in addition to the immediate anomaly patterns, the hot issue recognition model can also identify the gradually changing anomaly trends. For example, although the single CPU usage rate does not exceed 90%, in the past few months, the CPU usage rate has been continuously rising, resulting in a gradual deterioration of performance. At this time, the system needs to monitor and identify this anomaly trend in order to make timely adjustments. Through time series analysis and causal inference models, the system can discover the gradual upward trend of performance indicators. For example, the continuous increase in CPU usage rate is due to overloaded system, inefficient certain code, or uneven resource allocation. This trend is identified in the real-time monitoring of the system, providing a basis for the early warning system.
[0062] Next, based on the results of the second anomaly pattern (complex pattern recognition) and anomaly trend (trend analysis), the system dynamically adjusts the first anomaly pattern (known rule anomaly pattern). For example, if the hot issue recognition model finds that the impact of a specific event or trend is significant, it can adjust the thresholds or policies in the rule engine so that the alarm system can respond to system problems more promptly and accurately.
[0063] For example, the original rule was to trigger an alarm when the CPU usage exceeded 90% and lasted for 5 minutes. However, through the analysis of trends and complex patterns, it was adjusted to: trigger an alarm immediately when the CPU usage exceeds 85% and lasts for 3 minutes, or when certain specific events (such as high concurrent access) occur.
[0064] After dynamic adjustment, the system can more accurately identify potential hot issues. For example, through the dynamically adjusted rule engine, the system can timely detect the increase in CPU usage and associate it with other abnormal events (such as database response latency), and finally identify the hot issue of the system - the performance bottleneck of the database leads to the overuse of server resources.
[0065] Finally, through the integrated visualization tool, the system visually displays the hot issues to the operation and maintenance personnel in the form of charts, heat maps, etc., to help them quickly locate and handle the problems.
[0066] This example shows how this embodiment collaborates with the rule engine and the hot issue recognition model to perform real-time analysis of the influencing factors of system anomalies, quickly and accurately identify hot issues, and find the root causes affecting system performance, thereby improving the efficiency and accuracy of system hot issue recognition.
[0067] As can be seen from the above description, the intelligent recognition method for system hot issues provided by the embodiment of the present application can unify the timestamp formats of multi-source data of historical systems, compare them on the same time axis to obtain multi-source data of historical systems on the same time axis, input the data into a preset initial model, mine events strongly related to system performance anomalies in the multi-source data of the system through an association rule algorithm, map them into a knowledge graph, perform causal relationship reasoning operations on the knowledge graph according to a causal inference model to obtain the root cause events leading to system performance anomalies, calculate the loss between the root cause events and the preset root cause event labels to obtain a system hot issue recognition model; receive real-time multi-source data of the system, identify anomaly patterns according to system hot issue detection rules, dynamically adjust the anomaly patterns according to the system hot issue recognition model, and determine system hot issues, thereby improving the efficiency and accuracy of system hot issue recognition.
[0068] In an embodiment of the intelligent recognition method for system hot issues of the present application, refer to Figure 2, it may also specifically include the following content: Step S201: Perform a format unification operation on the timestamps in the historical system multi-source data according to a preset timestamp format; Step S202: And perform a comparison operation on the historical system multi-source data after the format unification operation to determine the corresponding historical system multi-source data on the same time axis.
[0069] Optionally, to ensure that the data can be aligned in chronological order, the timestamps of all data need to be in a unified format. Preferably, in this embodiment, it is unified into the UTC timestamp format. UTC time is not affected by time zone and daylight saving time changes and is the standard time globally.
[0070] Specifically, different data sources may use different time formats. For example: UNIX timestamp (the number of seconds / milliseconds represents the number of seconds since January 1, 1970) ISO 8601 format (2024-12-23T12:34:56Z) Custom time format (yyyy-mm-dd hh:mm:ss) To ensure that the timestamps of all data sources are in the same format, we perform a standardized conversion of the timestamp format in the data access platform or data processing pipeline to convert all timestamps into a unified timestamp format.
[0071] Next, to perform multi-source data comparison, we perform deviation adjustment and sorting on the multi-source data after timestamp unification. Specifically, since there will be time deviations in data from different sources (such as due to network latency, clock drift, etc.), we adopt a synchronization strategy. At the data access stage, according to the transmission time of the data packet, we adjust the received timestamp. For example, when receiving a data packet, we can record the current reception time and calculate the transmission delay to adjust the timestamp. At the same time, there is clock drift in different devices, resulting in inconsistent timestamps. We use the NTP (Network Time Protocol) protocol to synchronize the system time of each node regularly to ensure the consistency of timestamps. After deviation adjustment, the multi-source data is sorted and stored in timestamp order.
[0072] To align the sorted multi-source data on the time axis for easy multi-source data comparison, we group events into time windows, for example, classify events by hour, by minute, or by day, so as to unify the data within the same time period onto the same time axis. For the case where there are gaps between timestamps, a linear interpolation algorithm is used to fill the gap time periods.
[0073] If there is a certain delay in the timestamps of some data sources, set a reasonable delay tolerance threshold to ensure data consistency. For example, allow a certain delay time, preferably 0.3 seconds. As long as the data is still within the acceptable timeliness range, synchronization can be performed.
[0074] Through step S202, this embodiment obtains multi-source data on the same time axis, which is convenient for subsequent comparison between multi-source data and lays a solid data foundation for the machine learning algorithm to discover the associations between multi-source data.
[0075] In an embodiment of the intelligent identification method for system hot issues of this application, refer to Figure 3 , and it may specifically include the following content: Step S301: Perform deviation adjustment operation and timestamp sorting operation on the historical system multi-source data after the format unification operation, and determine the corresponding sequential historical system multi-source data; Step S302: Perform time zone division operation on the sequential historical system multi-source data according to the preset time window algorithm, and determine the historical system multi-source data within the corresponding multiple time regions; Step S303: Perform synchronization alignment operations on the historical system multi-source data within the multiple time regions respectively, and perform missing data filling operation and association fusion operation on the historical system multi-source data after the synchronization alignment operation, and determine the historical system multi-source data on the same time axis.
[0076] Optionally, in this embodiment, this step is a refinement of the step of comparing the historical multi-source data of the system after timestamp synchronization. Specifically, after the timestamps are unified, our multi-source data already has a unified timestamp format that can be used for comparison. At this time, when receiving data, perform deviation adjustment and sorting according to the unified timestamp to obtain the sequential historical system multi-source data.
[0077] Specifically, in the data access stage, the deviation adjustment is to adjust the received timestamp according to the time of data packet transmission. For example, when receiving a data packet, the current reception time can be recorded, and the transmission delay can be calculated to adjust the timestamp. At the same time, there are drifts in the clocks of different devices, resulting in inconsistent timestamps. We use the NTP (Network Time Protocol) protocol to regularly synchronize the system time of each node to ensure the consistency of timestamps. When data flows into the system, a buffer is used for temporary storage to facilitate batch processing and analysis of the data. Sort the data in the buffer according to the standardized timestamp to ensure that all events are arranged in chronological order in subsequent processing steps.
[0078] Next, perform data processing on the received sequential historical system multi-source data, including alignment between multi-source data. Specifically, use the time window technique to divide the data into time intervals (such as every minute or every hour). Within each time window, use timestamps to align the data. Use interpolation methods to identify and fill in missing data points found during the synchronization process to ensure the integrity of the data in each time window. Based on the alignment, associate and fuse the data from different sources according to time windows for subsequent unified analysis.
[0079] Through step S303, this embodiment obtains multi-source data on the same time axis, facilitating subsequent comparative analysis between multi-source data and laying a solid data foundation for the machine learning algorithm to discover the associations between multi-source data.
[0080] In an embodiment of the intelligent identification method for system hot issues in this application, refer to Figure 4 , and it may specifically include the following content: Step S401: Perform a recursive traversal operation on the multi-source data of the historical system on the same time axis to determine the frequent item sets corresponding to abnormal events; Step S402: Perform a confidence calculation operation on the frequent item sets to determine the corresponding frequent item confidence; Step S403: Screen the frequent item confidence according to a preset confidence threshold to determine the strong association events corresponding to abnormal events.
[0081] Optionally, in this embodiment, the purpose of step S401 is to find the frequent item sets with high correlation. First, scan the data set once to find the frequent 1-item sets and sort them in descending order of frequency to obtain list L. Then, based on list L, scan the data set again and process each original transaction: delete the items not in L and arrange them in the order of L to obtain the modified transaction set T'. Next, construct an FP-tree, sort and link the data in T' according to the frequent items to form a tree with NULL as the root node. Record the support degree of each node at each node. Then, starting from the bottom (leaf nodes) of the tree and going up, perform recursive mining of the conditional pattern basis for each node to find all the frequent item sets. Specifically, for each node, first find all its successor nodes (directly connected nodes), and then perform recursive mining on each successor node. During the recursive process, it is necessary to continuously update the conditional pattern basis and conditional FP-tree of each node until no more frequent item sets can be found.
[0082] Optionally, in this embodiment, the purpose of steps S402 and S403 is to find events strongly associated with system exception events based on frequent itemsets. Specifically, for each frequent itemset, the confidence levels between all the items it contains are calculated. The confidence level represents the probability of the result item occurring under the condition that the premise item appears. According to the set confidence threshold, the association rules with confidence levels higher than the threshold are filtered out, and these rules are the strongly associated events.
[0083] For example, assume the events collected by the system include: Event 1: CPU usage exceeds 90%.
[0084] Event 2: Memory usage exceeds 85%.
[0085] Event 3: Network traffic exceeds the predetermined threshold.
[0086] Using the association algorithm, the following associations will be found: There is a strong association between Event 1 and Event 2, that is, high CPU usage is usually accompanied by increased memory occupancy.
[0087] There is also a strong association between Event 2 and Event 3. When memory occupancy is high, network traffic is usually large.
[0088] Through step S403, this embodiment successfully discovers events strongly associated with system exception events and the relationships between events from the multi-source data of the historical system on the same time axis, laying a foundation for finding the root cause event through causal inference later.
[0089] In an embodiment of the intelligent identification method for system hot issues in this application, refer to Figure 5 , and it may specifically include the following content: Step S501: Perform variable definition operations on the multi-source data of the historical system on the same time axis according to preset system performance influencing factors to determine corresponding key variables, where the key variables include explicit variables and latent variables; Step S502: Perform causal relationship definition operations on the key variables, and perform causal model construction operations according to the causal paths obtained after the causal relationship definition operations to determine corresponding initial causal models; Step S503: Perform model tuning operations on the initial causal model according to preset goodness-of-fit criteria to determine corresponding causal inference models.
[0090] Optionally, this step is the construction process of the causal inference model. First, it is determined that the problem to be solved is to identify hot issues, abnormal patterns, and potential causal relationships of system performance. Then, within this framework, the key variables affecting system performance are defined, including latent variables and explicit variables.
[0091] Latent variables are variables that cannot be directly observed, such as system performance, abnormal patterns, causal relationships, etc. For example, "system performance" includes multiple latent variables (such as load, response time, etc.).
[0092] Manifest variables are data items that can be observed and directly measured. For example, CPU usage rate, network latency, etc. collected through a monitoring system.
[0093] Build an initial causal model through data modeling to define how manifest variables reflect latent variables and establish measurement indicators for each latent variable. For example, a latent variable such as system performance is measured through multiple manifest variables (such as system load, response time, CPU usage rate).
[0094] Then, define the causal relationships of the latent variables and set the causal paths. Among them, the causal paths between latent variables and between latent variables and manifest variables need to be set according to the theoretical framework. For example, "event correlation" can be set as a latent variable, and abnormal events identified through machine learning or rule engines can be used as manifest variables. For example, "system load" affects "response time", and "response time" in turn affects the latent variable "user experience".
[0095] Next, for the defined initial causal model, use maximum likelihood estimation (ML) for model estimation, use the lavaan package (structural equation analysis package) in AMOS for model fitting, evaluate whether the fitting indicators meet the standards, and check the path coefficients and significance between latent variables and between latent variables and manifest variables.
[0096] Preferably, the goodness of fit uses the TLI>0.90 standard to evaluate the goodness of fit of the model, ensure that the coefficient of each path is significant, and verify the causal relationship between latent variables. According to the test results of the goodness of fit and path coefficients, adjust the structure of the model, add or delete paths until the model goodness of fit reaches the best state, and obtain a causal inference model.
[0097] Through step S503, this embodiment successfully obtains a causal inference model, laying a solid foundation for finding the root cause of system performance abnormal events.
[0098] In an embodiment of the intelligent identification method for system hot issues of the present application, refer to Figure 6 , and it may specifically include the following content: Step S601: Perform an index construction operation on the latent variables in the key variables according to a preset causal relationship hypothesis diagram to determine the measurement indicators corresponding to each latent variable; Step S602: Perform a causal relationship definition operation based on the measurement metrics to determine the corresponding causal paths, and perform a causal model construction operation based on the causal paths to determine the corresponding initial causal model, where the causal paths include causal paths between latent variables and causal paths between latent variables and manifest variables.
[0099] Optionally, in this embodiment, the causal relationship hypothesis graph is a graphical representation method used to display the causal relationships between different variables. This graph consists of latent variables and manifest variables. By defining the relationships between latent variables and manifest variables, it helps to infer the causal paths between these variables and guides the subsequent construction of the causal model. In the causal relationship hypothesis graph, latent variables need to be quantified through some measurement metrics. The measurement metrics are manifest variables related to the latent variables and can reflect the characteristics of the latent variables. Based on prior domain knowledge and data analysis, variables that can accurately measure the latent variables are selected. For example, if the latent variable is "system load", then the corresponding measurement metrics are "CPU usage rate", "memory occupancy rate", etc.
[0100] Optionally, in this embodiment, the causal path is a causal chain between latent variables and manifest variables or between latent variables.
[0101] Causal path between latent variables: There is an interaction relationship between latent variables. For example, one latent variable (such as "network latency") affects another latent variable (such as "system response time").
[0102] Causal path between latent variable and manifest variable: The manifest variable, as a measurement metric of the latent variable, affects the latent variable. For example, the CPU usage rate (manifest variable) can be used as a measurement metric for the system load (latent variable) and in turn affects the system performance (latent variable).
[0103] By transforming the above causal relationship paths into a structural equation model to describe the relationships between variables, after the preliminary construction of the model, testing and adjustment can be carried out until the model can effectively describe the causal relationships in the data.
[0104] Through step S602, this embodiment successfully obtains the initial causal model, laying a solid foundation for subsequent optimization of the initial model to obtain a causal inference model for finding the root causes affecting system performance.
[0105] In an embodiment of the intelligent identification method for system hot issues in this application, refer to Figure 7 , and it may specifically include the following content: Step S701: Perform a causal relationship reasoning operation on the strongly correlated events according to the set causal inference model to determine the causal order of the strongly correlated events and the root cause events corresponding to the causal order; Step S702: Determine the corresponding strong correlation event causal chain according to the causal order of the strong correlation events and the root cause events corresponding to the causal order.
[0106] Optionally, in this embodiment, the events and their association relationships in the graph database are used as inputs, and a causal inference model is used for causal reasoning to obtain a causal chain. The goal of causal inference is to identify which events are the root causes that lead to other events. For example, whether the excessive memory usage directly causes the high CPU usage rate, or is affected by other factors.
[0107] For example, assume that the knowledge graph records events A, B, C and their relationships: Event A: The database response time exceeds the preset threshold.
[0108] Event B: The system CPU usage rate exceeds 90%.
[0109] Event C: The server restarts.
[0110] Relationship: There is a strong correlation between Event A and Event B, and there is also a causal relationship between Event B and Event C.
[0111] Based on the association relationships already identified in the graph, the causal inference model further analyzes whether there is a causal order between events: When Event A (the database response time is too long) occurs, it directly causes Event B (the CPU usage rate increases), and Event B triggers Event C (the server restarts). This causal chain describes the process of a system performance problem occurring.
[0112] It should be noted that when constructing the causal chain, the causal inference model not only considers the causal relationships between single events, but also combines other influencing factors (such as the external environment, system load changes, etc.) to ensure the accuracy and reliability of the inference results.
[0113] Causal chain: Event A → Event B → Event C Event A: High concurrent requests cause an increase in the database response time.
[0114] Event B: The increase in the database response time causes an increase in the CPU load.
[0115] Event C: The excessive CPU load ultimately causes the server to restart.
[0116] Through the reasoning of this causal chain, the system finds the root cause of the server restart, that is, the "high concurrent request" event. Subsequently, the system can give an early warning when a "high concurrent request" occurs and provide improvement suggestions for the operation and maintenance personnel according to the subsequent problems predicted by the causal chain.
[0117] Through step S702, the present embodiment successfully obtains the system performance causal chain, and through the causal chain, the root cause factors affecting the system performance can be effectively found, increasing the efficiency of identifying hot issues in the system.
[0118] To improve the efficiency and accuracy of identifying hot issues in the system, the present application provides an embodiment of an intelligent identification device for system hot issues that implements all or part of the content of the intelligent identification method for the system hot issues. Refer to Figure 8 The intelligent identification device for system hot issues specifically includes the following content: The multi-source data preprocessing module 10 is configured to obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to characterize the system performance; The hot issue identification model construction module 20 is configured to perform an abnormal event correlation relationship mining operation on the historical system multi-source data on the same time axis according to a preset abnormal correlation algorithm, determine strong correlation events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strong correlation events according to a set causal inference model, determine the corresponding strong correlation event causal chain, perform a loss calculation and iterative tuning operation on the root cause event obtained according to the strong correlation event causal chain and a preset actual root cause event label, and determine the corresponding system hot issue identification model; The hot issue identification module 30 is configured to receive real-time system multi-source data, perform an abnormal identification operation on the real-time system multi-source data according to a preset system hot issue detection rule, determine the corresponding abnormal pattern, perform a dynamic adjustment operation on the abnormal pattern according to the system hot issue identification model, and identify system hot issues according to the abnormal pattern obtained after the dynamic adjustment operation to determine the corresponding system hot issues.
[0119] As can be seen from the above description, the intelligent identification device for system hot issues provided by the embodiment of the present application can unify the timestamp format of historical system multi-source data, perform a comparison on the same time axis to obtain the historical system multi-source data on the same time axis, input the data into a preset initial model, mine events strongly related to system performance anomalies in the system multi-source data through an association rule algorithm, map them into a knowledge graph, perform a causal relationship reasoning operation on the knowledge graph according to a causal inference model, obtain the root cause event leading to system performance anomalies, perform a loss calculation on the root cause event and a preset root cause event label to obtain a system hot issue identification model; receive real-time system multi-source data, identify abnormal patterns according to the system hot issue detection rule, dynamically adjust the abnormal patterns according to the system hot issue identification model, and determine system hot issues, thereby improving the efficiency and accuracy of identifying system hot issues.
[0120] At the hardware level, to improve the efficiency and accuracy of system hot issue identification, an embodiment of an electronic device for implementing all or part of the intelligent identification method for the system hot issues provided by this application includes the following specific content: A processor, a memory, a communications interface, and a bus; wherein, the processor, the memory, and the communications interface complete communication with each other through the bus; the communications interface is used to implement information transmission between the intelligent identification method for system hot issues and related devices such as a core business system, a user terminal, and a related database, etc.; this logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., and this embodiment is not limited thereto. In this embodiment, this logic controller can be implemented with reference to the embodiments of the intelligent identification method for system hot issues in the embodiments, as well as the embodiments of the intelligent identification method for system hot issues, and its content is incorporated herein, and repeated parts will not be elaborated again.
[0121] It can be understood that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0122] In practical applications, part of the intelligent identification method for system hot issues can be executed on the electronic device side as described above, or all operations can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor.
[0123] The above-mentioned client device may have a communication module (i.e., a communication unit), and can be communicatively connected to a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and may also include a server on an intermediate platform in other implementation scenarios, such as a server on a third-party server platform communicatively linked to the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.
[0124] Figure 9 It is a schematic block diagram of the system composition of the electronic device 9600 according to an embodiment of this application. As Figure 9As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It should be noted that this Figure 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0125] In one embodiment, the function of the intelligent identification method for system hot - spot problems may be integrated into the central processing unit 9100. Among them, the central processing unit 9100 may be configured to perform the following controls: Step S101: Obtain historical system multi - source data, perform a timestamp format standardization operation on the historical system multi - source data, and perform a comparison operation on the historical system multi - source data after the timestamp format standardization operation to determine the corresponding historical system multi - source data on the same time axis, where the historical system multi - source data is used to characterize system performance; Step S102: Perform an abnormal event correlation relationship mining operation on the historical system multi - source data on the same time axis according to a preset abnormal correlation algorithm to determine strong correlation events corresponding to abnormal events, perform a causal relationship reasoning operation on the strong correlation events according to a set causal inference model to determine the corresponding causal chain of strong correlation events, perform a loss calculation and iterative optimization operation on the root cause event obtained according to the causal chain of strong correlation events and a preset actual root cause event label to determine the corresponding system hot - spot identification model; Step S103: Receive real - time system multi - source data, perform an abnormal identification operation on the real - time system multi - source data according to a preset system hot - spot problem detection rule to determine the corresponding abnormal pattern, perform a dynamic adjustment operation on the abnormal pattern according to the system hot - spot identification model, and identify system hot - spot problems according to the abnormal pattern obtained after the dynamic adjustment operation to determine the corresponding system hot - spot problems.
[0126] As can be seen from the above description, the electronic device provided in the embodiment of the present application unifies the timestamp format of historical system multi - source data, compares on the same time axis to obtain historical system multi - source data on the same time axis, inputs the data into a preset initial model, mines events strongly related to system performance anomalies in the system multi - source data through an association rule algorithm and maps them into a knowledge graph, performs a causal relationship reasoning operation on the knowledge graph according to a causal inference model to obtain the root cause event leading to system performance anomalies, performs a loss calculation on the root cause event and a preset root cause event label to obtain a system hot - spot identification model; receives real - time system multi - source data, identifies abnormal patterns according to system hot - spot problem detection rules, dynamically adjusts the abnormal patterns according to the system hot - spot identification model, and determines system hot - spot problems, thereby improving the efficiency and accuracy of system hot - spot problem identification.
[0127] In another embodiment, the intelligent identification method for system hot issues can be separately configured from the central processing unit 9100. For example, the intelligent identification method for system hot issues can be configured as a chip connected to the central processing unit 9100, and the function of the intelligent identification method for system hot issues can be realized through the control of the central processing unit.
[0128] As Figure 9 shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It should be noted that the electronic device 9600 does not necessarily have to include Figure 9 all the components shown in Figure 9 ; in addition, the electronic device 9600 may further include
[0129] As Figure 9 shown, the central processing unit 9100, sometimes also referred to as a controller or operation control, may include a microprocessor or other processor devices and / or logic devices. The central processing unit 9100 receives inputs and controls the operations of the various components of the electronic device 9600.
[0130] Among them, the memory 9140, for example, may be one or more of a buffer, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory, or other suitable devices. The above information related to failures can be stored, and in addition, programs for executing relevant information can also be stored. And the central processing unit 9100 can execute the program stored in the memory 9140 to implement information storage or processing, etc.
[0131] The input unit 9120 provides inputs to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to supply power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.
[0132] The memory 9140 may be a solid-state memory. For example, a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be such a memory that stores information even when powered off, can be selectively erased and has more data. Examples of such a memory are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, and the application / function storage unit 9142 is used to store application programs and function programs or the processes for operating the electronic device 9600 through the central processing unit 9100.
[0133] The memory 9140 may further include a data storage unit 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (such as a messaging application, an address book application, etc.).
[0134] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as in the case of a conventional mobile communication terminal.
[0135] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 is also coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby implementing normal telecommunication functions. The audio processor 9130 may include any suitable buffers, decoders, amplifiers, etc. Additionally, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine via the microphone 9132, and the sound stored on the local machine can be played via the speaker 9131.
[0136] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps in the intelligent recognition method for system hot issues where the execution subject in the above embodiments is a server or a client. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements all steps of the intelligent recognition method for system hot issues where the execution subject in the above embodiments is a server or a client. For example, when the processor executes the computer program, the following steps are implemented: Step S101: Obtain historical system multi-source data, perform a timestamp format standardization operation on the historical system multi-source data, and perform a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, where the historical system multi-source data is used to characterize system performance; Step S102: Perform an abnormal event correlation relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm, determine strongly correlated events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strongly correlated events according to a set causal inference model, determine the corresponding causal chain of the strongly correlated events, perform a loss calculation and iterative optimization operation on the root cause event obtained according to the causal chain of the strongly correlated events and a preset actual root cause event label, and determine the corresponding system hotspot identification model; Step S103: Receive multi-source data of the real-time system, perform an abnormal identification operation on the multi-source data of the real-time system according to a preset system hotspot problem detection rule, determine the corresponding abnormal pattern, perform a dynamic adjustment operation on the abnormal pattern according to the system hotspot identification model, and identify a system hotspot problem according to the abnormal pattern obtained after the dynamic adjustment operation.
[0137] As can be seen from the above description, the computer-readable storage medium provided by the embodiment of the present application unifies the timestamp formats of the multi-source data of the historical system, compares them on the same time axis to obtain the multi-source data of the historical system on the same time axis, and inputs the data into a preset initial model. By using an association rule algorithm to mine events strongly related to system performance anomalies in the multi-source data of the system and map them into a knowledge graph, a causal relationship reasoning operation is performed on the knowledge graph according to a causal inference model to obtain the root cause event leading to system performance anomalies, and a loss calculation is performed on the root cause event and a preset root cause event label to obtain a system hotspot identification model; receive multi-source data of the real-time system, identify an abnormal pattern according to a system hotspot problem detection rule, dynamically adjust the abnormal pattern according to the system hotspot identification model, and determine a system hotspot problem, thereby improving the efficiency and accuracy of identifying system hotspot problems.
[0138] An embodiment of the present application further provides a computer program product capable of implementing all steps in the intelligent identification method for system hotspot problems with the execution subject being a server or a client in the above embodiment. When the computer program / instructions are executed by a processor, the steps of the intelligent identification method for system hotspot problems are implemented. For example, the computer program / instructions implement the following steps: Step S101: Obtain multi-source data of the historical system, perform a timestamp format standardization operation on the multi-source data of the historical system, and perform a comparison operation on the multi-source data of the historical system after the timestamp format standardization operation to determine the corresponding multi-source data of the historical system on the same time axis, where the multi-source data of the historical system is used to characterize system performance; Step S102: Perform an abnormal event correlation relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal correlation algorithm to determine strongly correlated events corresponding to the abnormal events. Perform a causal relationship reasoning operation on the strongly correlated events according to a set causal inference model to determine the corresponding causal chain of the strongly correlated events. Perform a loss calculation and iterative optimization operation on the root cause events obtained according to the causal chain of the strongly correlated events and preset actual root cause event labels to determine the corresponding system hot spot identification model; Step S103: Receive multi-source data of the real-time system, perform an abnormal identification operation on the multi-source data of the real-time system according to a preset system hot spot problem detection rule to determine the corresponding abnormal pattern. Perform a dynamic adjustment operation on the abnormal pattern according to the system hot spot identification model, and identify system hot spot problems according to the abnormal pattern obtained after the dynamic adjustment operation.
[0139] As can be seen from the above description, the computer program product provided by the embodiment of the present application unifies the time stamp format of the multi-source data of the historical system, compares them on the same time axis to obtain the multi-source data of the historical system on the same time axis, and inputs the data into a preset initial model. By using an association rule algorithm to mine events strongly related to system performance anomalies in the multi-source data of the system and map them into a knowledge graph, perform a causal relationship reasoning operation on the knowledge graph according to a causal inference model to obtain the root cause events leading to system performance anomalies, and perform a loss calculation on the root cause events and preset root cause event labels to obtain a system hot spot identification model; receive multi-source data of the real-time system, identify abnormal patterns according to the system hot spot problem detection rule, dynamically adjust the abnormal patterns according to the system hot spot identification model, and determine system hot spot problems, thereby improving the efficiency and accuracy of identifying system hot spot problems.
[0140] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.
[0144] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An intelligent identification method for system hot spots, characterized in that: The method comprises: Acquire multi-source data of a historical system, perform a timestamp format standardization operation on the multi-source data of the historical system, and perform a comparison operation on the multi-source data of the historical system after the timestamp format standardization operation to determine corresponding multi-source data of the historical system on the same time axis, wherein the multi-source data of the historical system is used to characterize system performance; According to the preset abnormal association algorithm, the abnormal event association relationship mining operation is performed on the multi-source data of the historical system on the same time axis to determine the strongly associated events corresponding to the abnormal events, and the causal relationship reasoning operation is performed on the strongly associated events according to the set causal inference model to determine the corresponding strongly associated event causal chain, and the root cause event obtained according to the strongly associated event causal chain is compared with the preset actual root cause event label to perform loss calculation and iterative tuning operations to determine the corresponding system hot spot identification model; Receive real-time system multi-source data, perform anomaly recognition operations on the real-time system multi-source data according to preset system hot spot problem detection rules, determine corresponding abnormal patterns, dynamically adjust the abnormal patterns according to the system hot spot recognition model, identify system hot spot problems according to the abnormal patterns obtained after the dynamic adjustment operations, and determine corresponding system hot spot problems.
2. The intelligent identification method of system hot spots according to claim 1 is characterized in that: The performing of a timestamp format standardization operation on the historical system multi-source data, and performing a comparison operation on the historical system multi-source data after the timestamp format standardization operation to determine the corresponding historical system multi-source data on the same time axis, includes: Performing a format unification operation on the timestamps in the multi-source data of the historical system according to a preset timestamp format; A comparison operation is performed on the multi-source data of the historical system after the format unification operation to determine the corresponding multi-source data of the historical system on the same time axis.
3. The intelligent identification method of system hot spots according to claim 2 is characterized in that: The comparing operation is performed on the multi-source data of the historical system after the format unification operation to determine the corresponding multi-source data of the historical system on the same time axis, including: and performing deviation adjustment operation and timestamp sorting operation on the multi-source data of the historical system after the format unification operation, to determine the corresponding sequence of the multi-source data of the historical system; Performing a time zone division operation on the sequential historical system multi-source data according to a preset time window algorithm to determine the historical system multi-source data within corresponding multiple time zones; The historical system multi-source data in the multiple time zones are synchronously aligned respectively, and the missing filling operation and the associated fusion operation are performed on the historical system multi-source data after the synchronous alignment operation to determine the corresponding historical system multi-source data on the same time axis.
4. The intelligent identification method of system hot spots according to claim 1 is characterized in that: The method of performing an abnormal event association relationship mining operation on the multi-source data of the same time axis historical system according to a preset abnormal association algorithm to determine a strongly associated event corresponding to the abnormal event includes: Recursively traverse the multi-source data of the historical system on the same time axis to determine frequent item sets corresponding to abnormal events; Performing a confidence calculation operation on the frequent item set to determine the corresponding frequent item confidence; The confidence of the frequent items is screened according to a preset confidence threshold to determine a strongly associated event corresponding to the abnormal event.
5. The intelligent identification method of system hot spots according to claim 1 is characterized in that: Before performing causal relationship reasoning operations on the strongly correlated events according to the set causal inference model to determine the corresponding causal chain of the strongly correlated events, the method includes: According to the preset system performance influencing factors, a variable definition operation is performed on the multi-source data of the historical system on the same time axis to determine the corresponding key variables, wherein the key variables include explicit variables and latent variables; Performing a causal relationship definition operation on the key variables, performing a causal model construction operation according to the causal path obtained after the causal relationship definition operation, and determining a corresponding initial causal model; The initial causal model is subjected to a model tuning operation according to a preset fitness standard to determine a corresponding causal inference model.
6. The intelligent identification method of system hot spots according to claim 5 is characterized in that: The performing a causal relationship definition operation on the key variables, performing a causal model construction operation according to the causal path obtained after the causal relationship definition operation, and determining the corresponding initial causal model includes: Performing an indicator construction operation on the latent variables in the key variables according to a preset causal relationship hypothesis diagram, and determining a measurement indicator corresponding to each of the latent variables; A causal relationship definition operation is performed according to the measurement indicator to determine the corresponding causal path, and a causal model construction operation is performed according to the causal path to determine the corresponding initial causal model, wherein the causal path includes a causal path between latent variables and a causal path between latent variables and explicit variables.
7. The intelligent identification method of system hot spots according to claim 1 is characterized in that: The performing causal relationship reasoning operations on the strongly correlated events according to the set causal inference model to determine the corresponding causal chain of the strongly correlated events includes: Performing causal relationship reasoning operations on the strongly correlated events according to a set causal inference model to determine a causal sequence of the strongly correlated events and a root cause event corresponding to the causal sequence; A corresponding causal chain of strongly associated events is determined according to the causal sequence of the strongly associated events and the root cause events corresponding to the causal sequence.
8. An intelligent identification device for system hot spots, characterized in that: The device comprises: A multi-source data preprocessing module is used to obtain multi-source data of a historical system, perform a timestamp format standardization operation on the multi-source data of the historical system, and perform a comparison operation on the multi-source data of the historical system after the timestamp format standardization operation to determine corresponding multi-source data of the historical system on the same time axis, wherein the multi-source data of the historical system is used to reflect system performance; A hotspot identification model construction module is used to perform an abnormal event association relationship mining operation on the multi-source data of the historical system on the same time axis according to a preset abnormal association algorithm, determine the strongly associated events corresponding to the abnormal events, perform a causal relationship reasoning operation on the strongly associated events according to a set causal inference model, determine the corresponding strongly associated event causal chain, perform loss calculation and iterative tuning operations on the root cause events obtained according to the strongly associated event causal chain and the preset actual root cause event labels, and determine the corresponding system hotspot identification model; The hotspot identification module is used to receive multi-source data of a real-time system, perform anomaly identification operations on the multi-source data of the real-time system according to preset system hotspot problem detection rules, determine the corresponding abnormal pattern, dynamically adjust the abnormal pattern according to the system hotspot identification model, and identify system hotspot problems according to the abnormal pattern obtained after the dynamic adjustment operation.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the intelligent identification method of system hot spot problems described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent identification method of system hot spot problems described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Unit colleague commuting and sharing willingness influence factor evaluation method based on structural equation model
CN115496395A
Intelligent inspection decision-making method and device based on event detection
CN118229272A
Root cause recognition model training method, root cause recognition method, device and equipment
CN118827326A
Airport data service interface fault analysis method and system
CN119248560A
Abnormal data prediction method and system based on dynamic threshold and pattern mining
CN119475148A
Cited By
Abnormal event detection method and system based on multi-source operation and maintenance data fusion
CN120803804A