Fault automatic detection and diagnosis processing method and device based on data platform, and terminal

By realizing automatic fault detection, diagnosis and repair on the data platform, the problem of low efficiency and long response time of fault processing in the existing technology is solved, and the reliability and availability of the data platform are improved.

CN120216243APending Publication Date: 2025-06-27SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510323371.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, data platform fault handling relies on manual monitoring and manual intervention, and there are problems such as low processing efficiency, long response time, and high professional requirements for operation and maintenance personnel.

Method used

It provides a data platform-based fault automatic detection and diagnosis processing method, which can achieve self-healing of faults by timely collecting system resource data, analyzing and judging abnormalities, performing fault diagnosis and automatic repair.

Benefits of technology

It realizes automatic detection, accurate diagnosis and automatic repair of data platform faults, improves fault handling efficiency, shortens response time, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216243A_ABST
    Figure CN120216243A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic fault detection and diagnosis processing method and device based on a data platform and a terminal. The method comprises the following steps: collecting system resource data regularly according to a set time interval; analyzing and judging whether the collected system resource data is abnormal or not; when it is judged that the system resource data is abnormal, fault diagnosis is controlled to be carried out, a mode based on combination of preset rule diagnosis, log analysis diagnosis and / or correlation analysis diagnosis is adopted, a fault generation reason is analyzed, and fault generation reason data is generated; and based on the generated fault generation reason data, searching a preset automatic repair strategy corresponding to the generated fault generation reason data, performing fault repair according to a repair scheme recorded by the found automatic repair strategy, and outputting a fault repair result. The invention provides the method capable of automatically detecting the fault, accurately diagnosing the fault reason and automatically executing the repair measure, so that the stable operation of the data platform can be ensured, and the reliability and availability of the data platform can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data platform operation and maintenance, and particularly relates to a method, device, intelligent terminal and storage medium for automatic fault detection, diagnosis and processing based on a data platform. Background Art

[0002] With the deepening of enterprise digital transformation, data platforms play a crucial role in storing, processing and analyzing massive amounts of data. However, data platforms may face various faults during operation, such as system resource overload, data service anomalies, network faults, etc. If these faults cannot be processed in a timely manner, they will seriously affect the business operations and data analysis work of enterprises.

[0003] In the prior art, most traditional fault handling methods rely on manual monitoring and manual intervention, which have problems such as low processing efficiency, long response time, and high professional requirements for operation and maintenance personnel.

[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, device, intelligent terminal and storage medium for automatic fault detection, diagnosis and processing based on a data platform, aiming at the problems and defects of the above-mentioned prior art, and providing a self-healing method for automatically detecting faults, accurately diagnosing the causes of faults and automatically executing repair measures, which can ensure the stable operation of the data platform and improve the reliability and availability of the data platform.

[0006] The technical solution adopted by the present invention to solve the problem is as follows: A method for automatic fault detection, diagnosis and processing based on a data platform, including: Regularly collecting system resource data at a set time interval; Analyzing the collected system resource data to determine whether an anomaly occurs; When it is determined that the system resource data is abnormal, control is performed for fault diagnosis, and a combination of diagnosis based on preset rules, log analysis diagnosis and / or correlation analysis diagnosis is used to analyze the cause of the fault and generate cause data of the fault; Based on the generated cause data of the fault, a preset automatic repair strategy corresponding to the generated cause data of the fault is searched, and fault repair is performed according to the repair plan recorded in the found automatic repair strategy, and the fault repair result is output.

[0007] In the above-mentioned method for automatic fault detection, diagnosis and processing based on a data platform, the step of analyzing the collected system resource data to determine whether an anomaly occurs includes: Compare the collected system resource data with the threshold of the corresponding item to determine whether an abnormality occurs; When the collected system resource data exceeds the threshold of the corresponding item, it is determined that an abnormality has occurred.

[0008] The automatic fault detection, diagnosis and processing method based on a data platform, wherein the step of analyzing the collected system resource data to determine whether an abnormality occurs includes: Detect the connection status, query performance, master-slave synchronization data of the database service, and detect the startup situation, execution duration, error message data of the data processing task, and detect the message backlog, consumption rate, and connection number operation status of the message queue to determine whether an abnormal situation occurs; When one or more of the connection status, query performance, master-slave synchronization data of the detected database service, and the startup situation, execution duration, error message data of the detected data processing task, and the message backlog, consumption rate, and connection number operation status of the detected message queue do not conform to the corresponding preset rules, it is determined that an abnormal situation occurs currently.

[0009] The automatic fault detection, diagnosis and processing method based on a data platform, wherein the step of, when it is determined that the system resource data is abnormal, controlling to perform fault diagnosis, and analyzing the cause of the fault and generating the cause data of the fault by combining rule-based diagnosis, log analysis diagnosis and / or correlation analysis diagnosis includes: When it is determined that the system resource data is abnormal, control to perform fault diagnosis; When analyzing the cause of the fault by rule-based diagnosis, obtain the historical fault data and the corresponding monitoring index situation, use the machine learning rule learning algorithm to generate initial rules, and continuously update and optimize the rules in the knowledge base. When abnormal indicators of the system resource data are detected, control to match in the knowledge base, analyze the possible cause of the fault according to the matching result, and generate the cause data of the fault; When analyzing the cause of the fault by log analysis diagnosis, use the ELK stack log management platform to collect the log information of each component, and extract error information and exception stack from the collected log information according to the preset log parsing rules; assist in locating the root cause of the fault based on the extracted error information and exception stack, and generate the cause data of the fault; When analyzing the cause of the fault by correlation analysis diagnosis, construct an index correlation model, analyze the mutual influence between different monitoring indicators, combine the abnormal situation of the system resource data with the mutual influence situation between different monitoring indicators, diagnose the fault point and generate the cause data of the fault.

[0010] The described method for automatic fault detection, diagnosis, and processing based on a data platform, wherein the step of finding a preset automatic repair strategy corresponding to the generated fault cause data and performing fault repair according to the repair plan recorded in the found automatic repair strategy includes: When the automatic repair strategy found based on the generated fault cause data and corresponding to the generated fault cause data is for resource adjustment - type problem repair, for the case of running in a virtualized environment, a request is sent to the cluster scheduler to dynamically expand resources according to the resource shortage situation; in a physical server environment, system parameters are adjusted through scripts to optimize resource allocation; When the automatic repair strategy found based on the generated fault cause data and corresponding to the generated fault cause data is for service restart - type repair, control is performed to automatically execute a restart operation according to a pre - set service restart script when a software fault occurs in the service, and the service status is continuously monitored after the restart; When the automatic repair strategy found based on the generated fault cause data and corresponding to the generated fault cause data is for data repair - type repair, according to the pre - backed - up data, when a data consistency problem is found, data is automatically extracted from the pre - backed - up data for recovery, and relevant tasks are reasonably paused and restarted for real - time data processing scenarios; When the automatic repair strategy found based on the generated fault cause data and corresponding to the generated fault cause data is for network configuration repair, control is performed to use a network automation configuration tool to automatically re - configure network parameters and reset network connections according to a preset network recovery template when a network fault occurs; According to the above - mentioned automatic repair strategy, a repair operation is executed through an automation script or by calling the management API of each component, and a detailed repair log is recorded. The described method for automatic fault detection, diagnosis, and processing based on a data platform, wherein the step of finding a preset automatic repair strategy corresponding to the generated fault cause data and performing fault repair according to the repair plan recorded in the found automatic repair strategy further includes: Verify the fault repair result, and continuously detect relevant monitoring indicators and service running status after the fault repair according to the set verification period; If it returns to normal within the period, it is determined that the repair is successful; otherwise, an escalation alarm notification is triggered.

[0011] The described method for automatic fault detection, diagnosis, and processing based on a data platform, wherein after the step of outputting the fault repair result, it further includes: When the fault repair result is not repaired, control is performed to give a warning and notification.

[0012] An automatic fault detection, diagnosis and processing device based on a data platform, wherein the device includes: A fault monitoring module, configured to regularly collect system resource data at a set time interval; and analyze the collected system resource data to determine whether an abnormality occurs; A fault diagnosis module, configured to, when it is determined that the system resource data is abnormal, control the execution of fault diagnosis, and analyze the cause of the fault by combining preset rule-based diagnosis, log analysis diagnosis, and / or correlation analysis diagnosis, and generate fault cause data; A fault repair module, configured to, based on the generated fault cause data, search for a preset automatic repair strategy corresponding to the generated fault cause data, perform fault repair according to the repair plan recorded in the found automatic repair strategy, and output a fault repair result; An alarm and notification module, configured to verify the fault repair result, continuously detect relevant monitoring indicators and service running status after the fault is repaired according to a set verification period; when the fault repair result is not repaired, control the execution of warnings and notifications.

[0013] An intelligent terminal, which includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include instructions for executing the method according to any one of the above.

[0014] A computer-readable storage medium, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the method according to any one of the above.

[0015] Advantages of the present invention: The present invention provides an automatic fault detection, diagnosis and processing method, device, intelligent terminal and storage medium based on a data platform; the present invention provides a self-healing fault method that can automatically detect faults, accurately diagnose the cause of faults and automatically execute repair measures, which can ensure the stable operation of the data platform, so as to improve the reliability and availability of the data platform. The present invention basically does not rely on manual monitoring and manual intervention, improves the fault handling efficiency, shortens the response time, and reduces the operation and maintenance costs. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of the method for automatically detecting, diagnosing, and processing faults based on a data platform provided in Embodiment 1 of the present invention.

[0018] Figure 2 It is a schematic flowchart of the method for automatically detecting, diagnosing, and processing faults based on a data platform provided in Embodiment 2 of the present invention.

[0019] Figure 3 A principle block diagram of an embodiment of the apparatus for automatically detecting, diagnosing, and processing faults based on a data platform provided by the present invention.

[0020] Figure 4 It is a principle block diagram of the internal structure of an intelligent terminal provided in an embodiment of the present invention. Detailed implementation manners

[0021] To make the objectives, technical solutions, and advantages of the present invention clearer and more explicit, the following further elaborates on the present invention by way of examples with reference to the accompanying drawings. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0022] With the in-depth digital transformation of enterprises, data platforms play a crucial role in storing, processing, and analyzing massive amounts of data. However, data platforms may face various faults during operation, such as system resource overload, data service anomalies, network failures, etc. If these faults cannot be processed in a timely manner, they will seriously affect the business operations and data analysis work of enterprises.

[0023] In view of the problems that traditional fault handling methods in the prior art mostly rely on manual monitoring and manual intervention, such as low processing efficiency, long response time, and high professional requirements for operation and maintenance personnel. The present invention provides a method for self-healing data platform faults, aiming to achieve automatic detection, diagnosis, and repair of data platform faults and ensure the stable operation of the data platform.

[0024] As Figure 1 shown, a method for automatically detecting, diagnosing, and processing faults based on a data platform in Embodiment 1 of the present invention includes the following steps: Step S100: Regularly collect system resource data at a set time interval; In an embodiment of the present invention, the system resource data of the data platform is regularly collected at a set time interval. For example, the CPU usage rate is collected every 5 minutes; in an embodiment of the present invention, the collection and detection of system resource data include data operation status data and data service status data.

[0025] Among them, the running state data includes metric data such as CPU usage rate, memory usage, disk I / O, network bandwidth, etc. The running state data can be detected by using operating system tools (such as top, sar, etc. under Linux) and professional monitoring software (such as Zabbix, Prometheus, etc.).

[0026] For the detection of data service state data, the running states of database services (connection status, query performance, master-slave synchronization, etc.), data processing tasks (start-up situation, execution duration, error messages, etc.), and message queues (message backlog, consumption rate, number of connections, etc.) can be monitored respectively through the built-in monitoring function of the database, task scheduling system monitoring, message queue management tools, etc.

[0027] Specifically, in this step of the embodiment, regarding the set time interval, that is, the frequency of data collection needs to be determined. For example, data can be collected once per minute, per hour, or per day. The selection of this time interval usually depends on the system requirements and the purpose of monitoring.

[0028] For timed collection, the present invention can use timed tasks (such as cron jobs, Windows Task Scheduler, etc.) to automatically perform data collection; relevant scripts or programs can be automatically run at the set time points without manual intervention.

[0029] Regarding the resource data, that is, the collected data may include CPU usage rate, memory occupancy, disk I / O, network traffic, etc., and may also include the running state data of database services (connection status, query performance, master-slave synchronization, etc.), data processing tasks (start-up situation, execution duration, error messages, etc.), and message queues (message backlog, consumption rate, number of connections, etc.). These data can help understand the running state and performance bottlenecks of the system.

[0030] During specific implementation, the collected data can be stored in a database or log file for subsequent analysis and reporting. By analyzing these data, potential problems can be discovered, resource allocation can be optimized, or capacity planning can be carried out.

[0031] Step S200: Analyze the collected system resource data to determine whether an anomaly occurs; In the embodiment of the present invention, the collected system resource data will be analyzed to determine whether an anomaly occurs; for example, the collected system resource data, including metrics such as CPU usage rate, memory usage, disk I / O, network bandwidth, etc., is compared with the corresponding thresholds to determine whether an anomaly occurs.

[0032] And monitor the running status of the database service (connection status, query performance, master-slave synchronization, etc.), data processing tasks (start-up situation, execution duration, error messages, etc.), and message queues (message backlog, consumption rate, number of connections, etc.) respectively through the built-in monitoring function of the database, task scheduling system monitoring, message queue management tool, etc. Once an abnormal situation occurs, trigger the subsequent diagnosis process.

[0033] In specific implementation, the present invention can also organize the collected system resource data first, for example, organize the collected system resource data into a format easy to analyze, such as using a data framework (DataFrame) or a database table. Ensure the integrity and accuracy of the data for subsequent analysis.

[0034] And set abnormal determination rules: Before analysis, it is necessary to set abnormal determination rules. This can be determined through historical data, such as the average CPU usage rate, memory usage amount, etc. of the system in the normal working state. These abnormal determination rules will be used to compare the current data.

[0035] The method for judging whether an abnormality occurs can adopt the following methods: 1) Statistical analysis: Use statistical methods, such as mean, standard deviation, etc., to automatically judge whether the data exceeds the normal range. If the usage amount of a certain resource is much higher than the reference value, it may indicate an abnormality.

[0036] 2) Visualization tools: Use charts (such as line charts, bar charts, etc.) to visualize the data to help identify trends and abnormal fluctuations.

[0037] 3) Machine learning: In some complex scenarios, machine learning algorithms can be used to detect abnormalities in the data, and an unusual pattern can be automatically identified by training a model.

[0038] Judge the abnormality: According to the analysis result, automatically judge whether an abnormal situation occurs. The abnormality can be continuous high CPU utilization rate, memory leakage, abnormal network traffic, etc. When it is judged that the system resource data is abnormal, enter step S300.

[0039] Step S300: When it is judged that the system resource data is abnormal, control the fault diagnosis to be carried out in a combined manner of diagnosis based on preset rules, log analysis diagnosis, and / or correlation analysis diagnosis, analyze the cause of the fault, and generate the data on the cause of the fault; In the previous steps of the present invention, the abnormality judgment of the system resource data is carried out by real-time monitoring of the resource data (such as CPU usage rate, memory usage situation, network traffic, etc.), and whether there is an abnormality is judged through a set threshold or model. This process can adopt machine learning algorithms to learn the characteristics of normal and abnormal states through historical data.

[0040] In the embodiments of the present invention, when it is determined that the system resource data is abnormal, fault diagnosis is controlled; the system will automatically start the fault diagnosis program. This control process can ensure the system's rapid response and reduce the impact of faults on services.

[0041] Regarding the combination of specific fault diagnosis methods: Diagnosis based on preset rules: Specifically, by sorting out historical fault data and corresponding monitoring index situations, using machine learning rule learning algorithms to generate initial rules, and continuously updating and optimizing the rules in the knowledge base. When abnormal indicators are detected, they are matched in the knowledge base, and possible fault causes are analyzed according to the matching results; preliminary judgments are made using the updated rules (such as thresholds under specific conditions); this method is fast and easy to implement.

[0042] Diagnosis based on log analysis: The present invention will analyze the generated log data and identify faults by searching for error messages, warnings, abnormal events, etc. Log analysis can provide more in-depth context information to help locate problems.

[0043] Correlation analysis diagnosis: By analyzing the relationships between different system components, potential fault sources are identified. For example, an abnormality in one service may be related to a performance degradation in another service. Correlation analysis usually requires more complex data processing and models. An index correlation model can also be constructed (for example, using a directed graph structure, where nodes represent indicators and edges represent correlation relationships and weights), analyze the mutual influences between different monitoring indicators, and comprehensively consider the situations of multiple correlated indicators when a fault occurs to more accurately diagnose the fault point.

[0044] Regarding generating data on the causes of faults, finally, the embodiments of the present invention will integrate the above analysis results, generate detailed data on the causes of faults, and locate specific fault points. In further embodiments of the present invention, intelligent fault prediction can be adopted. Specifically, by introducing a machine learning model, potential fault risks can be predicted through historical data and real-time monitoring data. This can take preventive measures in advance to prevent faults from occurring.

[0045] In further embodiments of the present invention, an adaptive rule engine can be adopted. Specifically, an adaptive rule engine can be developed to dynamically adjust diagnostic rules according to the system operation status and historical fault cases to improve the accuracy and efficiency of fault detection.

[0046] In further embodiments of the present invention, a visual fault analysis tool can be set up. Design a graphical interface to help operation and maintenance personnel more intuitively view the system status, fault logs, and correlation analysis results, and improve the efficiency of fault diagnosis.

[0047] In a further embodiment of the present invention, a multi-level collaborative diagnosis can be set up. By combining artificial intelligence with the expert knowledge base of human experts, a multi-level collaborative diagnosis system is established. In the case of complex faults, the system can introduce the expert knowledge base for in-depth analysis, thereby increasing the success rate of fault resolution.

[0048] In a further embodiment of the present invention, a real-time feedback mechanism can be set up. Specifically, a real-time feedback mechanism can be established to feed back the results of the fault cause analysis to the system learning module, so as to continuously optimize the algorithms for fault detection and diagnosis and improve the intelligence level of the system.

[0049] Through the above steps, the present invention can further improve the stability of the system and the fault handling ability, reduce the operation and maintenance costs, and enhance the user experience.

[0050] Step S400: Based on the generated fault cause data, search for the preset automatic repair strategy corresponding to the generated fault cause data, perform fault repair according to the repair plan recorded in the found automatic repair strategy, and output the fault repair result.

[0051] In the embodiment of the present invention, regarding the fault cause data, that is, after the fault is diagnosed, a set of data about the fault cause will be generated first. These data may include information such as the fault type, affected system components, timeline, abnormal patterns, etc.

[0052] Then, based on the fault cause data, search for the automatic repair strategy. Specifically, the preset knowledge base or strategy library can be used to search for the automatic repair strategy corresponding to the generated fault cause data. These strategies can be formulated based on past fault handling experiences and include specific repair steps, commands, scripts, etc. This process can use pattern matching or machine learning techniques to automatically classify and match the fault causes and corresponding repair strategies, thereby improving the accuracy and efficiency of the search.

[0053] Then perform the fault repair. That is, in the embodiment of the present invention, once the appropriate automatic repair strategy is found, the system will automatically execute the fault repair operation according to these repair plans. This may include restarting the service, adjusting the configuration, replacing the faulty component, cleaning the log, etc. In the embodiment of the present invention, the process of fault repair is as seamless as possible to avoid causing a greater impact on users or other system components.

[0054] After the fault is repaired, output the fault repair result. That is, in the embodiment of the present invention, after the repair plan is executed, a report of the repair result will be generated, which includes whether the fault is successfully repaired, what operations are performed, and whether there are subsequent situations that need to be monitored, etc. These results can help the operation and maintenance personnel quickly understand the problem-solving situation and provide feedback for future fault handling.

[0055] In a further embodiment of the present invention, a self - learning repair function can be set. Specifically, by introducing a self - learning mechanism, after repairing a fault, the repair effect can be analyzed to update and optimize the automatic repair strategy for subsequent similar faults. This enables the system of the method of the present invention to continuously evolve and more effectively cope with unpredictable fault situations.

[0056] In a further embodiment of the present invention, a repair strategy recommendation based on AI can be provided. Specifically, through an analysis tool based on artificial intelligence, a dynamic repair strategy can be generated according to real - time monitoring data and historical events. Machine learning algorithms can be used to identify uncommon fault patterns and corresponding solutions.

[0057] In a further embodiment of the present invention, a multi - level automatic repair function can be set. For example, a multi - level automatic repair system can be designed to select a repair strategy according to the severity and scope of influence of the fault. For example, for minor faults, low - interference automatic repair measures can be used; for severe faults, more thorough means can be taken, but moderate manual intervention may be required.

[0058] In a further embodiment of the present invention, a real - time monitoring and feedback mechanism can be set. By establishing a real - time monitoring system, the performance of the system after repair can be continuously tracked to confirm the effectiveness of the repair. At the same time, by sending feedback to the operation and maintenance team, their understanding of the system health status can be improved.

[0059] In a further embodiment of the present invention, a user - defined repair strategy can be set. Specifically, users can be allowed to define and store their own repair strategies in specific situations and integrate these strategies into the automatic repair process. This can quickly respond to and handle difficult problems according to the needs of specific industries or users.

[0060] Through the above - mentioned further embodiments, the degree of automation and the level of intelligence of fault repair can be greatly improved, thereby enhancing the stability and reliability of the system, and reducing manual intervention and operation and maintenance costs.

[0061] In another embodiment of the present invention, for the method of automatically detecting, diagnosing and processing faults based on a data platform, the step of analyzing the collected system resource data to determine whether an abnormality occurs may further include: S201. Compare the collected system resource data with the threshold of the corresponding item to determine whether an abnormality occurs; S202. When the collected system resource data exceeds the threshold of the corresponding item, it is determined that an abnormality occurs.

[0062] For example, operating system tools (such as top, sar, etc. under Linux) and professional monitoring software (such as Zabbix, Prometheus, etc.) can be used to collect system resource data at set time intervals (CPU usage is collected every 5 minutes), including metrics such as CPU usage, memory usage, disk I / O, network bandwidth, etc., and compare them with corresponding thresholds to determine whether an anomaly occurs.

[0063] That is, in the embodiments of the present invention, the system will monitor and collect various resource data in real time, such as CPU usage, memory occupancy, disk read / write speed, network traffic, etc. These data are important bases for evaluating the system operation status.

[0064] Then, for each resource item to be monitored, the present invention will preset one or more thresholds (such as maximum limit or safe range). Each threshold represents the standard for normal resource usage. For example, the threshold for CPU usage may be set to 85%, and exceeding this value may be considered an anomaly.

[0065] Then, anomaly judgment is performed. That is, once the collected data exceeds the preset corresponding threshold, the present invention will immediately perform anomaly judgment. This usually involves conditional logic judgment and may combine different monitoring data for comprehensive evaluation. Of course, this anomaly detection does not rely solely on a single threshold and can also combine methods such as trend analysis and historical data comparison to more comprehensively determine whether the current state is abnormal.

[0066] For example, if the CPU usage of a web server exceeds 90% during collection, the system will determine that the current state is abnormal. This may mean that the server is under excessive requests or there is a potential resource leak. Accordingly, the system will start the monitoring process, analyze the cause of the anomaly and solve it.

[0067] In the embodiments of this step, through real-time monitoring and anomaly judgment, measures can be taken before the fault problem becomes serious, thereby reducing the risk of system downtime and performance degradation. Through such a monitoring mechanism, a higher level of system stability and availability can be achieved.

[0068] In other embodiments of the present invention, in the method for automatically detecting, diagnosing and processing faults based on a data platform, the step of analyzing the collected system resource data to determine whether an anomaly occurs further includes: S201. Detect the connection status, query performance, master-slave synchronization data of the database service, and detect the startup situation, execution duration, error message data of the data processing task, and detect the message backlog, consumption rate, and connection number operation status of the message queue to determine whether an abnormal situation occurs; S202. When one or more items among the running status of detecting the connection status, query performance, master-slave synchronization data of the database service, the startup situation, execution duration, error message data of the data processing task, and the message backlog, consumption rate, and connection count of the message queue do not conform to the corresponding preset rules, it is determined that an abnormal situation occurs currently.

[0069] That is, in the embodiments of the present invention, the running status of the database service (connection status, query performance, master-slave synchronization, etc.), data processing tasks (startup situation, execution duration, error message, etc.), and message queue (message backlog, consumption rate, connection count, etc.) can be monitored respectively through the built-in monitoring function of the database, task scheduling system monitoring, message queue management tool, etc. Once an abnormal situation occurs, the subsequent diagnosis process is triggered.

[0070] Regarding the built-in monitoring function of the database, in the database system (such as MySQL, PostgreSQL, etc.), there is a built-in monitoring function that can monitor key indicators such as the connection status, query performance, transaction processing, and master-slave synchronization of the database. These monitoring functions can help detect performance bottlenecks or connection problems in a timely manner and ensure data consistency and service reliability.

[0071] Regarding the task scheduling system monitoring, specifically, the task scheduling system (such as Cron, Apache Airflow, etc.) is responsible for managing regularly executed tasks. These systems can monitor the startup situation, execution duration, and whether there are error messages of the data processing task. Through these monitors, it can be ensured that the data processing task is completed on time and can quickly respond to the situation of processing failure.

[0072] Regarding the message queue management tool monitoring, specifically, the message queue system (such as RabbitMQ, Kafka, etc.) is used for processing asynchronous message passing. Monitoring the status of the message queue can include the message backlog situation, consumption rate, and connection count, etc. By monitoring these metrics, problems such as message delay, loss, or system overload can be effectively prevented.

[0073] In the embodiments of the present invention, once any one of these monitoring tools detects an abnormal state (such as database connection failure, task execution timeout, or slow message consumption speed, etc.), the subsequent diagnosis process will be immediately triggered. These processes may include automatically recording abnormal information, notifying the operation and maintenance personnel, scheduling a fault handling program, etc.

[0074] Such a monitoring mechanism can improve the robustness of the system, ensure a rapid response when problems occur, reduce downtime, and maintain the availability of the service and the validity of the data.

[0075] By integrating multiple monitoring tools, the present invention can provide the system with comprehensive status monitoring and anomaly response capabilities. Effective monitoring strategies not only facilitate the smooth progress of daily operation and maintenance.

[0076] In another embodiment of the present invention, in the automatic fault detection, diagnosis and processing method based on a data platform, when it is determined that the system resource data is abnormal, then control is performed to conduct a fault diagnosis. The step of analyzing the cause of the fault and generating cause - of - fault data by combining rule - based diagnosis, log analysis diagnosis and / or correlation analysis diagnosis includes: S301. When it is determined that the system resource data is abnormal, then control is performed to conduct a fault diagnosis; S302. When analyzing the cause of the fault by rule - based diagnosis, obtain historical fault data and the corresponding monitoring index conditions, use the machine - learning rule - learning algorithm to generate initial rules, and continuously update and optimize the rules in the knowledge base. When abnormal indicators of system resource data are detected, control is performed to match in the knowledge base, and analyze the possible cause of the fault according to the matching result, and generate cause - of - fault data; In this step of the embodiment, the historical fault data includes past fault events, fault occurrence times, types, handling measures and their results, etc. By analyzing this data, the system can identify the patterns and trends of faults. The monitoring index conditions refer to the system resource usage conditions (such as CPU, memory, disk I / O, network traffic data, etc.) recorded when the fault occurs, which are incorporated as monitoring indexes. These indexes can help analyze under what circumstances the fault appears, so as to formulate diagnostic criteria.

[0077] Regarding generating the initial rules, using the machine - learning rule - learning algorithm, the present invention can analyze historical fault data, extract effective patterns and correlations from it. Based on this information, preliminary fault diagnosis rules are generated. These rules may include thresholds under specific conditions, identification logics of fault types, etc. In the embodiment of the present invention, the generation of the initial rules is automated and can continuously learn and improve according to historical data, further enhancing the accuracy of diagnosis.

[0078] Regarding the update and optimization of the rules in the knowledge base, in the embodiment of the present invention, by continuously obtaining new monitoring data and fault cases, dynamic update of the rules in the knowledge base can be realized. This process ensures that the rules always reflect the current operating state of the system and specific contexts. The optimization process can adopt methods such as online learning or incremental learning to enable the system to improve its diagnostic ability during actual operation.

[0079] Regarding the monitoring of abnormal indicators and matching, specifically: In the embodiments of the present invention, once an abnormal indicator of system resource data is detected, the existing rules in the knowledge base will be automatically consulted. At this time, based on the currently monitored abnormal situation, a match will be made with the rules in the knowledge base. The matching result will reveal the possible causes of failure corresponding to this state, the failure situation can be inferred from the rules, and which failures are most likely to occur can be identified.

[0080] Then the present invention will generate data on the cause of the failure, that is, in the embodiments of the present invention, finally, based on the matching result, data on the cause of the failure will be generated. This may include information such as the failure category, the scope of influence, and the possible root cause, providing a basis for subsequent failure handling and repair.

[0081] This step S302 effectively improves the intelligent level of fault diagnosis by combining historical data, machine learning, dynamic rule update, and monitoring metrics. It can transform historical experience into dynamic rules, enabling the system to autonomously adapt to the new operating environment and improving the fault detection and handling capabilities. The core of this method lies in promoting intelligent operation and maintenance and enhancing the robustness of the system through a data-driven approach.

[0082] S303. When diagnosing the cause of the failure using log analysis, use a log management platform such as the ELK stack (Elasticsearch, Logstash, Kibana) to collect and store the log information of each component, and extract error information and exception stacks from the collected log information according to preset log parsing rules; assist in locating the root cause of the failure based on the extracted error information and exception stacks, and generate data on the cause of the failure; In the embodiments of the present invention, regarding log collection and storage, the ELK stack (Elasticsearch, Logstash, Kibana) is used. Among them, Elasticsearch: A distributed search and analysis engine for storing and quickly retrieving log data. Logstash: A tool for data processing that can collect log information from different sources, filter and parse it, and finally send the data to Elasticsearch. Kibana: A tool for data visualization and analysis, through which users can display and query the log data stored in Elasticsearch.

[0083] The present invention collects log information from each component (such as application servers, databases, network devices, etc.) through Logstash and stores it in Elasticsearch. These log information may include error logs, access logs, operation logs, etc.

[0084] Then, log parsing and extraction are performed. Specifically, during implementation, according to the preset log parsing rules, during the log collection process, the present invention will parse the log according to the preset parsing rules (such as regular expressions, field delimiters, etc.). These rules help the system identify the key information in the log.

[0085] Then, error information and exception stack are extracted. That is, from the parsed log, the present invention can extract error information (such as error codes, error descriptions) and exception stack (stack trace). The exception stack usually contains the call path when the program goes wrong and can provide the context information of the fault occurrence.

[0086] Then, it helps to locate the root cause of the fault. Specifically, the extracted information can be analyzed. For example, according to the extracted error information and exception stack, the present invention can assist in locating the root cause of the fault. This process includes: finding the meaning of specific error codes; analyzing the call chain in the exception stack to determine the specific location where the fault occurs. Comprehensive analysis is carried out in combination with other monitoring data (such as performance metrics, resource usage, etc.).

[0087] Then, the embodiment of the present invention will generate data on the cause of the fault: Specifically, the present invention will generate data on the cause of the fault according to the analysis results, including: fault type (such as application exception, network fault, database error, etc.); affected system components or services; possible root causes and recommended repair measures.

[0088] In this way, through the log analysis of the ELK stack, the embodiment of the present invention can efficiently collect, store and parse log information, and extract valuable error data and exception stack from it. This method not only improves the accuracy of fault diagnosis, but also speeds up the fault location, helping the operation and maintenance team to respond quickly and solve problems. Log analysis, as an important means of fault diagnosis, can provide strong support for the stable operation of the system.

[0089] S304. When diagnosing the cause of the fault by using correlation analysis, construct an index correlation model (for example, adopt a directed graph structure, where nodes represent indexes and edges represent correlation relationships and weights), analyze the mutual influence between different monitoring indexes, and combine the abnormal situation of system resource data with the mutual influence situation between different monitoring indexes to diagnose the fault point and generate data on the cause of the fault.

[0090] In this embodiment, when analyzing the causes of faults based on association analysis, an index association model is constructed. Regarding the directed graph structure, it is used to establish the relationships between monitoring indexes. Nodes represent different monitoring indexes (such as CPU usage rate, memory usage rate, disk I / O, etc.), and edges represent the association relationships and weights between these indexes. The weight of each edge can reflect the degree of mutual influence between two indexes, and the higher the weight, the stronger the correlation between the two indexes.

[0091] The present invention analyzes the mutual influence between monitoring indexes. Specifically, by analyzing the edges in the directed graph, the present invention can identify which monitoring indexes influence each other when a fault occurs. For example, when the CPU usage rate rises sharply, it may lead to an increase in memory usage rate or an extension of the database query response time. This kind of analysis helps the system quickly identify potential fault points. Especially in complex systems, an abnormality in one index may affect multiple other indexes.

[0092] Then, anomaly monitoring and fault diagnosis are carried out. Regarding anomaly monitoring, in the present invention, when an abnormal index appears in the system resource data (such as a certain monitoring index exceeding the threshold), combined with the index association model, the system will analyze the mutual influence between this abnormal index and other indexes.

[0093] Regarding fault diagnosis, through analysis, the present invention can identify possible fault sources. For example, if the memory usage rate is extremely high and the CPU usage rate is also rising, the system may judge that it is a performance problem caused by memory leakage.

[0094] Finally, the fault cause data generated by the present invention can include the type of fault, the affected components, and the possible root causes. Among them: the types of faults such as performance degradation, service interruption, etc.; the affected system components (such as a certain database or service); the specific fault root cause and recommended solutions (for example, restarting the service, increasing resource allocation, etc.); When the present invention is specifically implemented, it can also adopt machine learning-based dynamic adjustment of the weights between indexes. For example, introducing machine learning algorithms to automatically adjust the weights between indexes according to historical data. This dynamic update ability can make the model adapt to the system usage pattern over time and improve the accuracy.

[0095] Then, time series prediction technology can be used in combination with association analysis to identify potential fault risks in advance. For example, by observing the trends in historical data, the present invention can predict possible future anomalies and take measures in advance to prevent faults from occurring.

[0096] In another embodiment, the present invention relates to cross - layer linkage analysis. An extended model can be adopted to make it not only limited to monitoring metrics, but also capable of performing correlation analysis in combination with business aspects (such as user behavior, transaction volume, etc.), providing a more comprehensive perspective for fault diagnosis. For example, analyze whether the sharp drop in sales data on the same day is related to the increased load of the backend service. And establish a fault repair feedback mechanism so that after each fault is processed, the processing result can be updated back to the model to further optimize the relationship between metrics and improve the intelligence of fault diagnosis.

[0097] Through the above steps, the present invention can significantly improve the accuracy and response speed of the fault diagnosis process based on correlation analysis, thereby enhancing the stability and reliability of the entire system.

[0098] In another embodiment of the present invention, for the automatic fault detection, diagnosis and processing method based on a data platform, the step of finding a preset automatic repair strategy corresponding to the generated fault cause data and performing fault repair according to the repair plan recorded in the found automatic repair strategy specifically includes: S401. When the preset automatic repair strategy corresponding to the generated fault cause data found based on the generated fault cause data is a problem repair of resource adjustment, for the case of running in a virtualized environment, send a request to the cluster scheduler to dynamically expand resources according to the resource shortage situation; in a physical server environment, optimize resource allocation by adjusting system parameters through a script; That is, in the embodiment of the present invention, for resource adjustment in a virtualized environment, dynamic resource expansion is adopted. For example, when the system detects a shortage of certain resources (such as CPU, memory, etc.), the present invention can send a request to the cluster scheduler to dynamically expand resources (such as increasing the container CPU limit, memory quota) according to the generated fault cause data. In this way, the cluster scheduler can automatically allocate more virtual machine instances or increase the resource configuration of existing virtual machines according to the load situation.

[0099] For example, in a cloud environment running an online shopping platform, if it is monitored that the backup processing task causes a sharp increase in CPU usage, thereby affecting the response time of other services, the present invention can send a request to the cluster scheduler to add more virtual machines to process the backup task, thus releasing the resource pressure on the main service.

[0100] Regarding resource optimization in a physical server environment, the present invention uses a script to adjust system parameters. Specifically, in a physical server environment, the present invention may directly adjust relevant system parameters (such as memory limit, CPU scheduling policy, etc.) to optimize resource allocation. This is usually achieved through pre - configured scripts to ensure fast response and reduce manual intervention.

[0101] For example, when a database server detects high memory usage during operation, resulting in slow query response. At this time, the present invention can execute a preset script to adjust the memory allocation parameters of the operating system, such as increasing the swap space or adjusting the cache size, to optimize the database performance.

[0102] It can be seen that through this automatic repair strategy based on the data of the cause of the fault, the present invention can quickly respond to the problem of insufficient resources and ensure the stability and high availability of the service. This method not only improves the operation and maintenance efficiency but also reduces the potential need for manual intervention, thereby accelerating fault handling and service recovery.

[0103] S402. When the automatically repaired policy corresponding to the generated cause - of - fault data found preset is service restart - type repair, control to automatically execute a restart operation according to a preset service restart script when a software fault occurs in the service. For example, first stop the dependent services in the pre - configured order, restart the target service, and then start the dependent services, and control to continuously monitor the service status after the restart; In the embodiment of this step, when the automatic repair policy is service restart - type repair, control according to a preset service restart script. Specifically, when the system detects a software fault (such as service crash, abnormal response, etc.), it will automatically search for a preset automatic repair policy corresponding to the cause - of - fault data. This policy usually includes the restart process for specific services. The present invention automatically executes the restart operation according to the pre - configured service restart script. For example, in the pre - configured order (first stop the dependent services, restart the target service, and then start the dependent services).

[0104] In the embodiment of the present invention, the order control for restarting the service: first stop the dependent services - restart the target service - then start the dependent services; 1) Stop the dependent services: That is, before restarting the target service, the present invention will first stop all other services that depend on this service in the set order. This can avoid errors or service - to - service conflicts during the restart process.

[0105] 2) Restart the target service: After stopping the dependent services, the present invention will control to restart the target service. This process may include clearing the cache of the service, resetting the status, etc., to ensure that the service restarts in a clean state.

[0106] 3) Start the dependent services: Finally, the present invention will control to restart the previously stopped dependent services to restore the normal workflow of the service.

[0107] The present invention continuously monitors the service status, that is, it monitors after restart. Specifically, after restart, the system of the present invention will continuously monitor the status of the target service to ensure the successful operation and normal recovery of the service. If it is found that the service still has abnormalities, the present invention will trigger further processing or alarms.

[0108] For example, take the case where a software failure occurs in the payment service of an e-commerce platform, resulting in the inability to process orders normally. When the present invention detects this failure, the following steps will be executed: Find the script: The control system of the present invention automatically finds the restart script related to the payment service.

[0109] Stop dependent services: The payment service may depend on the user authentication service and the order processing service. The control system will first stop these two services.

[0110] Restart the payment service: Execute the restart operation to restart the payment service.

[0111] Restart dependent services: After the payment service is restarted, the control system will then enable the user authentication service and the order processing service in sequence.

[0112] Monitor the status: Finally, the control system continuously monitors the payment service to ensure its normal operation. If there are abnormalities, alarms will be issued or other automatic repair measures will be triggered.

[0113] It can be seen that through this automatic service restart mechanism based on the data of the cause of the failure, the present invention can quickly respond to software failures, reduce the need for manual intervention, improve the availability and reliability of the service, and ensure that the user experience is not affected.

[0114] S403. When, based on the generated data of the cause of the failure, the preset automatic repair strategy corresponding to the generated data of the cause of the failure is data repair type repair, according to the pre-backup data, when data consistency problems are found, automatically extract data from the pre-backup data for recovery, and reasonably pause and restart relevant tasks for real-time data processing scenarios; That is, in the embodiments of the present invention, when, based on the generated data of the cause of the failure, the preset automatic repair strategy corresponding to the generated data of the cause of the failure is data repair type repair, it is possible to, according to the data backup strategy (full backup combined with incremental backup), automatically extract data from the backup storage for recovery when data consistency problems are found, and at the same time reasonably pause and restart relevant tasks for real-time data processing scenarios to ensure the integrity and accuracy of the data.

[0115] S404. When the automatically repaired policy corresponding to the generated fault cause data found from the preset ones is network configuration repair, control is performed to use the network automation configuration tool to automatically reconfigure network parameters and reset network connections according to the preset network recovery template during a network fault. In the embodiment of this step, when the system identifies the cause of a network fault and generates fault cause data, it will automatically search for the preset automatic repair policies related to the fault. These policies usually include the specific steps for using the network automation configuration tool to restore network functions.

[0116] The present invention can use pre-configured network recovery templates. These templates contain the specific network parameter configurations and steps required when a specific network fault occurs, such as IP addresses, router settings, switch port configurations, etc. The network automation configuration tool can read these templates and automatically perform network configuration according to the settings in the templates.

[0117] Then when a network fault occurs, the embodiment of the present invention can automatically adjust network parameters (such as reconfiguring IP addresses, subnet masks, DNS settings, etc.) to ensure that network devices can reconnect to the network.

[0118] Of course, in specific implementation, in addition to reconfiguring settings, the network devices can also be restarted to clear the state before the fault and ensure that the devices can be normally connected to the network.

[0119] For example, take the case where a certain router in the company's internal network cannot connect to the Internet due to a configuration error. When the system detects this network fault, the possible automatic repair process is as follows: 1) Identify the fault: When it is identified that the cause of the fault is an incorrect configuration parameter of the router (such as an incorrect IP address configuration).

[0120] 2) Search for the template: The control system automatically searches for the recovery template for the router configuration, and the template contains the standard configuration parameters for this type of device.

[0121] 3) Apply the template: Using the network automation configuration tool, the control system reconfigures the IP address, subnet mask, and other parameters of the router according to the settings in the template.

[0122] 4) Reset the device: The control system resets the router to apply the new configuration and ensure that all changes take effect.

[0123] 5) Restore the connection: After the configuration is completed, the connection between the router and the external network is restored to ensure that all internal users can normally access the Internet.

[0124] It can be seen that through the network configuration automatic repair mechanism based on fault cause data, the present invention can quickly respond to network faults and automatically repair network problems. This method reduces the need for manual intervention, improves the reliability and availability of the network, and ensures business continuity.

[0125] S405. According to the above automatic repair strategy, execute the repair operation through an automated script or by calling the management APIs of each component, and record detailed repair logs. In the embodiment of this step, an automated script is adopted, that is, a program script is written to automatically execute repair tasks, such as restarting services, clearing caches, etc. Or call the API interfaces of the system or application to perform repair operations. For example, stop a service through the API and then restart it.

[0126] Moreover, during the process of the present invention executing the repair operation, the execution situation of each step will be recorded, including success or failure, timestamp, error information, etc., for subsequent analysis and tracking.

[0127] For example, when it is detected that there is a connection problem with the database of a website, the automatic repair strategy may include the following steps: write a script to detect the database connection status. If it is found that the connection fails, automatically restart the database service. The API provided by the database management system can be used in the script. For example, call the POST / restart interface to restart the database. And add a logging function to the script to record the time, result (success or failure), any error information, etc. of the restart operation, and generate a log file. The advantage of this is that it can quickly respond to problems, reduce manual intervention, and at the same time help subsequent troubleshooting and optimization through log records.

[0128] In another embodiment of the present invention, in the automatic fault detection, diagnosis and processing method based on a data platform, the method further includes the steps of: S501. Verify the fault repair result, and continuously detect the relevant monitoring metrics and service running status after the fault is repaired according to the set verification period; S502. If it returns to normal within the period, it is determined that the repair is successful; otherwise, trigger an escalation alarm notification.

[0129] That is, in the embodiment of the present invention, after the repair, a verification period can be set (for example, a specified time every day or every week), and the relevant monitoring metrics and service running status are continuously observed. If it returns to normal within the period, it is determined that the repair is successful; otherwise, trigger an escalation alarm notification for manual intervention.

[0130] S503. When the fault repair result is not repaired, control to give a warning and notification.

[0131] For example, alarms can be classified into different levels such as urgent, important, and general, and classified according to the severity and urgency of the impact of the fault on the data platform. For example, the unavailability of the core database service is classified as an urgent alarm, and the slightly overdue execution of the ETL task is classified as a general alarm.

[0132] The notification method can be integrated with common enterprise communication tools (such as text messages, phone calls, instant messaging tools such as WeCom and DingTalk, and emails), and different notification frequencies and notification objects can be configured for different levels of alarms to achieve accurate and timely push of alarm messages.

[0133] The following further describes the present invention in detail through another specific application embodiment, as Figure 2 shown, a method for automatically detecting, diagnosing and processing faults based on a data platform in this specific application embodiment includes the following steps: S10. Set up a fault monitoring module and enter step S11; S11. Collect system resource data using operating system tools and professional monitoring software; and enter step S12; In this step, use operating system tools (such as top, sar, etc. under Linux) and professional monitoring software (such as Zabbix, Prometheus, etc.) to collect system resource data at a set time interval (CPU usage is collected every 5 minutes); S12. Compare the collected system resource data with the corresponding thresholds to determine whether an abnormality occurs; and enter step S13; S13. Monitor the database service respectively through the built-in monitoring function of the database, task scheduling system monitoring, message queue management tool, etc., to determine whether an abnormality occurs; when an abnormality occurs, enter step S14; S14. Trigger the fault diagnosis process and enter step S15; S15. Set up a fault diagnosis module for fault diagnosis and analysis; and enter steps S16, S17, and S18 respectively; S16. Rule-based diagnosis, using machine learning to generate initial rules and continuously update and optimize them; Specifically, historical fault data and corresponding monitoring index situations can be sorted out, initial rules can be generated using machine learning rule learning algorithms, and the rules in the knowledge base can be continuously updated and optimized. When abnormal indicators are detected, they are matched in the knowledge base, and step S19 is entered to analyze possible fault causes according to the matching results; S17, perform log analysis and diagnosis, collect and centrally store log information of each component with the help of log management platforms such as ELK stack (Elasticsearch, Logstash, Kibana), write log parsing rules to extract key content, and assist in locating the root cause of the fault based on the extracted error information, exception stack, etc.; then proceed to step S19; S18, perform correlation analysis and diagnosis: construct an indicator correlation model (for example, using a directed graph structure, where nodes represent indicators and edges represent correlation relationships and weights), analyze the mutual influence between different monitoring indicators, and comprehensively consider multiple correlation indicators when a fault occurs to more accurately diagnose the fault point. Enter step S19 by combining multiple correlation indicators; S19, analyzing possible fault causes according to the matching results, and then entering S20; S20, setting a fault repair module, and then respectively entering S21, S22, S23, and S24 to perform fault repair according to the automatic repair strategy; S21, the automatic repair strategy is resource adjustment repair. Specifically, for the case of running in a virtualized environment, a request is sent to the cluster scheduler to dynamically expand resources (such as increasing container CPU limits and memory quotas) according to insufficient resources; in a physical server environment, system parameters are adjusted through scripts (such as ulimit settings) to optimize resource allocation; and enter S25; S22, the automatic repair strategy is a service restart type repair, which specifically adopts a service restart script written in a pre-configured order (first stop the dependent service, restart the target service, and then restart the dependent service). When a software failure occurs in the service, the restart operation is automatically executed, and the service status is continuously monitored after the restart, and then enters S25; S23, the automatic repair strategy is data repair type repair. According to the data backup strategy (full backup combined with incremental backup), when data consistency problems are found, data can be automatically extracted from the backup storage for recovery. At the same time, for real-time data processing scenarios, related tasks can be reasonably suspended and restarted to ensure data integrity and accuracy, and then enter S25; S24, the automatic repair strategy is network configuration repair, which can use network automation configuration tools (such as Ansible's network module) according to the preset network recovery template to automatically reconfigure network parameters and reset network connections in the event of a network failure to ensure network connectivity and enter S25; S25, performing repair operations through automated scripts (such as Shell scripts, Python scripts, etc.) or calling each component management API, recording detailed repair logs; and entering S26.

[0134] S26. Set a verification period after repair, continuously observe relevant monitoring indicators and service running status. If it returns to normal within the period, enter S27. Determine that the repair is successful; otherwise, if it does not return to normal, enter S28. Trigger an upgrade warning to notify manual intervention; and enter S29; S29. Set up an alarm and notification module, and enter S30 and S31; S30. Alarm policy setting: Classify alarms into different levels such as emergency, important, and general, and classify them according to the severity and urgency of the impact of the failure on the data platform. For example, the unavailability of the core database service is classified as an emergency alarm, and the slightly overdue execution of the ETL task is classified as a general alarm; and configure the notification frequency and object, then enter S32; S31. Integration of notification methods, integrate with common enterprise communication tools (such as SMS, phone, instant messaging tools like WeCom, DingTalk, and emails, etc.), configure different notification frequencies and notification objects for different levels of alarms, and enter step S32; S32. Achieve accurate and timely push of alarm messages.

[0135] As can be seen from the above, the present invention provides a data platform fault self-healing system and a corresponding method, aiming to realize automatic detection, diagnosis, and repair of data platform faults, and ensure the stable operation of the data platform.

[0136] Exemplary device As Figure 3 shown, an embodiment of the present invention provides a fault automatic detection, diagnosis, and processing device based on a data platform. The device includes: A fault monitoring module 310, configured to periodically collect system resource data at a set time interval; and analyze the collected system resource data to determine whether an abnormality occurs; A fault diagnosis module 320, configured to, when it is determined that the system resource data is abnormal, control to perform fault diagnosis, and analyze the cause of the fault by combining diagnosis based on preset rules, log analysis diagnosis, and / or correlation analysis diagnosis, and generate cause-of-fault data; A fault repair module 330, configured to, based on the generated cause-of-fault data, search for a preset automatic repair strategy corresponding to the generated cause-of-fault data, perform fault repair according to the repair plan recorded in the found automatic repair strategy, and output a fault repair result; An alarm and notification module 340, configured to verify the fault repair result, continuously detect relevant monitoring indicators and service running status after fault repair according to a set verification period; when the fault repair result is not repaired, control to give a warning and notification, as described above specifically.

[0137] Based on the above embodiments, the present invention also provides an intelligent terminal, and its principle block diagram can be asFigure 4 As shown in the figure. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a database connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for automatically detecting, diagnosing, and processing faults based on a data platform. The database of the intelligent terminal is used to store a program for automatically detecting, diagnosing, and processing faults based on a data platform.

[0138] Those skilled in the art can understand that Figure 4 the block diagram of the principle shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0139] In one embodiment, an intelligent terminal is provided, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Regularly collect system resource data at set time intervals; Analyze the collected system resource data to determine whether an abnormality occurs; When it is determined that the system resource data is abnormal, control the fault diagnosis to be performed, and analyze the cause of the fault by combining diagnosis based on preset rules, log analysis diagnosis, and / or correlation analysis diagnosis, and generate cause data of the fault; Based on the generated cause data of the fault, search for a preset automatic repair strategy corresponding to the generated cause data of the fault, perform fault repair according to the repair plan recorded in the found automatic repair strategy, and output the fault repair result; Verify the fault repair result, and continuously detect relevant monitoring indicators and service running status after the fault repair according to the set verification period; When it returns to normal within the period, it is determined that the repair is successful, otherwise an upgrade warning notification is triggered, as described above.

[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

Claims

1. A method for automatic fault detection and diagnosis based on a data platform, characterized in that: include: Collect system resource data regularly at set time intervals; Analyze the collected system resource data to determine whether any abnormality occurs; When it is determined that the system resource data is abnormal, the control performs fault diagnosis, adopts a combination of preset rule diagnosis, log analysis diagnosis and / or correlation analysis diagnosis to analyze the cause of the fault and generate fault cause data; Based on the generated fault cause data, a preset automatic repair strategy corresponding to the generated fault cause data is searched, the fault is repaired according to the repair plan recorded in the found automatic repair strategy, and the fault repair result is output.

2. The method for automatic fault detection and diagnosis based on a data platform according to claim 1 is characterized in that: The step of analyzing the collected system resource data to determine whether an abnormality occurs includes: Compare the collected system resource data with the threshold of the corresponding item to determine whether an abnormality occurs; When the collected system resource data exceeds the threshold of the corresponding item, it is judged that an abnormality has occurred.

3. The method for automatic fault detection and diagnosis based on a data platform according to claim 1 is characterized in that: The step of analyzing the collected system resource data to determine whether an abnormality occurs includes: Detect the connection status, query performance, master-slave synchronization data of the database service, as well as the startup status, execution time, error information data of the data processing task, and the message backlog, consumption rate, and number of connections of the message queue to determine whether there are any abnormalities; When one or more items in the connection status, query performance, master-slave synchronization data of the database service, the startup status, execution time, error information data of the data processing task, and the message backlog, consumption rate, and number of connections of the message queue do not meet the corresponding preset rules, it is determined that an abnormal situation has occurred.

4. The method for automatic fault detection and diagnosis based on a data platform according to claim 1 is characterized in that: When it is determined that the system resource data is abnormal, the control performs fault diagnosis, adopts a combination of preset rule diagnosis, log analysis diagnosis and / or correlation analysis diagnosis to analyze the cause of the fault, and generates the fault cause data, the steps include: When it is determined that the system resource data is abnormal, the control performs fault diagnosis; When using rule-based diagnosis to analyze the cause of a fault, historical fault data and corresponding monitoring indicators are obtained, and the machine learning rule learning algorithm is used to generate initial rules, and the rules in the knowledge base are continuously updated and optimized. When abnormal indicators are detected in the system resource data, the control is matched in the knowledge base, and the possible causes of the fault are analyzed based on the matching results, and the cause of the fault data is generated; When using log analysis to diagnose and analyze the cause of a fault, the ELK stack log management platform is used to collect log information of each component, and the error information and exception stack are extracted from the collected log information according to the preset log parsing rules; the root cause of the fault is located based on the extracted error information and exception stack, and the fault cause data is generated; When using correlation analysis-based diagnosis to analyze the cause of a fault, an indicator correlation model is built to analyze the mutual influence between different monitoring indicators. When abnormal system resource data is combined with the mutual influence between different monitoring indicators, the fault point is diagnosed and the cause of the fault data is generated.

5. The method for automatic fault detection and diagnosis based on a data platform according to claim 1, characterized in that: The step of searching for a preset automatic repair strategy corresponding to the generated fault cause data based on the generated fault cause data, and performing fault repair according to the repair solution recorded in the found automatic repair strategy comprises: Based on the generated fault cause data, the preset automatic repair strategy corresponding to the generated fault cause data is searched. When repairing resource adjustment problems, for the case of running in a virtualized environment, a request is sent to the cluster scheduler to dynamically expand resources according to the insufficient resources; in a physical server environment, system parameters are adjusted through scripts to optimize resource allocation; Based on the generated fault cause data, when the preset automatic repair strategy corresponding to the generated fault cause data is searched for a service restart type repair, the control automatically executes the restart operation when a software fault occurs in the service according to the preset service restart script, and controls the continuous monitoring of the service status after the restart; Based on the generated fault cause data, the preset automatic repair strategy corresponding to the generated fault cause data is searched. When the data repair type is a data repair, according to the pre-backed up data, when a data consistency problem is found, the data is automatically extracted from the pre-backed up data for recovery, and related tasks are reasonably suspended and restarted for real-time data processing scenarios; Based on the generated fault cause data, when the preset automatic repair strategy corresponding to the generated fault cause data is found to be network configuration repair, the network automatic configuration tool is controlled to automatically reconfigure network parameters and reset network connections according to the preset network recovery template when the network fails; According to the above automatic repair strategy, the repair operation is performed through automated scripts or by calling the management API of each component, and a detailed repair log is recorded.

6. The method for automatic fault detection and diagnosis based on a data platform according to claim 5 is characterized in that: The step of searching for a preset automatic repair strategy corresponding to the generated fault cause data based on the generated fault cause data, and performing fault repair according to the repair solution recorded in the found automatic repair strategy also includes: Verify the fault repair results and continuously detect the relevant monitoring indicators and service operation status after the fault repair according to the set verification cycle; If it returns to normal within the period, the repair is considered successful, otherwise an upgrade alarm notification is triggered.

7. The method for automatic fault detection and diagnosis based on a data platform according to claim 1, characterized in that: After the step of outputting the fault repair result, the following steps are further included: When the fault repair result is not repaired, the control will warn and notify.

8. A data platform-based automatic fault detection, diagnosis and processing device, characterized in that: The device comprises: The fault monitoring module is used to collect system resource data at set time intervals and analyze the collected system resource data to determine whether there is any abnormality; A fault diagnosis module is used to control fault diagnosis when it is determined that the system resource data is abnormal, analyze the cause of the fault by combining diagnosis based on preset rules, log analysis diagnosis and / or correlation analysis diagnosis, and generate fault cause data; A fault repair module is used to find a preset automatic repair strategy corresponding to the generated fault cause data based on the generated fault cause data, perform fault repair according to the repair solution recorded in the found automatic repair strategy, and output the fault repair result; The alarm and notification module is used to verify the fault repair results. According to the set verification cycle, it continuously detects the relevant monitoring indicators and service operation status after the fault is repaired; when the fault repair result is not repaired, the control will issue a warning and notification.

9. An intelligent terminal, characterized in that: The device comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include being used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Cross-system fault diagnosis method and system combined with multi-dimensional anomaly detection

    CN120892238A

  • Fault repairing method and device for full flash memory storage system

    CN121029472A

  • Method and apparatus for failure recovery of all-flash storage system

    CN121029472B

  • Data processing method and device, electronic equipment, storage medium and program product

    CN121597461A