Fault self-healing method and device, electronic equipment and computer readable storage medium

By acquiring and analyzing host hardware monitoring metrics and service logs, fault prediction and self-healing are performed, solving the data processing reliability problem of distributed file systems in the event of hardware failure, and improving the stability and reliability of the system.

CN120973570APending Publication Date: 2025-11-18PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511072309.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing distributed file systems cannot predict hardware failures, leading to reduced data processing reliability.

Method used

By acquiring host hardware monitoring metrics and service logs, we analyze and process them to obtain fault prediction information. When the disk lifespan is below the threshold, we migrate data blocks to healthy nodes, generate data service fault information, migrate data service containers based on recovery strategies, and finally synchronize data when the disaster recovery system detects path changes.

Benefits of technology

It enables accurate prediction and self-healing of hardware failures, improving the reliability of data processing and the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973570A_ABST
    Figure CN120973570A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fault processing and the field of insurance business and smart medical treatment, and provides a fault self-recovery method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining a host hardware monitoring index and a service log; analyzing and processing the host hardware monitoring index and the service log to obtain fault prediction information; when the fault prediction information represents that the service life of the disk is lower than a preset threshold value, migrating a data block on the fault node to a healthy node, and generating data service fault information; migrating a data service container to a healthy node based on the data service fault information and a service recovery strategy to recover the data service and change a metadatabase path; and when the disaster recovery system monitors that the path of the metadatabase is changed, performing data synchronization operation on the disaster recovery cluster in the disaster recovery system. According to the technical scheme, fault prediction and fault self-healing repair can be accurately carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to, but are not limited to, the field of fault handling technology, and in particular to a fault self-healing method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] In the current fields of smart healthcare and insurance, there is often a need to store and process large amounts of data. Distributed file systems can store and process large-scale datasets, with high throughput and high fault tolerance, making them very suitable for storing and computing large-scale data. Distributed file systems use a replication mechanism by default, achieving hardware failure tolerance by storing data through multiple borrowing points. However, hardware monitoring is lagging behind, relying on hardware failures to trigger alarms, and cannot predict and handle failures, thus reducing the reliability of data processing. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0004] To address the problems mentioned in the background section, this application provides a fault self-healing method, apparatus, electronic device, and computer-readable storage medium, which can accurately predict and repair faults, thereby improving the reliability of data processing.

[0005] In a first aspect, embodiments of this application provide a fault self-healing method, including:

[0006] Obtain host hardware monitoring metrics and service logs;

[0007] The host hardware monitoring metrics and service logs are analyzed and processed to obtain fault prediction information.

[0008] When the fault prediction information indicates that the disk lifespan is lower than a preset threshold, the data blocks on the faulty node corresponding to the disk are migrated to healthy nodes, and data service fault information is generated.

[0009] Based on the data service failure information and the preset service recovery strategy, the data service container corresponding to the disk is migrated to the healthy node to restore the data service and change the metadata database path.

[0010] If the preset disaster recovery system detects a change in the metadata database path, it will perform a data synchronization operation on the disaster recovery cluster in the disaster recovery system.

[0011] Secondly, embodiments of this application also provide a fault self-healing device, the device comprising:

[0012] The acquisition unit is used to acquire host hardware monitoring metrics and service logs.

[0013] The analysis unit is used to analyze and process the host hardware monitoring indicators and the service logs to obtain fault prediction information.

[0014] The first migration unit is used to migrate data blocks on the faulty node corresponding to the disk to a healthy node when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, and to generate data service fault information.

[0015] The second migration unit is used to migrate the data service container corresponding to the disk to the healthy node based on the data service failure information and the preset service recovery strategy, so as to restore the data service and change the metadata database path.

[0016] The synchronization unit is used to perform data synchronization operations on the disaster recovery cluster in the preset disaster recovery system when the metadata database path is detected to have changed.

[0017] Thirdly, embodiments of this application also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the fault self-healing method described in the first aspect above.

[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for performing the fault self-healing method described in the first aspect above.

[0019] The fault self-healing method according to the embodiments provided in this application has at least the following beneficial effects: During the fault self-healing process, host hardware monitoring indicators and service logs are first acquired; then, fault prediction information is obtained by analyzing and processing the host hardware monitoring indicators and service logs; next, when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, the data blocks on the faulty node corresponding to the disk are migrated to healthy nodes, and data service fault information is generated; then, based on the data service fault information and a preset service recovery strategy, the data service container corresponding to the disk is migrated to healthy nodes to restore data services and cause a change in the metadata path; when the pre-set disaster recovery system monitors the change in the metadata path, data synchronization operations can be performed on the disaster recovery cluster in the disaster recovery system. Through the above technical solution, fault prediction and fault self-healing repair can be accurately performed, improving the reliability of data processing. Attached Figure Description

[0020] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0021] Figure 1 This is a flowchart illustrating a fault self-healing method provided in one embodiment of this application;

[0022] Figure 2 yes Figure 1 A schematic diagram of a specific implementation method of step S200;

[0023] Figure 3 yes Figure 1 A schematic diagram of a specific implementation of step S300;

[0024] Figure 4 yes Figure 1 A schematic diagram of a specific implementation of step S400;

[0025] Figure 5 yes Figure 1 A schematic diagram of a specific implementation of step S500;

[0026] Figure 6 This is a flowchart illustrating a fault self-healing method provided in another embodiment of this application;

[0027] Figure 7 Is it completed? Figure 1 A flowchart illustrating a specific implementation method following step S200;

[0028] Figure 8 This is a schematic diagram of a fault self-healing device provided in one embodiment of this application;

[0029] Figure 9 This is a schematic diagram of an electronic device provided in one embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0031] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0032] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0033] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0034] AI is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Artificial intelligence can simulate the information processes of human consciousness and thought. Furthermore, artificial intelligence utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results—the theories, methods, technologies, and application systems available for use.

[0035] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0036] Artificial intelligence, or AI, is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0037] The servers involved in artificial intelligence technology can be standalone servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0038] This application provides a fault self-healing method, apparatus, electronic device, and computer-readable storage medium. In the fault self-healing process, firstly, host hardware monitoring indicators and service logs are acquired; then, fault prediction information is obtained by analyzing and processing the host hardware monitoring indicators and service logs; next, when the fault prediction information indicates that the disk's lifespan is below a preset threshold, data blocks on the faulty node corresponding to the disk are migrated to healthy nodes, and data service fault information is generated; then, based on the data service fault information and a preset service recovery strategy, the data service container corresponding to the disk is migrated to healthy nodes to restore data service and change the metadata database path; when the pre-defined disaster recovery system detects the change in the metadata database path, data synchronization operations can be performed on the disaster recovery cluster in the disaster recovery system. Through the above technical solution, fault prediction and fault self-healing repair can be accurately performed, improving the reliability of data processing.

[0039] The fault self-healing method provided in this application relates to the field of fault handling technology. This fault self-healing method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0040] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0041] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0042] The embodiments of this application will be further described below with reference to the accompanying drawings.

[0043] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a fault self-healing method provided in one embodiment of this application. The fault self-healing method includes the following steps:

[0044] Step S100: Obtain host hardware monitoring metrics and service logs.

[0045] The fault self-healing method provided in this application first obtains host hardware monitoring indicators and service logs during the fault self-healing process. After obtaining the host hardware monitoring indicators and service logs, it can prepare for subsequent fault self-healing processing.

[0046] It's worth noting that host hardware monitoring metrics can include read / write error rate, remapped sector count, temperature, and power-on time. Read / write error rate refers to the ratio of the number of errors that occur during data read or write operations on hardware devices (such as hard drives and memory) to the total number of read / write operations. It reflects the reliability of the hardware during data transmission. For example, for a hard drive, when the operating system sends a read command, the hard drive may fail to read the data correctly due to head damage, disk surface scratches, signal interference, etc. Similarly, when writing data, errors may occur due to head write failures, disk sector damage, etc. Read / write error rate is usually expressed as a percentage or the number of errors per million operations. Remapped sector count refers to the number of times the hard drive migrates the data in a sector to a spare sector (spare storage area) after detecting a read / write error. When a sector on the hard drive suffers physical damage (such as disk surface scratches, head damage, etc.), the hard drive marks that sector as a "bad sector" and rewrites the data to a spare sector. Host hardware temperature is one of the important indicators for measuring the operating status of a computer system. Monitoring hardware temperature can help detect potential heat dissipation problems in a timely manner, avoiding performance degradation, system instability, or even hardware damage caused by overheating. Host hardware power-on time refers to the total amount of time that hardware devices have been powered on and running since they left the factory. For example, the power-on time of a hard drive can be viewed through self-monitoring, analysis, and reporting technologies. Host hardware power-on time reflects the total amount of time that the hard drive has been driven by a power supply since it left the factory.

[0047] It's worth noting that service logs can include runtime logs, audit logs, process startup logs, and garbage collection logs. Runtime logs are files or collections of data that record various events, operations, and state changes that occur during the operation of a system, application, or device; they are crucial for system management, troubleshooting, performance optimization, and security monitoring. Audit logs are a special type of runtime log used to record detailed events related to security, compliance, and operational procedures within a system, application, or organization; they are typically used to track user behavior, system operations, data access, and changes to ensure system security and compliance. Process startup logs record detailed information about process startup events in a system or application, typically including the process startup time, the user who started the process, and the command-line parameters used; they are also crucial for system management, troubleshooting, security monitoring, and performance optimization. Garbage collection logs are log files that record detailed information about garbage collection operations; garbage collection is a mechanism used by the programming language runtime environment to automatically manage memory, reclaiming memory space occupied by unused objects to prevent memory leaks and insufficient memory.

[0048] It is worth noting that in the medical field, various systems can be used to store, manage, and process relevant medical data. Each system can correspond to one or more devices. Therefore, during system maintenance in the medical field, the host hardware monitoring indicators and service logs of each device in the system can be monitored, acquired, and processed. Similarly, in the insurance business, various systems can be used to store, manage, and process relevant insurance business data. Each system can also correspond to one or more devices. Therefore, during system maintenance in the insurance business, the host hardware monitoring indicators and service logs of each device in the system can be monitored, acquired, and processed to prepare for subsequent fault self-healing.

[0049] It is worth noting that user permission or consent is obtained before acquiring host hardware monitoring metrics and service logs. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. Additionally, when this application embodiment needs to acquire sensitive user personal information, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after explicitly obtaining the user's separate permission or consent is the necessary user-related data acquired to enable the normal operation of this application embodiment.

[0050] Step S200: Analyze and process the host hardware monitoring indicators and service logs to obtain fault prediction information.

[0051] The fault self-healing method provided in this application, after obtaining the host hardware monitoring indicators and service logs, can analyze and process the host hardware monitoring indicators and service logs to obtain fault prediction information, so as to detect and prevent potential faults in the system.

[0052] It is worth noting that after obtaining fault prediction information, if the fault prediction information indicates that the disk's lifespan is lower than a pre-set threshold, the data blocks on the faulty node corresponding to the disk can be migrated to a healthy node, and corresponding data service fault information can be generated. Subsequently, based on the data service fault information and a pre-set service recovery strategy, the data service container corresponding to the disk can be migrated to a healthy node to restore the data service and change the metadata database path, thereby achieving fault detection and self-healing.

[0053] It is worth noting that in the process of analyzing and processing host hardware monitoring indicators and service logs to obtain fault prediction information, the read / write error rate, number of remapped sectors, temperature, and power-on time are determined from the host hardware monitoring indicators; and the operation log, audit log, process startup log, and garbage collection log are determined from the service logs. Then, based on a preset historical detection standard library, the read / write error rate, number of remapped sectors, temperature, and power-on time are subjected to a first detection process to obtain first detection information; and based on a preset historical fault log knowledge base, the operation log, audit log, process startup log, and garbage collection log are subjected to a second detection process to obtain second detection information; finally, the first and second detection information are comprehensively judged and processed to obtain fault prediction information, which prepares for subsequent fault self-healing.

[0054] like Figure 2 As shown, analyzing and processing host hardware monitoring metrics and service logs to obtain fault prediction information can include the following steps:

[0055] Step S210: Determine the read / write error rate, number of remapped sectors, temperature, and power-on time from the host hardware monitoring metrics; and determine the operation log, audit log, process startup log, and garbage collection log from the service log;

[0056] Step S220: Based on a preset historical detection standard library, perform first detection processing on read / write error rate, number of remapped sectors, temperature and power-on time to obtain first detection information; and based on a preset historical fault log knowledge base, perform second detection processing on running log, audit log, process startup log and garbage collection log to obtain second detection information.

[0057] Step S230: The first detection information and the second detection information are combined and processed to obtain fault prediction information.

[0058] For steps S210 to S230, in the process of analyzing and processing the host hardware monitoring indicators and service logs to obtain fault prediction information, firstly, the read / write error rate, number of remapped sectors, temperature, and power-on time are determined from the host hardware monitoring indicators; and secondly, the operation log, audit log, process startup log, and garbage collection log are determined from the service logs. Then, based on a preset historical detection standard library, the read / write error rate, number of remapped sectors, temperature, and power-on time are subjected to a first detection process to obtain first detection information. And, based on a preset historical fault log knowledge base, the operation log, audit log, process startup log, and garbage collection log are subjected to a second detection process to obtain second detection information. Finally, the first detection information and the second detection information are comprehensively judged and processed to obtain fault prediction information, which prepares for subsequent fault self-healing.

[0059] It is worth noting that the historical detection standard library stores standard values ​​for read / write error rates, standard remapped sector counts, standard temperatures, and standard power-on times. By comparing the read / write error rate with the standard values, the standard remapped sector count with the standard values, the temperature with the standard temperature, and the power-on time with the standard power-on time, the first detection information can be determined. The historical fault log knowledge base stores historical runtime fault logs, historical audit fault logs, historical process startup fault logs, and historical garbage collection fault logs. By comparing the runtime logs with historical runtime fault logs, historical audit fault logs with audit logs, historical process startup fault logs with process startup logs, and historical garbage collection fault logs with garbage collection logs, the second detection information can be determined. Finally, by comprehensively judging and processing the first and second detection information, fault prediction information can be obtained, thereby achieving comprehensive and intelligent fault detection and processing.

[0060] For example, in the medical field, read / write error rate, remapped sector count, temperature, and power-on time are determined from the host hardware monitoring indicators associated with the medical system; and operation logs, audit logs, process startup logs, and garbage collection logs are determined from the service logs associated with the medical system. Subsequently, the host hardware monitoring indicators associated with the medical system can be compared against historical testing standards, and the service logs associated with the medical system can be compared against historical operational failure logs to finally determine fault prediction information. Alternatively, in the insurance business field, read / write error rate, remapped sector count, temperature, and power-on time are determined from the host hardware monitoring indicators of the insurance business system; and operation logs, audit logs, process startup logs, and garbage collection logs are determined from the service logs of the insurance business system. Subsequently, the host hardware monitoring indicators of the insurance business system can be compared against historical testing standards, and the service logs of the insurance business system can be compared against historical operational failure logs to finally determine fault prediction information.

[0061] Step S300: When the fault prediction information indicates that the disk lifespan is lower than a preset threshold, the data blocks on the faulty node corresponding to the disk are migrated to the healthy node, and data service fault information is generated.

[0062] The fault self-healing method provided in this application, after obtaining fault prediction information, can migrate data blocks on the faulty node corresponding to the disk to healthy nodes when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, in order to prepare for subsequent fault recovery; and will also generate data service fault information to prepare for subsequent data service container migration.

[0063] It is worth noting that if the disk lifespan is lower than a preset threshold, the data blocks on the faulty node corresponding to the disk need to be migrated to the healthy node in advance to achieve fault prevention. In some embodiments of this application, the distributed file system includes faulty nodes and healthy nodes.

[0064] like Figure 3 As shown, when the fault prediction information indicates that the disk lifespan is below a preset threshold, migrating the data blocks on the faulty node corresponding to the disk to a healthy node may include the following steps:

[0065] Step S310: When the fault prediction information indicates that the disk lifespan is lower than a preset threshold, the node corresponding to the disk in the preset distributed file system is designated as the fault node.

[0066] Step S320: Migrate the data blocks on the faulty node to the healthy node, where the healthy node is a node that is running normally in the distributed file system.

[0067] For steps S310 to S320, when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, during the process of migrating the data blocks on the faulty node corresponding to the disk to the healthy node, the node corresponding to the disk in the pre-defined distributed file system will be designated as the faulty node when the fault prediction information indicates that the disk lifespan is lower than the preset threshold; then the data blocks on the faulty node will be migrated to the healthy node to repair the potential fault.

[0068] It is worth noting that in some embodiments of this application, the distributed file system includes healthy nodes and faulty nodes. Healthy nodes are nodes that are running normally in the distributed file system, and faulty nodes are nodes that are running abnormally in the distributed file system. The preset threshold can be set according to actual needs. Migrating data blocks from faulty nodes to healthy nodes allows the migrated data blocks to continue providing relevant data services, ensuring the system can operate normally.

[0069] Step S400: Based on the data service failure information and the preset service recovery strategy, migrate the data service container corresponding to the disk to a healthy node to restore the data service and change the metadata database path.

[0070] The fault self-healing method provided in this application, when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, migrates the data blocks on the faulty node corresponding to the disk to a healthy node and generates data service fault information. After obtaining the data service fault information, the data service container corresponding to the disk can be migrated to a healthy node based on the data service fault information and a pre-set service recovery strategy, so that the data service can be restored. When the data service container is migrated to the healthy node, the metadata path will change, which can trigger the subsequent data synchronization operation in the disaster recovery system to achieve fault repair.

[0071] It is worth noting that, in the process of migrating the data service container corresponding to the disk to a healthy node based on data service failure information and a pre-set service recovery strategy, matching the data service failure information with the service recovery strategy yields the target recovery strategy. Then, the target data service container is determined from the disk according to the target recovery strategy, and finally, the target data service container can be migrated to the corresponding healthy node. This technical solution makes self-healing and repair of faults simpler and faster.

[0072] It is worth noting that the service recovery strategy includes multiple recovery strategies. Therefore, by matching the data service failure information with the service recovery strategy, the target recovery strategy can be obtained. Then, the target data service container can be determined from the disk according to the target recovery strategy. Finally, the target data service container is migrated to the corresponding healthy node to achieve self-healing repair of the failure.

[0073] like Figure 4 As shown, based on data service failure information and preset service recovery strategies, migrating the data service container corresponding to the disk to a healthy node may include the following steps:

[0074] Step S410: Match the data service failure information with the service recovery strategy to obtain the target recovery strategy;

[0075] Step S420: Determine the target data service container from the disk according to the target recovery strategy;

[0076] Step S430: Migrate the target data service container to a healthy node.

[0077] For steps S410 to S430, during the process of migrating the data service container corresponding to the disk to a healthy node based on data service fault information and a pre-set service recovery strategy, the data service fault information and the service recovery strategy are matched to obtain the target recovery strategy. Then, the target data service container is determined from the disk according to the target recovery strategy, and finally, the target data service container is migrated to the corresponding healthy node. This technical solution makes self-healing and repair of faults simpler and faster.

[0078] It is worth noting that the service recovery strategy includes multiple recovery strategies. By matching the data service failure information with the multiple recovery strategies in the service recovery strategy, the target recovery strategy can be determined. Subsequently, the target data service container can be determined from the disk according to the target recovery strategy. Then, the target data service container can be migrated to a healthy node, thereby achieving stable and reliable self-healing and repair of the failure.

[0079] Step S500: If the meta database path changes as detected by the preset disaster recovery system, perform data synchronization operation on the disaster recovery cluster in the disaster recovery system.

[0080] The fault self-healing method provided in this application, based on data service fault information and preset service recovery strategies, migrates the data service container corresponding to the disk to a healthy node to restore the data service and causes the metadata path to change. When the disaster recovery system detects the change in the metadata path, it can perform data synchronization operations on the disaster recovery cluster in the disaster recovery system, thereby ensuring data integrity, improving system availability, enhancing data security, and supporting disaster recovery.

[0081] It's worth noting that a disaster recovery system is a comprehensive solution that uses hardware, software, network, and storage technologies, combined with management strategies and processes, to enable information systems to quickly resume operation after a disaster. A disaster recovery cluster is the core component of a disaster recovery system. It uses a cluster architecture composed of multiple nodes (usually servers) to achieve redundant data storage and high business availability. The design goal of a disaster recovery cluster is to quickly switch to a backup node when the primary node fails, ensuring business continuity, while simultaneously ensuring data integrity and consistency through data synchronization and other technologies. For example, in the healthcare industry, if the disaster recovery system in a healthcare system detects a change in the metadata database path, data synchronization operations need to be performed on the disaster recovery cluster within that system. Similarly, in the insurance business, if the disaster recovery system in an insurance business system detects a change in the metadata database path, data synchronization operations need to be performed on the disaster recovery cluster within that system.

[0082] like Figure 5 As shown, when the preset disaster recovery system detects a change in the metadata database path, the data synchronization operation on the disaster recovery cluster in the disaster recovery system may include the following steps:

[0083] Step S510: After the disaster recovery system detects a change in the metadata database path, determine the variable data;

[0084] Step S520: Perform data synchronization operation on the disaster recovery cluster in the disaster recovery system based on the variable data.

[0085] For steps S510 to S520, when the preset disaster recovery system detects a change in the metadata database path and performs data synchronization operations on the disaster recovery cluster in the disaster recovery system, the variable data can be determined when the disaster recovery system detects a change in the metadata database path; then, based on the variable data, the disaster recovery cluster in the disaster recovery system is synchronized to ensure data integrity and improve system availability.

[0086] It's worth noting that disaster recovery clusters typically consist of multiple nodes, and data synchronization ensures data consistency between the master and slave nodes. For example, in financial systems, transaction data needs to be synchronized to all nodes in the disaster recovery cluster in real time. This ensures that both master and slave nodes accurately reflect user account balances, transaction records, and other information, preventing business errors caused by data inconsistencies. Data synchronization also supports load balancing; by synchronizing data across multiple nodes, business traffic can be rationally distributed across different nodes, preventing excessive load on a single node from causing system performance degradation or even crashes.

[0087] like Figure 6 As shown, the fault self-healing method in this application embodiment may include the following steps:

[0088] Step S610: Monitor the disaster recovery data in the disaster recovery system;

[0089] Step S620: If the time during which disaster recovery data is not used reaches a preset time threshold, the corresponding disaster recovery data is transferred to the target cold database of the disaster recovery system for storage processing.

[0090] For steps S610 to S620, in some embodiments of this application, disaster recovery data in the disaster recovery system can be monitored and processed; when the time during which the disaster recovery data is not used reaches a preset time threshold, the corresponding disaster recovery data can be transferred to the target cold database of the disaster recovery system for storage processing, thereby reducing storage costs, reducing backup costs, releasing hot storage resources, and improving system performance.

[0091] It is worth noting that the time threshold in this embodiment can be set according to actual needs. When the time during which disaster recovery data has not been used reaches the preset time threshold, the corresponding disaster recovery data can be transferred to the target cold database of the disaster recovery system for storage processing.

[0092] like Figure 7 As shown, after analyzing and processing the host hardware monitoring metrics and service logs to obtain fault prediction information, the following steps can be included:

[0093] Step S330: If the fault prediction information indicates that the disk usage rate is greater than a preset usage threshold, generate disk expansion suggestion information;

[0094] Step S340: Send disk expansion suggestion information to the preset control terminal.

[0095] For steps S330 to S340, after analyzing and processing the host hardware monitoring indicators and service logs to obtain fault prediction information, if the fault prediction information indicates that the disk usage rate is greater than the preset usage threshold, disk expansion suggestion information will be generated. Then, the disk expansion suggestion information can be sent to the pre-set control terminal to remind the user that the corresponding disk needs to be expanded.

[0096] It is worth noting that the control terminal in this embodiment can be a mobile phone, computer, tablet, or other control terminal, and is not limited here. The preset usage threshold can be set according to actual needs, and is also not limited here. For example, in the medical industry, when fault prediction information indicates that disk usage exceeds a preset usage threshold, disk expansion suggestion information is generated; then, the disk expansion information is sent to the control terminal of the medical system to notify relevant system maintenance personnel. Alternatively, in the insurance industry, when fault prediction information indicates that disk usage exceeds a preset usage threshold, disk expansion suggestion information is generated; then, the disk expansion information is sent to the control terminal of the insurance system to notify relevant system maintenance personnel.

[0097] In addition, such as Figure 8 As shown, one embodiment of this application also provides a fault self-healing device 10, the device comprising:

[0098] Acquisition unit 100 is used to acquire host hardware monitoring metrics and service logs;

[0099] Analysis unit 200 is used to analyze and process host hardware monitoring indicators and service logs to obtain fault prediction information;

[0100] The first migration unit 300 is used to migrate data blocks on the faulty node corresponding to the disk to a healthy node when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, and to generate data service fault information.

[0101] The second migration unit 400 is used to migrate the data service container corresponding to the disk to a healthy node based on the data service failure information and the preset service recovery strategy, so as to restore the data service and change the metadata database path.

[0102] Synchronization unit 500 is used to perform data synchronization operations on the disaster recovery cluster in the disaster recovery system when the meta database path is detected to have changed in the preset disaster recovery system.

[0103] It should be noted that during the fault self-healing process, the first step is to acquire host hardware monitoring metrics and service logs. Then, analyzing and processing these metrics and logs yields fault prediction information. Next, if the fault prediction indicates that the disk's lifespan is below a preset threshold, the data blocks on the faulty node corresponding to the disk are migrated to healthy nodes, generating data service fault information. Based on this fault information and a preset service recovery strategy, the data service container corresponding to the disk is migrated to a healthy node to restore data service and change the metadata path. Once the pre-defined disaster recovery system detects the change in the metadata path, data synchronization can be performed on the disaster recovery cluster within the system. This technical solution enables accurate fault prediction and self-healing, improving the reliability of data processing.

[0104] The specific implementation of the fault self-healing device 10 is basically the same as the specific embodiment of the fault self-healing method described above, and will not be repeated here.

[0105] In addition, such as Figure 9 As shown, one embodiment of this application also provides an electronic device 700, which includes: a memory 720, a processor 710, and a computer program stored on the memory 720 and executable on the processor 710.

[0106] The processor 710 and memory 720 can be connected via a bus or other means.

[0107] The non-transient software program and instructions required to implement the fault self-healing method of the above embodiments are stored in the memory 720. When executed by the processor 710, the fault self-healing method of each of the above embodiments is executed.

[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0109] Furthermore, one embodiment of this application provides a computer-readable storage medium storing computer-executable instructions that are executed by a processor 710 or a controller, for example, by a processor 710 in the above-described device embodiment, such that the processor 710 performs the fault self-healing method in the above-described embodiment.

[0110] The above embodiments can be used in combination, and modules with the same name in different embodiments may be the same or different.

[0111] The foregoing has described specific embodiments of this application; other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily have to follow the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0112] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and computer-readable storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0113] The apparatus, device, computer-readable storage medium and method provided in the embodiments of this application are corresponding. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and computer storage medium will not be described again here.

[0114] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0115] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91 SAM, Microchip PIC18F26K20, and Silicon Labs C8051 F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0116] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0117] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing the embodiments of this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0118] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0119] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0122] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0123] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0124] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0125] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0126] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0127] The embodiments of this application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. The embodiments of this application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.

[0128] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0129] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A fault self-healing method, characterized in that, The method includes: Obtain host hardware monitoring metrics and service logs; The host hardware monitoring metrics and service logs are analyzed and processed to obtain fault prediction information. When the fault prediction information indicates that the disk lifespan is lower than a preset threshold, the data blocks on the faulty node corresponding to the disk are migrated to healthy nodes, and data service fault information is generated. Based on the data service failure information and the preset service recovery strategy, the data service container corresponding to the disk is migrated to the healthy node to restore the data service and change the metadata database path. If the preset disaster recovery system detects a change in the metadata database path, it will perform a data synchronization operation on the disaster recovery cluster in the disaster recovery system.

2. The fault self-healing method according to claim 1, characterized in that, The analysis and processing of the host hardware monitoring metrics and the service logs to obtain fault prediction information includes: The read / write error rate, number of remapped sectors, temperature, and power-on time are determined from the host hardware monitoring metrics; and the operation log, audit log, process startup log, and garbage collection log are determined from the service log. Based on a preset historical detection standard library, the read / write error rate, the number of remapped sectors, the temperature, and the power-on time are subjected to a first detection process to obtain first detection information; and based on a preset historical fault log knowledge base, the operation log, the audit log, the process startup log, and the garbage collection log are subjected to a second detection process to obtain second detection information. The first detection information and the second detection information are combined and processed to obtain fault prediction information.

3. The fault self-healing method according to claim 1, characterized in that, When the fault prediction information indicates that the disk lifespan is below a preset threshold, the step of migrating the data blocks on the faulty node corresponding to the disk to a healthy node includes: When the fault prediction information indicates that the disk lifespan is lower than the preset threshold, the node corresponding to the disk in the preset distributed file system is identified as the fault node. The data blocks on the faulty node are migrated to a healthy node, wherein the healthy node is a node that is running normally in the distributed file system.

4. The fault self-healing method according to claim 1, characterized in that, The step of migrating the data service container corresponding to the disk to the healthy node based on the data service failure information and the preset service recovery strategy includes: The data service failure information is matched with the service recovery strategy to obtain the target recovery strategy; The target data service container is determined from the disk according to the target recovery strategy; Migrate the target data service container to the healthy node.

5. The fault self-healing method according to claim 1, characterized in that, The step of detecting a change in the metadata database path in the preset disaster recovery system and performing data synchronization operations on the disaster recovery cluster in the disaster recovery system includes: The disaster recovery system detects a change in the metadata database path and identifies the variable data. The disaster recovery cluster in the disaster recovery system is synchronized based on the variable data.

6. The fault self-healing method according to claim 1, characterized in that, The method further includes: Monitor the disaster recovery data in the disaster recovery system; If the disaster recovery data is not used for a preset time threshold, the corresponding disaster recovery data will be transferred to the target cold database of the disaster recovery system for storage processing.

7. The fault self-healing method according to claim 1, characterized in that, After analyzing and processing the host hardware monitoring metrics and the service logs to obtain fault prediction information, the method further includes: When the fault prediction information indicates that the disk usage rate is greater than a preset usage threshold, disk expansion suggestion information is generated. The disk expansion suggestion information is sent to a preset control terminal.

8. A self-healing fault device, characterized in that, The device includes: The acquisition unit is used to acquire host hardware monitoring metrics and service logs. The analysis unit is used to analyze and process the host hardware monitoring indicators and the service logs to obtain fault prediction information. The first migration unit is used to migrate data blocks on the faulty node corresponding to the disk to a healthy node when the fault prediction information indicates that the disk lifespan is lower than a preset threshold, and to generate data service fault information. The second migration unit is used to migrate the data service container corresponding to the disk to the healthy node based on the data service failure information and the preset service recovery strategy, so as to restore the data service and change the metadata database path. The synchronization unit is used to perform data synchronization operations on the disaster recovery cluster in the preset disaster recovery system when the metadata database path is detected to have changed.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the fault self-healing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the fault self-healing method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Safe atomic power fault detection self-healing method and device, electronic equipment and storage medium

    CN121435222A