Log management method, electronic equipment, storage medium and product

By determining the fault level in the server cluster and selecting the appropriate log transmission interface, the problem of log collection instability in complex network environments is solved, and the hierarchical transmission and recovery of fault logs is realized, which improves operation and maintenance reliability.

CN120045434AActive Publication Date: 2025-05-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510519964.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing technology is difficult to ensure the stability and reliability of log collection in complex network environments, resulting in the loss of key alarm information and increasing the difficulty of operation and maintenance.

Method used

By determining the server fault level, selecting the corresponding log transmission interface, and transmitting the fault log to the target server through multiple transmission channels, ensuring the storage and recovery of logs.

Benefits of technology

It realizes the hierarchical transmission of fault logs in the server cluster environment, avoids the risk of single points of fault, ensures the integrity and recovery of fault logs, and improves the operation and maintenance reliability of server clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045434A_ABST
    Figure CN120045434A_ABST
Patent Text Reader

Abstract

The invention discloses a log management method, electronic equipment, a storage medium and a product, and relates to the technical field of computers, the method comprises the following steps: when a first server breaks down, determining a fault level of the first server; determining a log transmission interface corresponding to the fault level; transmitting the fault log in the first server to a target server through a transmission channel corresponding to the log transmission interface, so that the target server stores the fault log; and when fault recovery of the first server is completed, log recovery is performed based on the fault log stored in the target server, so that the first server stores the recovered fault log, fault log hierarchical transmission in a server cluster environment is realized, a single-point fault risk is effectively avoided, and the fault recovery efficiency is improved even if a complex network environment fluctuates or is interrupted. And the integrity and the restorability of the fault log can still be ensured, so that the operation and maintenance reliability of the server cluster is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a log management method, an electronic device, a storage medium, and a product. Background Art

[0002] With the continuous development of server technology, the scale of server clusters has become increasingly large. While this large-scale cluster architecture improves system performance and processing capabilities, it also brings many problems to the maintainability of the cluster. When a server fails, a fault log is generated and transmitted to the management server via a network or a specific interface. The management server integrates the information and then feedbacks it to the operation and maintenance personnel for fault troubleshooting.

[0003] Currently, related technologies mainly focus on decentralized log storage and data verification. In terms of data transmission, related technologies still rely on the cluster management server and the central network node to build a data transmission channel. It can be seen that the transmission path of related technologies is single, relying on the cluster management server and the central network node. In the face of a complex network environment, it is difficult to ensure the stability and reliability of log collection, increasing the operation and maintenance difficulty. Summary of the Invention

[0004] This application provides a log management method, an electronic device, a storage medium, and a product to at least solve the problems of single data transmission and easy loss of key alarm information in related technologies.

[0005] This application provides a log management method, including: When a first server fails, determining the fault level of the first server; Determining a log transmission interface corresponding to the fault level; Transmitting the fault log in the first server to a target server through a transmission channel corresponding to the log transmission interface, so that the target server stores the fault log; When the fault repair of the first server is completed, performing log recovery based on the fault log stored in the target server, so that the first server stores the recovered fault log.

[0006] This application also provides a log management device, including: A level determination unit, configured to determine the fault level of the first server when the first server fails; An interface determination unit, configured to determine a log transmission interface corresponding to the fault level; A transmission unit, configured to transmit the fault log in the first server to a target server through a transmission channel corresponding to the log transmission interface, so that the target server stores the fault log; A recovery unit, configured to perform log recovery based on the fault logs stored in the target server when the repair of the first server's fault is completed, so that the first server stores the recovered fault logs.

[0007] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above log management methods when executing the computer program.

[0008] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above log management methods when executed by a processor.

[0009] This application also provides a computer program product including a computer program, which implements the steps of any of the above log management methods when executed by a processor.

[0010] Through this application, when a fault occurs in the first server, the fault level of the first server is determined; the log transmission interface corresponding to the fault level is determined; the fault logs in the first server are transmitted to the target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the fault logs; when the repair of the first server's fault is completed, log recovery is performed based on the fault logs stored in the target server, so that the first server stores the recovered fault logs, realizing hierarchical transmission of fault logs in a server cluster environment, effectively avoiding the risk of single-point failure, and ensuring the integrity and recoverability of fault logs even when fluctuations or interruptions occur in a complex network environment, thereby significantly improving the operation and maintenance reliability of the server cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a schematic flowchart of a log management method provided by an embodiment of this application; Figure 2 It is a schematic flowchart of a log management method provided by an embodiment of this application; Figure 3 It is a schematic flowchart of a log management method provided by an embodiment of this application; Figure 4 It is a schematic flowchart of a log management method provided by an embodiment of this application; Figure 5Schematic diagram of a log management system provided by an embodiment of the present application; Figure 6 Flowchart of a specific log management provided by an embodiment of the present application; Figure 7 Schematic structural diagram of a log management device provided by an embodiment of the present application. Detailed implementation manners

[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0015] With the continuous development of server technology, the scale of server clusters is becoming increasingly large. While this large-scale cluster architecture improves system performance and processing capabilities, it also brings many problems to the maintainability of the cluster. When a server fails, a fault log is generated and transmitted to the management server via the network or a specific interface. After the management server integrates the information, it is fed back to the operation and maintenance personnel for fault troubleshooting.

[0016] Currently, related technologies mainly focus on decentralized storage of logs and data verification. In terms of data transmission, related technologies still rely on the cluster management server and the central network node to build a data transmission channel.

[0017] It can be seen that the transmission path of related technologies is single, relying on the cluster management server and the central network node. In the face of a complex network environment, it is difficult to ensure the stability and reliability of log collection. Especially in some high-threat scenarios, the sending time window of server fault logs is extremely limited. Usually, there is only one chance to transmit fault information before the server crashes. Once a network or system anomaly occurs, even if the management server successfully captures the fault log, it cannot obtain the log information again through repeated handshakes or other means, resulting in the loss of key fault information and increasing the difficulty of operation and maintenance.

[0018] Next, several solutions of the log management method in related technologies will be briefly introduced: Solution A proposes a data storage method, including data identification and verification; the data identification and verification include data classification and data verification; data encrypted storage; the data encrypted storage includes data encryption and data storage; real-time behavior monitoring and anti-tampering mechanism. The combination of data encryption and decentralized storage effectively reduces the risks of data leakage and tampering; by reasonably combining symmetric encryption and asymmetric encryption technologies and introducing decentralized storage at the same time, the computational pressure in traditional public key encrypted storage solutions is alleviated.

[0019] Solution B proposes a method and system for recording inter-module call logs based on blockchain, including deploying a blockchain platform: building a decentralized blockchain database on a server; writing call logs to the chain: when there is an inter-module call, the call log data will be written to the chain to ensure the security and integrity of the data; anomaly location: when there are problems with inter-module data interaction, the abnormal module can be quickly located by comparing the log data of two organizations on the blockchain; log data mining. The system includes a blockchain platform deployment unit, a call log writing unit, an anomaly location unit, and a log data mining unit.

[0020] Solution C proposes a fault detection method and a server. It relates to the field of blockchain technology. It includes: obtaining a heartbeat message to be verified sent by a terminal to be verified from a blockchain network; obtaining a heartbeat log file sent by a heartbeat server from the blockchain network, where the heartbeat log file includes N preset heartbeat keep-alive messages, and N is an integer greater than or equal to 1; detecting whether the terminal to be verified is in a fault state based on the heartbeat message to be verified and the preset heartbeat keep-alive messages. By utilizing the decentralized feature of the blockchain network, the immutability of the heartbeat log file is ensured. Detecting whether the terminal to be verified is in a fault state based on the heartbeat message to be verified and the preset heartbeat keep-alive messages can obtain heartbeat information through different channels, monitor the security of the terminal to be verified in real time, improve the security of the terminal to be verified, and make the management behavior of the heartbeat server traceable through the heartbeat log file.

[0021] In the above solutions, there are the following defects: (1) Only considering the decentralized storage and data verification of logs, the data acquisition path still relies on the cluster management server and the central network node, and it is impossible to avoid log collection in a complex network environment.

[0022] (2) The decentralization in data storage and verification cannot avoid the influence of centralized management of the hardware entity link. The premise of the designed function lies in the normal operation of the server network environment, especially the central node. The path is single, and the overall maintainability of the cluster is weak.

[0023] (3) Although the overall log distribution and retrieval efficiency have been improved through the introduction of various algorithms, it cannot offset the exponential consumption of network resources compared to traditional log supervision. The lack of a reasonable hierarchical storage system has led to high decentralized storage and verification overheads, resulting in additional resource occupation.

[0024] Therefore, none of the above solutions can achieve the preservation of cluster logs in the event of network fluctuations or management machine anomalies, and cannot reproduce the failures in such situations, reducing the maintainability of the cluster.

[0025] To solve the problems existing in related solutions, an embodiment of the present application provides a log management method, including: when a first server fails, determining the failure level of the first server; determining the log transmission interface corresponding to the failure level; transmitting the failure logs in the first server to a target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the failure logs; when the failure of the first server is repaired, performing log recovery based on the failure logs stored in the target server, so that the first server stores the recovered failure logs, realizing hierarchical transmission of failure logs in a server cluster environment, effectively avoiding the risk of single-point failures, and ensuring the integrity and recoverability of failure logs even when there are fluctuations or interruptions in a complex network environment, thereby significantly improving the operation and maintenance reliability of the server cluster.

[0026] The log management method provided by an embodiment of the present disclosure can be applied to fields such as cloud computing, financial transactions, industrial Internet of Things, vehicle Internet of Things, and medical data management, which have strict requirements for high availability, strong consistency, and complex network disaster tolerance.

[0027] To enable those skilled in the art of this technology to better understand the solution of the present application, the following further elaborates on the present application in conjunction with the accompanying drawings and specific embodiments.

[0028] Figure 1 It is a schematic flowchart of a log management method provided by an embodiment of the present disclosure.

[0029] As Figure 1 shown, the method includes the following steps: Step 101, when a first server fails, determining the failure level of the first server.

[0030] In some embodiments, the present disclosure can utilize the monitoring functions built into the firmware or operating system of the server to automatically detect and identify the specific type of the current failure (such as hardware failure, software crash, network interruption, etc.). According to the pre-configured mapping relationship between the failure type and the failure level, the identified failure type is converted into the corresponding failure level.

[0031] Step 102, determining the log transmission interface corresponding to the failure level.

[0032] In some embodiments, the present disclosure may select an appropriate log transmission interface according to the fault level to ensure that critical logs can still be reliably transmitted to the target server in a fault scenario. The log transmission interfaces in the present disclosure include a first interface, a second interface, and a third interface. The first interface is specifically a BMC (Baseboard Management Controller) interface, the second interface is specifically an OS (Operating System) interface, and the third interface is specifically a hardware-level interface.

[0033] If the fault level is less than or equal to the first fault level, the log transmission interface is determined to be the first interface; if the fault level is greater than the first fault level and the fault is less than or equal to the second fault level, the log transmission interface is determined to be the first interface, the second interface, and the third interface; if the fault level is greater than the second fault level, the log transmission interface is determined to be the first interface.

[0034] The first fault level and the second fault level in the present disclosure can be preset according to experience and are not limited in the embodiments of the present disclosure.

[0035] Step 103: Transmit the fault logs in the first server to the target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the fault logs.

[0036] In some embodiments, when a fault occurs, the server firmware and the server operating system will record relevant alarm logs, and the relevant alarm logs are the fault logs in the present disclosure.

[0037] When a fault occurs, the server firmware and the operating system will generate alarm logs (i.e., fault logs) to record the detailed information of the fault. Through the interface determined in step 102, the logs are transmitted to the target server for storage for subsequent analysis and recovery, including transmitting the logs through the transmission channel of the BMC interface of the server (main channel (BMC / IPMI, bandwidth 10 Mbps)); transmitting the logs through the transmission channel of the OS interface (alternate channel (OS network stack TCP / IP, bandwidth 1 Gbps)); transmitting the logs through the transmission channel of the hardware-level interface (emergency channel (hardware-level NetConsole, bandwidth 1 Mbps)).

[0038] After receiving the fault logs, the target server stores them in a local storage device (such as a hard disk) for subsequent analysis.

[0039] In the present disclosure, the target server may include a management server and a second server. The present disclosure may directly transmit the fault log to the management server; or may transmit the fault log to the second server, and the second server transmits the fault log to the management server.

[0040] Step 104, when the first server fault repair is completed, perform log recovery based on the fault log stored in the target server, so that the first server stores the recovered fault log.

[0041] In some embodiments, after the first server fault repair is completed, it is necessary to perform log recovery through the fault log stored in the target server to ensure that the first server can completely retain the log records during the fault, which is crucial for subsequent fault analysis, auditing, and problem reproduction.

[0042] Specifically, when the first server detects that its own status has returned to "normal" or "repaired", it automatically triggers the log recovery process, or the administrator manually initiates the recovery operation.

[0043] The first server requests to obtain the stored fault log from the target server, that is, sends a data collection request to the target server to obtain the fault log sent by the target server based on the data collection request.

[0044] Among them, during the log recovery process, the first server can also perform integrity verification (such as MD5 verification) on the recovered fault log to ensure that the log has not been tampered with or lost.

[0045] In the embodiments of the present disclosure, since the target server includes a management server and a second server, there are some differences when performing log recovery.

[0046] If the fault log is transmitted to the management server through the transmission channel of the first interface, the fault log is only stored in the management server. When the first server performs log recovery, it can directly send a data collection request to the management server to obtain the fault log sent by the management server based on the data collection request.

[0047] If the fault log is transmitted to the target server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface, the fault log is stored in the management server and / or the second server. When the first server performs log recovery, it can directly send data collection requests to the management server and the second server to obtain the fault logs sent by the management server and the second server based on the data collection requests.

[0048] When the fault log is transmitted to the second server through the transmission channel corresponding to the first interface (i.e., the broadcast storage mechanism), the fault log is only stored in the second server. When performing log recovery on the first server, a data collection request can be sent to each second server; obtain the data blocks sent by each second server based on the data collection request, and perform hash verification based on the hash values stored in the blockchain and the hash values of each data block to obtain data blocks with consistent hash values; based on the erasure code technology, perform data recombination on the data blocks with consistent hash values, and use the recombined data as the fault log.

[0049] In addition, for the case where the server has not failed, the present disclosure can also perform real-time fault prediction on the server that has not failed. In order to more accurately predict server failures and take corresponding measures in a timely manner, the present disclosure can also introduce the MTBF / MTTR metric system. Combining the data collected in real time by hardware sensors, such as temperature, voltage, ECC error rate, etc., a fault prediction model is constructed. The fault prediction model can evaluate the health status of the server in real time and predict the probability of server downtime (i.e., the fault probability in the present disclosure). When it is predicted that the fault probability exceeds the preset probability threshold (for example, 70%), the present disclosure will immediately initiate firmware-level emergency transmission to ensure that as much critical log information as possible is transmitted before the server is about to go down, providing sufficient data support for subsequent fault troubleshooting. Specifically, it includes: when the first server has not failed, collect the sensor data of the first server; based on the sensor data, combined with the preset fault prediction model, predict the fault probability of the first server; if the fault probability of the first server is greater than or equal to the preset probability threshold, determine the log transmission interface as the first interface; transmit the fault log to the target server through the transmission channel corresponding to the first interface, that is, the broadcast storage mechanism in the present disclosure.

[0050] Among them, in non-emergency situations (i.e., the first server in the present disclosure does not fail), the cluster management server penetrates the traffic conditions of each network segment through the traffic monitoring system, and monitors the quality status of each channel in real time, including key indicators such as bandwidth, latency, and packet loss rate. Based on these real-time monitoring data, the present disclosure dynamically adjusts the transmission priority. For example, when the network quality of the main channel (BMC / IPMI) is good, this channel is preferentially used for log transmission; if the main channel fails or the network quality deteriorates, the system automatically switches the transmission task to the alternative channel (OS network stack); if the alternative channel cannot work properly either, the emergency channel (hardware-level NetConsole) is enabled to ensure that critical logs can be successfully transmitted. Specifically, it includes: when the first server does not fail, determining the network quality data of the transmission channel corresponding to each interface among the first interface, the second interface, and the third interface based on a preset channel quality scoring model; adjusting the transmission channel priority based on the network quality data of the transmission channel corresponding to each interface; and adjusting the current log transmission channel according to the adjusted transmission channel priority, so as to transmit the fault log to the target server according to the adjusted log transmission channel, and the adjusted log transmission channel is any one of the transmission channels corresponding to the first interface, the transmission channels corresponding to the second interface, and the transmission channels corresponding to the third interface. Among them, the channel quality scoring model can be pre-constructed by real-time monitoring of the packet loss rate (packet loss rate ≤ 5% is excellent), latency (latency ≤ 50 ms is excellent), and bandwidth utilization rate (bandwidth utilization rate ≤ 80% is excellent) of each channel. In the present disclosure, the switching delay of adjusting the log transmission channel needs to be controlled within 200 ms to ensure the continuity of log transmission.

[0051] In summary, through the present application, when the first server fails, determine the failure level of the first server; determine the log transmission interface corresponding to the failure level; transmit the fault log in the first server to the target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the fault log; when the failure of the first server is repaired, perform log recovery based on the fault log stored in the target server, so that the first server stores the recovered fault log, realizing hierarchical transmission of fault logs in a server cluster environment, effectively avoiding the risk of single-point failure, and ensuring the integrity and recoverability of fault logs even when fluctuations or interruptions occur in a complex network environment, thereby significantly improving the operation and maintenance reliability of the server cluster.

[0052] Figure 2 Further shows a flowchart of a log management method proposed by the present disclosure. Based on Figure 1 the shown embodiment, further explain step 103. The target server includes a management server, Figure 2 and may include the following steps.

[0053] Step 201, if the fault level is less than or equal to the first fault level, transmit the fault log to the management server through the transmission channel corresponding to the first interface.

[0054] In some embodiments, the first fault level indicates that the current fault is relatively minor.

[0055] At this time, when the fault level of the first server is less than or equal to the first fault level, the first server will, according to the established rules, transmit the fault log recording the detailed information of the fault to the management server through the transmission channel specifically associated with the first interface (i.e., the main channel corresponding to the BMC interface). The management server is usually used to centrally collect, store, and analyze the operation logs and fault information of various devices or systems, so that the operation and maintenance personnel can timely understand the system operation status and perform corresponding processing.

[0056] Step 202, if the confirmation information feedback by the management server is not received within the preset time period, transmit the fault log to the management server through the first interface until the confirmation information is received within the preset time period.

[0057] In the present disclosure, the confirmation information is feedback by the management server based on the transmission channel corresponding to the first interface.

[0058] In some embodiments, if the number of times of not receiving the confirmation information within the preset time period meets the preset number of times, determine that the log transmission interfaces are the first interface, the second interface, and the third interface; transmit the fault log to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface. In other words, after transmitting the fault log to the management server through the transmission channel corresponding to the first interface, the present disclosure will set a preset time period (such as 10 seconds, 30 seconds, etc.) and wait for the management server to feedback the confirmation information. The confirmation information is handshake confirmation information, which is a notification returned by the management server to the sender through the transmission channel corresponding to the first interface after successfully receiving the fault log.

[0059] If the confirmation information is not received within the preset time period, the present disclosure will consider that the fault log may not have been successfully transmitted to the management server, so the fault log will be sent again through the first interface. This process will be repeated until the confirmation information feedback by the management server is successfully received within the preset time period to ensure that the fault log can be reliably transmitted to the management server.

[0060] Further, if the number of times of not receiving the confirmation information within the preset time period meets the preset number of times, determine that the log transmission interface is the first interface; divide the fault log into multiple data blocks, and determine the second server for each data block; transmit each data block to the second server through the transmission channel corresponding to the first interface respectively.

[0061] When the number of times that the confirmation information feedback from the management server is not received within the preset time period reaches the preset number of times (such as 3 times, 5 times, etc.), the present disclosure will consider that there may be a greater risk of transmitting the fault log through a single first interface, which may lead to log loss. In order to improve the reliability of fault log transmission, the present disclosure will simultaneously select the first interface, the second interface, and the third interface as the log transmission interfaces. Then, the fault log is transmitted to the management server and the second server simultaneously through the transmission channels corresponding to these three interfaces. The second server refers to the server in the server cluster other than the first server. In this way, the present disclosure can ensure that the fault log can be received by multiple servers as much as possible, reducing the risk of log loss.

[0062] If the number of times that the confirmation information is still not received within the preset time period meets the preset number of times, the broadcast storage mechanism is started, that is, the log transmission interface is determined to be the first interface; the fault log is divided into multiple data blocks, and the second server for each data block is determined; each data block is transmitted to the second server through the transmission channel corresponding to the first interface respectively, so as to further improve the flexibility and reliability of fault log transmission, and is also conducive to more detailed storage and management of the fault log.

[0063] In addition, the present disclosure can monitor the network delay situation in real time when transmitting the fault log through the transmission channel corresponding to the log transmission interface. Network delay refers to the time required for data to be transmitted from the sending end to the receiving end, that is, the network round-trip time RTT. When it is monitored that the continuous duration of the network delay is greater than or equal to the preset delay duration (such as 50 milliseconds, etc.), it indicates that there may be a problem with the current network condition, which may lead to the failure or excessive delay of fault log transmission. In order to cope with this situation, the present disclosure will start the broadcast storage mechanism and select the first interface as the log transmission interface to transmit the fault log to the second server through the first interface. Specifically, the present disclosure can detect the network delay of transmitting the fault log through the transmission channel corresponding to the log transmission interface; when the continuous duration of the network delay is greater than or equal to the preset delay duration, the log transmission interface is determined to be the first interface; the fault log is divided into multiple data blocks, and the second server for each data block is determined; each data block is transmitted to the second server through the transmission channel corresponding to the first interface respectively.

[0064] In summary, the present disclosure can select the corresponding transmission result based on the current fault level for a lower fault level, and flexibly adjust the log transmission interface and transmission strategy according to the network condition and transmission failure situation, ensuring that the fault log can be reliably and timely transmitted to the target server, providing strong support for the operation and maintenance and management of the system.

[0065] Figure 3 Further shows a flowchart of a log management method proposed by the present disclosure. Based onFigure 1 For the embodiments shown, step 103 is further explained. The target server includes a management server and a second server. Figure 3 The following steps may be included.

[0066] Step 301: If the fault level is greater than the first fault level and less than or equal to the second fault level, then transmit the fault log to the management server through the transmission channels corresponding to the first interface and the second interface, and transmit the fault log to the second server through the transmission channel corresponding to the third interface. The second fault level is greater than the first fault level.

[0067] In some embodiments, the first fault level represents a minor fault, and the second fault level represents a relatively serious fault (such as a sudden spike in the server CPU temperature). In the present disclosure, the fault level that is greater than the first fault level and less than or equal to the second fault level is regarded as the medium fault level, representing a medium fault.

[0068] In the present disclosure, the fault log can be transmitted to the management server through the transmission channels corresponding to the first interface and the second interface simultaneously. Using multiple interfaces for transmission can improve the reliability of data transmission and avoid the situation where the fault log cannot be transmitted to the management server due to a single interface failure. At the same time, transmit the fault log to the second server through the transmission channel corresponding to the third interface to ensure that the fault information can be obtained by multiple servers for more comprehensive fault handling and analysis.

[0069] Step 302: If the confirmation information feedback from the management server is not received within the preset time period, then transmit the fault log to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface until the confirmation information is received within the preset time period.

[0070] In the present disclosure, the confirmation information is feedback by the management server through any one of the transmission channels corresponding to the first interface and the second interface and / or by the second server through the transmission channel corresponding to the third interface.

[0071] In some embodiments, if the confirmation information feedback from the management server is not received within the preset time period, the present disclosure can adopt a repeated transmission strategy to transmit the fault log to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface simultaneously until the confirmation information is received within the preset time period. This repeated transmission strategy increases the probability of successfully transmitting the fault log to the target server and ensures that the fault information will not be lost due to transmission problems. The preset time period is to give the receiving party enough time to process the fault log and feedback the confirmation information. If the confirmation information is not received after exceeding this time period, it is considered that there may be a transmission problem.

[0072] Furthermore, after multiple attempts, there may be relatively serious transmission problems, and more stringent processing of the transmission method of the fault log is required, that is, a broadcast storage mechanism is adopted. Specifically, if the number of times of not receiving the confirmation information within the preset time period meets the preset number, it is determined that the log transmission interface is the first interface; the fault log is divided into multiple data blocks, and the second server for each data block is determined; each data block is respectively transmitted to the second server through the transmission channel corresponding to the first interface.

[0073] Data block division can improve the flexibility and reliability of fault log transmission. If a certain data block fails to be transmitted, only that data block needs to be retransmitted instead of the entire fault log. Determining the second server for each data block can be reasonably allocated according to the content of the data block and the processing capacity of the second server. Each data block is respectively transmitted to the second server through the transmission channel corresponding to the first interface. This transmission method further increases the reliability of data transmission, ensuring that each data block can be successfully transmitted to the second server. It should be noted that in the present disclosure, each data block can be repeatedly transmitted to a preset number of second servers.

[0074] In addition, in order to still ensure the transmission of the fault log in the case of a large network delay and avoid the problem that the fault information cannot be transmitted in time due to the network delay, the present disclosure can detect the network delay of transmitting the fault log through the transmission channel corresponding to the log transmission interface; when the duration of the network delay is greater than or equal to the preset delay duration, the broadcast storage mechanism is adopted, and it is determined that the log transmission interface is the first interface; the fault log is divided into multiple data blocks, and the second server for each data block is determined; each data block is respectively transmitted to the second server through the transmission channel corresponding to the first interface.

[0075] In summary, the present disclosure flexibly adjusts the transmission method and transmission channel of the fault log according to factors such as the fault level, confirmation information reception, preset number, and network delay, ensuring that the fault log can be reliably transmitted to the management server and the second server, so that the operation and maintenance personnel can timely understand the system operation status and perform corresponding processing. It comprehensively considers the severity of the fault, the reliability and timeliness of transmission, and improves the stability and maintainability of the system.

[0076] Figure 4 Further, a flowchart of a log management method proposed by the present disclosure is shown. Based on Figure 1 the embodiments shown, step 103 is further explained. The target server includes the second server, Figure 4 and it can include the following steps.

[0077] Step 401: If the fault level is greater than the second fault level, divide the fault log into multiple data blocks and determine the second server for each data block.

[0078] In the present disclosure, the second fault level is greater than the first fault level.

[0079] Step 402: Transmit each data block to the second server respectively through the transmission channel corresponding to the first interface.

[0080] In some embodiments, the present disclosure regards faults greater than the second fault level as high-risk level faults. At this time, the present disclosure uses BMC to send relevant logs to all receivable devices in the network domain in the form of network segment broadcasting. Considering the situation that the server network segment and the BMC network segment in the server cluster are separated, BMC will also use broadcast transmission in the network segment where the operating system is located, that is, the broadcast storage mechanism in the present disclosure.

[0081] The broadcast storage mechanism of the present disclosure is to use a blockchain-like storage structure to store key logs. The specific method is to divide the fault log into multiple data blocks (i.e., N data blocks) according to a preset size, and use the cross-rack storage topology awareness algorithm to automatically calculate the topology relationship between different server nodes and management machine nodes in the cluster (i.e., the preset cluster node relationship in the present disclosure). Based on the network latency and remaining storage space between server nodes in the preset cluster node relationship, determine the preset number of second servers corresponding to each data block, and send each data block to different preset numbers (for example, at least 3) of second servers. Among them, to further improve the disaster tolerance effect, the preset number of second servers corresponding to each data block in the present disclosure needs to ensure that these second servers are not in the same computer room or the same row.

[0082] Among them, the second server in the present disclosure refers to a physical node, and the physical node can be either a management server, a second server, or the BMC storage space in the transmission channel corresponding to the first interface. This distributed storage method greatly improves the reliability and availability of log data. Even if individual nodes fail, it will not affect the integrity of log data.

[0083] When the present disclosure transmits each data block to the corresponding preset number of second servers, the present disclosure collaboratively ensures the recoverability and credibility of data through the immutability of the blockchain and the redundancy of distributed storage, and can also store the hash value corresponding to each data block in the blockchains corresponding to multiple data blocks at the same time.

[0084] Specifically, the log file is first split into fixed-size data blocks (such as 256KB). The IPFS sharding storage mechanism is used to distribute each data block to multiple second servers, and a unique hash value (such as SHA-3) is generated and written into the blockchain as evidence. To establish the logical association between blocks, the present disclosure can also construct a chained hash pointer structure. For example, the hash values of adjacent data blocks are embedded in the data header, or the root hash is generated by aggregating multiple data blocks through a Merkle tree and uploaded to the chain to ensure that the log order is irreversible and any tampering will cause the hash chain to break.

[0085] To accelerate the location of the dispersedly stored data blocks, the present disclosure can also maintain a key-value index table (such as LevelDB) outside the chain, recording the storage node addresses, timestamps, and redundant copy locations of each data block. The overall hash value of this index table is calculated by a smart contract every fixed data block period (such as every 100 data blocks) and written into the blockchain, forming a double anti-tampering protection: attackers need to modify the content of the index table and its corresponding on-chain hash simultaneously to destroy the data addressability, and the distributed consensus mechanism makes such attacks almost impossible to implement in a decentralized network.

[0086] In addition, when determining the second server for each data block, the present disclosure also needs to consider whether the second server fails. If the second server of the first data block fails, then based on the election protocol, a third server is determined from multiple alternative servers connected to the first server, and the number of node votes of the third server within the preset election period is greater than or equal to the preset vote count.

[0087] Specifically, to ensure that the present disclosure can still work properly in the case of some second servers failing, the present disclosure uses the election protocol (Raft protocol) to implement dynamic storage server election. By combining the preset cluster node relationship, the settings of the leader master node, Follower slave node, and Candidate candidate node are realized, and a new Leader node (i.e., the third server) is dynamically elected to achieve decentralized storage of logs in different scenarios, ensuring that even if 50% of the nodes fail, the present disclosure can still maintain the integrity of the logs. The election process requires the candidate node to obtain more than 50% of the node votes within a random timeout of 150 - 300ms.

[0088] In summary, for high-risk fault levels, a log storage broadcast mechanism can be based on the current fault level to ensure that the fault logs can be reliably and timely transmitted to the target server, providing strong support for the operation and maintenance of the system.

[0089] Based on the above Figures 1 to 4 illustrated embodiments, as Figure 5 shown, the present disclosure provides a schematic diagram of a log management system.

[0090] In an embodiment of the present disclosure, the log management system of the present disclosure includes a first server (i.e., the failed server in Figure 5 ), and a target server. The target server includes a management server and a second server (i.e., the other servers in Figure 5 ).

[0091] Both the first server and the second server include a first interface (i.e., the BMC interface in Figure 5 ), a second interface (i.e., the OS interface in Figure 5 ), and a third interface (i.e., the hardware-level interface in Figure 5 ). The management server includes a first interface and a second interface.

[0092] When the failure level of the first server is less than or equal to the first failure level, the failure log is transmitted to the management server through the transmission channel of the first interface (i.e., the network channel of the main channel corresponding to the BMC interface in Figure 5 ) in combination with a switch.

[0093] When the failure level of the first server is greater than the first failure level and less than or equal to the second failure level, the failure log is transmitted to the management server through the transmission channels corresponding to the first interface and the second (i.e., the network channel of the main channel corresponding to the BMC interface and the network channel of the alternative channel corresponding to the OS interface in Figure 5 ) in combination with a switch, and at the same time, the failure log is transmitted to the second server through the transmission channel corresponding to the third interface (i.e., the hardware-level connection of the emergency channel corresponding to the hardware-level interface in Figure 5 ).

[0094] When the failure level of the first server is greater than the second failure level, a log broadcast storage mechanism is adopted, that is, the failure log is divided into multiple data blocks, and the second server corresponding to each data block is selected, and each data block is transmitted to the corresponding second server through the transmission channel corresponding to the first interface (i.e., the network channel of the main channel corresponding to the BMC interface in Figure 5 ) in combination with a switch.

[0095] For the specific log management method of the log management system, reference can be made to the embodiment shown in Figures 1 to 4 , which will not be elaborated here.

[0096] Based on the embodiment shown in Figures 1 to 5 , as shown in Figure 6 , the present disclosure provides a flowchart of a specific log management.

[0097] In an embodiment of the present disclosure, it is determined whether there is a server failure in the server cluster. If the first server fails, the failure level of the first server is determined; if the first server does not fail, the log is transmitted to the management server through the transmission channel corresponding to the first interface (i.e.,Figure 6 (The general logs are transmitted to the management server through the main channel). Determine whether the failure level of the first server is greater than the first failure level (i.e., Figure 6 whether the failure level is greater than P4 in the figure). If the failure level is less than or equal to the first failure level, the failure logs are transmitted to the management server through the transmission channel corresponding to the first interface (i.e., Figure 6 the failure logs are transmitted to the management server through the main channel in the figure). If the failure level is greater than the first failure level, further determine whether the failure level is greater than the second failure level (i.e., Figure 6 whether the failure level is greater than P2 in the figure). If the failure level is greater than the first failure level and less than or equal to the second failure level, the failure logs are transmitted to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface (i.e., Figure 6 the logs are transmitted through the main channel, the backup channel, and the emergency channel in the figure). If the failure level is greater than the second failure level, the failure logs are divided into multiple data blocks through the broadcast storage mechanism, and the second server for each data block is determined, and each data block is transmitted to the second server through the transmission channel corresponding to the first interface respectively (i.e., Figure 6 the failure logs are transmitted to the cluster through the broadcast mechanism in the figure).

[0098] Among them, when the failure logs are transmitted to the management server through the transmission channel corresponding to the first interface, and when the failure logs are transmitted to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface, it is possible to further determine whether the management server and / or the second server normally feedback the confirmation information (i.e., Figure 6 judging whether the return value of the management machine correctly returns the correct handshake information and judging whether the return value is correct to return the correct handshake information in the figure). If the confirmation information is not normally feedback, a higher-level log transmission method is further adopted until the broadcast storage mechanism is adopted.

[0099] For Figure 6In the illustrated embodiment, for the sake of easy understanding, the present disclosure takes the failure of a sudden CPU temperature spike in a server as an example. The specific log management method is as follows: The CPU temperature of the first server suddenly spikes. The log management system determines that the failure level of the first server is a serious failure (P2 level). At this time, the failure level is the same as the second failure level, and the failure level is within the range less than or equal to the second failure level (P2 level). First, the failure log is sent through the transmission channels corresponding to the first interface, the second interface, and the third interface. However, no response is obtained due to network congestion. After a preset time period (5 seconds), it automatically switches to the transmission channel corresponding to the first interface and successfully splits the log into 3 data blocks. According to the intelligent location selection rule, the 3 data blocks are stored in the servers in Cabinet No. 1 of Machine Room A, Cabinet No. 3 of Machine Room B, and Cabinet No. 5 of Machine Room C respectively. A power outage occurred in Machine Room B that afternoon, resulting in data loss of two servers. The log management system can automatically collect data fragments from Machine Rooms A and C and the other two normal servers. The complete log is successfully restored through a mathematical algorithm, and the authenticity verification of 10 GB of data is completed.

[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware. However, in many cases, the former is a better implementation method.

[0101] An embodiment of the present application also provides a log management device 700. Figure 7 As a schematic structural diagram of a log management device provided by an embodiment of the present disclosure, as Figure 7 shown, it includes: A level determination unit 710, configured to determine the failure level of the first server when a failure occurs in the first server; An interface determination unit 720, configured to determine the log transmission interface corresponding to the failure level; A transmission unit 730, configured to transmit the failure log in the first server to a target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the failure log; A recovery unit 740, configured to perform log recovery based on the failure log stored in the target server when the failure of the first server is repaired, so that the first server stores the recovered failure log.

[0102] The log management device of the present disclosure, when a first server fails, determines the failure level of the first server; determines the log transmission interface corresponding to the failure level; transmits the failure logs in the first server to a target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the failure logs; when the failure of the first server is repaired, performs log recovery based on the failure logs stored in the target server, so that the first server stores the recovered failure logs, realizing hierarchical transmission of failure logs in a server cluster environment, effectively avoiding the risk of single-point failure, and ensuring the integrity and recoverability of failure logs even when there are fluctuations or interruptions in a complex network environment, thereby significantly improving the operation and maintenance reliability of the server cluster.

[0103] Further, in a possible implementation manner of the embodiment of the present disclosure, the interface determination unit 720 is configured to, if the failure level is less than or equal to a first failure level, determine the log transmission interface as a first interface; if the failure level is greater than the first failure level, determine the log transmission interface as the first interface, a second interface, and a third interface; if the failure level is greater than a second failure level, determine the log transmission interface as the first interface.

[0104] Further, in a possible implementation manner of the embodiment of the present disclosure, the target server includes a management server, and the transmission unit 730 is configured to, if the failure level is less than or equal to the first failure level, transmit the failure logs to the management server through the transmission channel corresponding to the first interface; if the confirmation information fed back by the management server is not received within a preset time period, transmit the failure logs to the management server through the first interface until the confirmation information is received within the preset time period, and the confirmation information is fed back by the management server based on the transmission channel corresponding to the first interface.

[0105] Further, in a possible implementation manner of the embodiment of the present disclosure, the target server includes a management server and a second server, and the transmission unit 730 is configured to, if the failure level is greater than the first failure level and less than or equal to a second failure level, transmit the failure logs to the management server through the transmission channel corresponding to the first interface and the transmission channel corresponding to the second interface, and transmit the failure logs to the second server through the transmission channel corresponding to the third interface, where the second failure level is greater than the first failure level; if the confirmation information fed back by the management server is not received within a preset time period, transmit the failure logs to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface until the confirmation information is received within the preset time period, and the confirmation information is fed back by the management server through any one of the transmission channels corresponding to the first interface and the second interface and / or by the second server through the transmission channel corresponding to the third interface.

[0106] Further, in a possible implementation manner of the embodiment of the present disclosure, the target server includes a second server, and the transmission unit 730 is configured to divide the fault log into multiple data blocks if the fault level is greater than the second fault level, and determine the second server of each data block, where the second fault level is greater than the first fault level; and transmit each data block to the second server through the transmission channel corresponding to the first interface.

[0107] Further, in a possible implementation manner of the embodiment of the present disclosure, the transmission unit 730 is further configured to, after transmitting the fault log to the target server through the first interface, if the number of times of not receiving the confirmation information within the preset time period meets the preset number of times, determine that the log transmission interface is the first interface, the second interface, and the third interface; and transmit the fault log to the management server and the second server through the transmission channels corresponding to the first interface, the second interface, and the third interface.

[0108] Further, in a possible implementation manner of the embodiment of the present disclosure, the transmission unit 730 is further configured to, after transmitting the fault log to the target server through the transmission channels corresponding to the first interface, the second interface, and the third interface, if the number of times of not receiving the confirmation information within the preset time period meets the preset number of times, determine that the log transmission interface is the first interface; divide the fault log into multiple data blocks; and determine the second server of each data block; and transmit each data block to the second server through the transmission channel corresponding to the first interface.

[0109] Further, in a possible implementation manner of the embodiment of the present disclosure, the apparatus further includes: a fault prediction unit, configured to collect sensor data of the first server when the first server does not fail; based on the sensor data and in combination with a preset fault prediction model, predict the fault probability of the first server; if the fault probability of the first server is greater than a preset probability threshold, determine that the log transmission interface is the first interface; and transmit the fault log to the target server through the transmission channel corresponding to the first interface.

[0110] Further, in a possible implementation manner of the embodiment of the present disclosure, the transmission unit 730 is configured to divide the fault log into multiple data blocks according to a preset size; based on the network latency and remaining storage space between server nodes in the preset cluster node relationship, determine a preset number of second servers corresponding to each data block, and store the hash value corresponding to each data block in the blockchain corresponding to the multiple data blocks.

[0111] Further, in a possible implementation manner of the embodiments of the present disclosure, the transmission unit 730 is configured to detect the network latency of transmitting the fault log through the transmission channel corresponding to the log transmission interface; when the duration of the network latency is greater than a preset latency duration, determine that the log transmission interface is the first interface; divide the fault log into multiple data blocks, and determine the second server for each data block; and transmit each data block to the second server through the transmission channel corresponding to the first interface respectively.

[0112] Further, in a possible implementation manner of the embodiments of the present disclosure, the transmission unit 730 is configured to, if the second server of the first data block fails, determine a third server from multiple alternative servers connected to the first server based on an election protocol, where the number of node votes of the third server within a preset election period is greater than a preset number of votes.

[0113] Further, in a possible implementation manner of the embodiments of the present disclosure, the recovery unit 740 is configured to send a data collection request to each second server; obtain the data blocks sent by each second server based on the data collection request, and perform hash verification based on the hash value stored in the blockchain and the hash value of each data block to obtain data blocks with consistent hash values; and perform data recombination on the data blocks with consistent hash values based on erasure code technology, and use the recombined data as the fault log.

[0114] For the descriptions of the features in the embodiments corresponding to the log management device, reference may be made to the relevant descriptions in the embodiments corresponding to the log management method, which will not be elaborated here one by one.

[0115] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the embodiments of the above log management method.

[0116] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the embodiments of the above log management method when running.

[0117] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (abbreviated as ROM), random access memory (abbreviated as RAM), mobile hard disk, magnetic disk, or optical disc, and other media that can store computer programs.

[0118] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the log management method are implemented.

[0119] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the log management method are implemented.

[0120] Those skilled in the art can further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0121] The above has introduced in detail a log management method, an electronic device, a storage medium, and a product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A log management method, characterized in that: The method comprises: When a first server fails, determining a failure level of the first server; Determine the log transmission interface corresponding to the fault level; Transmitting the fault log in the first server to the target server through the transmission channel corresponding to the log transmission interface, so that the target server stores the fault log; When the fault repair of the first server is completed, log recovery is performed based on the fault log stored in the target server, so that the first server stores the restored fault log.

2. The method according to claim 1, characterized in that: The log transmission interface corresponding to the fault level is determined as follows: If the fault level is less than or equal to the first fault level, determining the log transmission interface to be the first interface; If the fault level is greater than the first fault level, and the fault level is less than or equal to the second fault level, determining the log transmission interface to be the first interface, the second interface, and the third interface; If the fault level is greater than the second fault level, the log transmission interface is determined to be the first interface.

3. The method according to claim 2, characterized in that The target server includes a management server, and transmitting the fault log in the first server to the target server through the transmission channel corresponding to the log transmission interface includes: If the fault level is less than or equal to the first fault level, transmitting the fault log to the management server through the transmission channel corresponding to the first interface; If the confirmation information fed back by the management server is not received within the preset time period, the fault log is transmitted to the management server through the first interface until the confirmation information is received within the preset time period. The confirmation information is fed back by the management server based on the transmission channel corresponding to the first interface.

4. The method according to claim 2, characterized in that: The target server includes a management server and a second server, and transmitting the fault log in the first server to the target server through the transmission channel corresponding to the log transmission interface includes: If the fault level is greater than the first fault level and the fault level is less than or equal to the second fault level, the fault log is transmitted to the management server through the transmission channel corresponding to the first interface and the transmission channel corresponding to the second interface. and transmitting the fault log to the second server through a transmission channel corresponding to the third interface, the second fault level being greater than the first fault level; If the confirmation information fed back by the management server is not received within the preset time period, the fault log is transmitted to the management server and the second server through the transmission channels corresponding to the first interface, the second interface and the third interface until the confirmation information is received within the preset time period. The confirmation information is fed back by the management server through any one of the transmission channel corresponding to the first interface and the transmission channel corresponding to the second interface and / or the second server through the transmission channel corresponding to the third interface.

5. The method according to claim 2, characterized in that: The target server includes a second server, and transmitting the fault log in the first server to the target server through the transmission channel corresponding to the log transmission interface includes: If the fault level is greater than the second fault level, dividing the fault log into a plurality of data blocks, and determining a second server for each data block, the second fault level being greater than the first fault level; Each of the data blocks is transmitted to the second server through the transmission channel corresponding to the first interface.

6. The method according to claim 3, characterized in that After transmitting the fault log to the target server through the first interface, the method includes: If the number of times that the confirmation information is not received within the preset time period meets the preset number of times, determining that the log transmission interface is the first interface, the second interface and the third interface; The fault log is transmitted to the management server and the second server through the transmission channels corresponding to the first interface, the second interface and the third interface.

7. The method according to any one of claims 4 or 6, characterized in that After transmitting the fault log to the target server through the transmission channels corresponding to the first interface, the second interface, and the third interface, the method includes: If the number of times that the confirmation information is not received within the preset time period meets the preset number of times, determining that the log transmission interface is the first interface; Dividing the fault log into a plurality of data blocks, and determining a second server for each data block; Each of the data blocks is transmitted to the second server through the transmission channel corresponding to the first interface.

8. The method according to claim 1, characterized in that: The method further comprises: When the first server is not faulty, collecting sensor data of the first server; Based on the sensor data and in combination with a preset fault prediction model, predict the failure probability of the first server; If the failure probability of the first server is greater than or equal to a preset probability threshold, determining the log transmission interface to be the first interface; The fault log is transmitted to the target server through the transmission channel corresponding to the first interface.

9. The method according to claim 5, characterized in that The step of dividing the fault log into a plurality of data blocks and determining a second server for each data block comprises: Dividing the fault log into multiple data blocks according to a preset size; Based on the network delay and remaining storage space between the server nodes in the preset cluster node relationship, a preset number of second servers corresponding to each data block is determined, and the hash value corresponding to each data block is stored in the blockchain corresponding to the multiple data blocks.

10. The method according to any one of claims 3 or 4, characterized in that: The transmitting the fault log in the first server to the target server through the transmission channel corresponding to the log transmission interface comprises: Detecting a network delay in transmitting the fault log through a transmission channel corresponding to the log transmission interface; When the duration of the network delay is greater than or equal to the preset delay duration, determining that the log transmission interface is the first interface; Dividing the fault log into a plurality of data blocks, and determining a second server for each data block; Each of the data blocks is transmitted to the second server through the transmission channel corresponding to the first interface.

11. The method according to claim 5, characterized in that The determining of the second server for each data block comprises: If the second server of the first data block fails, a third server is determined from multiple candidate servers connected to the first server based on the election protocol, and the number of node votes of the third server within the preset election period is greater than or equal to the preset number of votes.

12. The method according to claim 9, characterized in that The log recovery based on the fault log stored in the target server includes: sending a data collection request to each second server; Obtaining the data blocks sent by each second server based on the data collection request, and performing hash verification based on the hash value stored in the blockchain and the hash value of each data block to obtain data blocks with consistent hash values; Based on the erasure code technology, data blocks with consistent hash values ​​are reorganized, and the reorganized data is used as the fault log.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the log management method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the log management method according to any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the log management method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Log information processing method and system

    CN109492045A

  • Multi-channel communication method and system

    CN116545922A