Hardware fault monitoring method and related device

By using software indicator data and preset abnormal conditions in the hardware fault monitoring system to trigger the target log collection action, and timely obtain the hardware log data, the problem of fault monitoring delay in the existing technology is solved, and faster fault detection and maintenance is achieved.

CN119938433APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311460906.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, hardware failure monitoring delay is high, which makes it difficult to detect hardware failures in a timely manner, affects the timeliness of fault repairs, and may even lead to system downtime or business damage.

Method used

By obtaining software indicator data and preset abnormal conditions, determine whether there is a target log acquisition action that needs to be triggered, control the hardware log acquisition device to perform the target log acquisition action, obtain the log data of the target hardware object, and determine whether there is a fault based on the log data.

Benefits of technology

It realizes more timely detection of hardware failures, reduces hardware failure monitoring delays, improves fault repair time, and avoids the risk of system downtime and business damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938433A_ABST
    Figure CN119938433A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a hardware fault monitoring method and related device.The method comprises the steps that whether a target log collection action needing to be triggered exists or not is determined according to obtained software index data and preset abnormal conditions, and the target log collection action is used for collecting log data of a target hardware object; the target hardware object is a hardware object capable of enabling the software index data to meet a preset abnormal condition, controlling the hardware log acquisition device to execute a target log acquisition action under the condition of determining that the target log acquisition action needing to be triggered exists, and obtaining log data of the target hardware object, and determining whether the target hardware object has a fault according to the log data of the target hardware object. Wherein the software index data is data which needs to be acquired by related services, so that no additional io load is generated, the existing hardware fault can be detected in time, the hardware fault monitoring delay is reduced, and the fault maintenance timeliness is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a hardware fault monitoring method and related devices. Background Art

[0002] In actual applications, hardware fault monitoring is used to monitor whether there are faults in the operating system and related hardware, and when a fault is detected, a fault ticket is created and submitted to the hardware fault repair system so that relevant maintenance personnel can repair the existing hardware fault according to the fault ticket.

[0003] In the related art, hardware log collection is usually triggered periodically, and the collected hardware logs are reported to the hardware monitoring system, which detects whether there is a hardware fault based on the hardware logs. However, collecting hardware logs will cause the operating system to generate a certain input / output (IO) load. In order to avoid the generated IO load affecting the normal operation of the business, the collection period of the hardware log is usually set to be longer, which will lead to a higher delay in fault monitoring, making it difficult for hardware faults to be discovered in time, thereby affecting the timeliness of fault repair, and in severe cases may cause system downtime and damage to related businesses. Summary of the invention

[0004] The embodiments of the present application provide a hardware fault monitoring method and related devices, which solve the problem of high fault detection delay and difficulty in timely detection of hardware faults.

[0005] The first aspect of the present application provides a hardware fault monitoring method, the method comprising:

[0006] Get software indicator data;

[0007] According to the software indicator data and the preset abnormal conditions, determine whether there is a target log collection action that needs to be triggered; the target log collection action is used to collect log data of the target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions;

[0008] When it is determined that there is a target log collection action that needs to be triggered, controlling the hardware log collection device to execute the target log collection action to obtain log data of the target hardware object;

[0009] Determine whether the target hardware object has a fault based on the log data of the target hardware object.

[0010] A second aspect of the present application provides a hardware fault monitoring device, the device comprising:

[0011] A software data acquisition module is used to acquire software indicator data;

[0012] The log collection trigger module is used to determine whether there is a target log collection action that needs to be triggered according to the software indicator data and the preset abnormal conditions; the target log collection action is used to collect the log data of the target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions;

[0013] The log collection module is used to control the hardware log collection device to execute the target log collection action and obtain the log data of the target hardware object when it is determined that there is a target log collection action to be triggered;

[0014] The fault detection module is used to determine whether the target hardware object has a fault according to the log data of the target hardware object.

[0015] Optionally, the software indicator data includes access data of at least one data interface; the log collection trigger module includes:

[0016] An acquisition unit, for acquiring, for each data interface, a hardware monitoring trigger rule corresponding to the data interface; the hardware monitoring trigger rule is used to indicate an abnormal threshold of access data corresponding to the data interface and a reference log collection action, and the reference log collection action is used to collect log data of hardware objects that affect the access of the data interface;

[0017] The log collection trigger unit is used to determine, for each data interface, whether the access data of the data interface meets the preset abnormal conditions based on the access data of the data interface and the access data abnormality threshold indicated by the hardware monitoring trigger rule corresponding to the data interface; if so, determine the reference log collection action indicated by the hardware monitoring trigger rule corresponding to the data interface as the target log collection action.

[0018] Optionally, the access data of the data interface includes the number of accesses, the number of erroneous accesses, and the access delay of the data interface in the current cycle; and the log collection trigger unit includes:

[0019] A calculation subunit is used to calculate the access failure rate corresponding to the statistical interface according to the number of accesses and the number of erroneous accesses in the access data; and to calculate the average access delay and the maximum access delay corresponding to the statistical interface according to the access delay in the access data;

[0020] The log collection trigger subunit is used to determine whether the access data meets the preset abnormal condition based on at least one of the relationship between the access failure rate and the failure rate threshold in the access data abnormal threshold, the relationship between the average access delay and the average delay threshold in the access data abnormal threshold, and the relationship between the maximum access delay and the maximum delay threshold in the access data abnormal threshold.

[0021] Optionally, the device further comprises:

[0022] An acquisition module, used to acquire the latest change time corresponding to the acquired hardware monitoring trigger rule from a rule library for storing the hardware monitoring trigger rule corresponding to the data interface;

[0023] The update module is used to determine whether the acquired hardware monitoring trigger rule has been updated based on the latest change time; if it is determined that the acquired target hardware monitoring trigger rule has been updated, the updated target hardware monitoring trigger rule is obtained from the rule library.

[0024] Optionally, the log collection module includes:

[0025] A target control instruction generation and sending unit, used to generate a target control instruction according to a target log collection action, and send the target control instruction to a hardware log collection device;

[0026] The receiving unit is used to receive the log data of the target hardware object collected by the hardware log collection device in response to the target control instruction.

[0027] Optionally, a log collection command set is stored in the hardware log collection device, wherein the log collection command set includes a plurality of log collection commands, and the log collection commands are used to collect log data of the corresponding hardware objects;

[0028] The hardware log collection device collects log data of the target hardware object in the following ways:

[0029] A parsing unit, used to parse the target control instruction and obtain the target log collection action;

[0030] The collection unit is used to search for a target log collection command corresponding to a target log collection action in a log collection command set, execute the target log collection command, and obtain log data of a target hardware object.

[0031] Optionally, the device further comprises:

[0032] The periodic acquisition module is used to acquire periodic hardware log data; the periodic hardware log data is collected by the hardware acquisition device performing periodic log collection actions, and the periodic hardware log data includes the log data of each reference type of hardware object;

[0033] The periodic fault detection module is used to determine whether a hardware object of each reference type has a fault based on periodic hardware log data.

[0034] Optionally, a log collection command set is stored in the hardware log collection device, wherein the log collection command set includes a plurality of log collection commands, the log collection commands are used to collect log data of the corresponding hardware objects, and the log collection commands have corresponding execution cycles;

[0035] The hardware log collection device collects periodic hardware log data in the following ways:

[0036] The periodic collection unit is used to periodically execute the log collection command according to the execution period corresponding to each log collection command, and obtain the periodic hardware log data of the hardware object corresponding to the log collection command.

[0037] A third aspect of the present application provides a computer device, the device comprising a processor and a memory:

[0038] The memory is used to store computer programs;

[0039] The processor is used to execute the steps of the hardware fault monitoring method as described in the first aspect according to the computer program.

[0040] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the steps of the hardware fault monitoring method described in the first aspect above.

[0041] In a fifth aspect, the present application provides a computer program product or a computer program, the computer program product or the computer program includes a computer instruction, the computer instruction is stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device performs the steps of the hardware fault monitoring method described in the first aspect above.

[0042] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0043] The hardware fault monitoring method provided by the embodiment of the present application determines whether there is a target log collection action that needs to be triggered by the acquired software indicator data and the preset abnormal conditions, and the target log collection action is used to collect the log data of the target hardware object. The target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions. When it is determined that there is a target log collection action that needs to be triggered, the hardware log collection device is controlled to perform the target log collection action to obtain the log data of the target hardware object, and the target hardware object is determined to have a fault according to the log data of the target hardware object. Among them, different from the periodic hardware log collection in the related art, the embodiment of the present application determines whether there is a target hardware object that may have a fault by detecting whether the software indicator data is abnormal. When it is determined that there is a target hardware object, the hardware log collection device is controlled to collect the log data of the target hardware object, so as to further confirm whether the target hardware object has actually failed. Since software indicator data is data that needs to be collected by the relevant business itself, the present application will not generate additional IO load due to the collection of software indicator data. On this basis, the present application can obtain the collected software indicator data based on a shorter cycle or in real time, and detect whether there is a target hardware object that may have a fault based on this more real-time software indicator data, and trigger the collection of log data of the target hardware object accordingly to further determine whether the target hardware object has a fault. Existing hardware faults can be detected more promptly, reducing hardware fault monitoring delays and further ensuring the timeliness of fault repair. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of a hardware monitoring system provided for related technologies;

[0045] Figure 2 A schematic diagram of a scenario of a hardware fault monitoring method provided in an embodiment of the present application;

[0046] Figure 3 A flowchart of a hardware fault monitoring method provided in an embodiment of the present application;

[0047] Figure 4 A schematic diagram of a hardware monitoring structure provided for related technologies;

[0048] Figure 5 A schematic diagram of a hardware fault monitoring structure is provided for an embodiment of the present application;

[0049] Figure 6 A schematic diagram of a software indicator data reporting module provided in an embodiment of the present application;

[0050] Figure 7 A schematic diagram of a real-time monitoring module provided in an embodiment of the present application;

[0051] Figure 8 A schematic diagram of a hardware log data collection module provided in an embodiment of the present application;

[0052] Fig. 9 A schematic diagram of a hardware monitoring process provided in an embodiment of the present application;

[0053] Fig.10 A schematic diagram of the structure of a hardware fault monitoring device provided in an embodiment of the present application;

[0054] Fig.11 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;

[0055] Fig.12 A schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0057] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0058] Figure 1 A schematic diagram of a hardware monitoring system provided for related technologies, combined with Figure 1As shown in the figure, the HAPoxy service, Node service, data server, relational database, intelligent platform management interface and message middleware are all deployed with the hardware indicator collection tool Exporter. The hardware monitoring system periodically collects hardware logs through the deployed Exporter and uploads them to the hardware monitoring system. Then the hardware monitoring system analyzes the logs to determine whether there is a hardware failure.

[0059] However, the hardware log collection cycle is usually set to be longer, for example, the collection cycle of relational database logs is set to once every 10 minutes, in order to avoid the IO load generated by collecting hardware logs affecting the normal business operations. This will lead to a higher fault monitoring delay, and hardware faults will be difficult to be discovered in time, which will affect the fault repair timeliness, and may even cause system downtime and damage to related businesses.

[0060] In order to solve the above technical problems, an embodiment of the present application provides a hardware fault monitoring method and related devices, which determine whether there is a target log collection action that needs to be triggered based on the acquired software indicator data and preset abnormal conditions. The target log collection action is used to collect log data of the target hardware object. The target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions. When it is determined that there is a target log collection action that needs to be triggered, the hardware log collection device is controlled to execute the target log collection action to obtain the log data of the target hardware object, and determine whether the target hardware object has a fault based on the log data of the target hardware object.

[0061] In this way, by detecting whether the software indicator data is abnormal, it is determined whether there is a target hardware object that may fail. When it is determined that there is a target hardware object, the hardware log collection device is controlled to collect the log data of the target hardware object, so as to further confirm whether the target hardware object actually fails.

[0062] Since software indicator data is data that needs to be collected by the relevant business itself, the present application will not generate additional IO load due to the collection of software indicator data. On this basis, the present application can obtain the collected software indicator data based on a shorter cycle or in real time, and detect whether there is a target hardware object that may have a fault based on this more real-time software indicator data, and trigger the collection of log data of the target hardware object accordingly to further determine whether the target hardware object has a fault. Existing hardware faults can be detected more promptly, reducing hardware fault monitoring delays and further ensuring the timeliness of fault repair.

[0063] See also Figure 2 , this figure is a scenario schematic diagram of a hardware fault monitoring method provided in an embodiment of the present application, which may include a server 201 or a terminal device 202.

[0064] The server 201 or the terminal device 202 obtains the software indicator data. As an example, the server 200 can monitor the software indicator data of the device (such as a server, IoT hardware, etc.) in real time.

[0065] The server 201 or the terminal device 202 determines whether there is a target log collection action for collecting log data of the target hardware object that needs to be triggered by the detection result through the acquired software indicator data and the preset abnormal conditions. The target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions. As an example, assuming that the monitored equipment includes hardware objects such as a central processing unit (CPU), a graphics card, a hard disk, and a motherboard, the server can detect whether the acquired software indicator data meets the preset abnormal conditions. If it is detected that one or several software indicator data meet the preset abnormal conditions (such as abnormal access data of a data interface), the hardware object that affects the access of the data interface can be determined as the target hardware object, and the target log collection action that needs to be triggered to collect the log data of the target hardware object is determined accordingly.

[0066] When determining that there is a target log collection action that needs to be triggered, the server 201 or the terminal device 202 controls the hardware log collection device to execute the target log collection action to obtain the log data of the target hardware object.

[0067] The server 201 or the terminal device 202 determines whether the target hardware object has a fault based on the log data of the target hardware object. Specifically, the server 201 or the terminal device 202 can analyze whether the log data meets the fault standard based on the acquired log data of the target hardware object and the fault standard corresponding to the target hardware object, thereby determining whether the target hardware object actually has a fault.

[0068] In this way, by detecting whether the software indicator data is abnormal, it is determined whether there is a target hardware object that may have a fault. When it is determined that there is a target hardware object, the hardware log collection device is controlled to collect the log data of the target hardware object, so as to further confirm whether the target hardware object has actually failed. Since the software indicator data is data that the relevant business itself needs to collect, this application will not generate additional IO load due to the collection of software indicator data. On this basis, this application can obtain the collected software indicator data based on a shorter cycle or in real time, and detect whether there is a target hardware object that may have a fault based on this software indicator data with higher real-time performance, and trigger the collection of the log data of the target hardware object accordingly to further determine whether the target hardware object has a fault. The existing hardware fault can be detected more timely, the hardware fault monitoring delay can be reduced, and the fault repair timeliness can be further guaranteed.

[0069] The hardware fault monitoring method provided in the embodiment of the present application can be applied to a terminal device or server with data processing capabilities, and the terminal device includes but is not limited to a mobile phone, a tablet, a computer, a computer, a vehicle terminal, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0070] The hardware fault monitoring method provided in the embodiment of the present application involves cloud technology and database.

[0071] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0072] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, which needs to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.

[0073] In short, a database can be seen as an electronic filing cabinet - a place where electronic files are stored, and users can add, query, update, delete, and other operations on the data in the files. The so-called "database" is a collection of data that is stored together in a certain way, can be shared with multiple users, has as little redundancy as possible, and is independent of the application.

[0074] A database management system (DBMS) is a computer software system designed for managing databases. It generally has basic functions such as storage, retrieval, security, and backup. Database management systems can be classified according to the database model they support, such as relational, XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters, mobile phones; or according to the query language used, such as SQL (Structured Query Language), XQuery; or according to performance focus, such as maximum scale, maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMS can cross categories, for example, supporting multiple query languages ​​at the same time.

[0075] The relevant data collection and processing in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, obtain the informed consent or separate consent of the subject of personal information, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the subject of personal information.

[0076] See also Figure 3 , which is a flowchart of a hardware fault monitoring method provided in an embodiment of the present application.

[0077] Combination Figure 3 As shown, the hardware fault monitoring method provided in the embodiment of the present application may include:

[0078] S301: Obtain software indicator data.

[0079] Software indicator data is indicator data related to the business provided by the current monitoring object. Exemplarily, assuming that the current monitoring object is a server used to provide specific services to users, that is, it is currently necessary to monitor whether the server has hardware failures, accordingly, the software indicator data to be obtained can be the access data of the data interface of the server, and the user can experience the service provided by the server by accessing the data interface. Exemplarily, assuming that the current monitoring object is an Internet of Things device, that is, it is currently necessary to monitor whether the Internet of Things device has hardware failures, accordingly, the software indicator data to be obtained can be the access data of the data interface of the Internet of Things service.

[0080] It should be understood that the current monitoring object may include multiple hardware objects, such as processor, memory, chipset, I / O (RAID card, network card, HBA card), hard disk, chassis (power supply, fan), etc., and the embodiments of the present application do not impose any limitations on this.

[0081] It should be understood that in actual applications, the method provided in the embodiments of the present application can also be used to monitor whether other objects have hardware failures, and the embodiments of the present application do not impose any limitations on the monitored objects. In addition, the embodiments of the present application can also obtain other types of software indicator data, and the embodiments of the present application do not impose any limitations on the obtained software indicator data.

[0082] Among them, software indicator data is the indicator data that needs to be collected by the monitored object itself. Therefore, in the embodiment of the present application, no additional IO load will be generated due to the collection of software indicator data. Based on this, the software indicator data can be obtained in a shorter period, and even in real time.

[0083] S302: Determine, based on the software indicator data and the preset abnormal conditions, whether there is a target log collection action that needs to be triggered based on the detection result.

[0084] In the case of detecting and determining that the software indicator data meets the preset abnormal condition, the hardware object in the monitored object that can make the software indicator data meet the preset abnormal condition can be determined as the target hardware object, and then the target log collection action that needs to be triggered to collect the log data of the target hardware object is determined. Exemplarily, assuming that the access data of the data interface a in the acquired software indicator data is abnormal, the hardware object that affects the access situation of the data interface a can be determined as the target hardware object. For example, assuming that the hardware object that can affect the access situation of the data interface a is determined to be a hard disk, the hard disk can be determined as the target hardware object, and it is determined that the collection of the log data of the hard disk needs to be triggered.

[0085] S303: When it is determined that there is a target log collection action that needs to be triggered, control the hardware log collection device to execute the target log collection action to obtain log data of the target hardware object.

[0086] Hardware log collection devices refer to devices used to collect logs of hardware objects, such as Logstash, Filebeat, Fluentd, Logagent, rsyslog, and so on.

[0087] It should be understood that when there is an abnormality in the software indication data and it is determined based on the detection results that there is a target log collection action that needs to be triggered, the target hardware object can be log collected through the hardware log collection device. This is different from the periodic collection of hardware logs in the related technology. Since the software indicator data is data that needs to be collected by the related business itself, the present application will not generate additional IO load due to the collection of software indicator data.

[0088] S304: Determine whether the target hardware object has a fault according to the log data of the target hardware object.

[0089] It should be understood that after obtaining the log data of the target hardware object, the log data can be analyzed to determine whether the target hardware object has a hardware fault. For example, if the target hardware object is a memory, and the log data collected for the memory is the number of failed storages, assuming that the number of failed storages exceeds a preset threshold within a preset period, it is considered that the memory has a fault.

[0090] Different from the periodic hardware log collection in the related technology, software indicator data is data that needs to be collected by the related business itself. Therefore, this application will not generate additional IO load due to the collection of software indicator data. On this basis, this application can obtain the collected software indicator data based on a shorter period or in real time, and detect whether there are target hardware objects that may have faults based on this more real-time software indicator data. Existing hardware faults can be detected more promptly, reducing hardware fault monitoring delays and further ensuring the timeliness of fault repair.

[0091] Based on the hardware fault monitoring method provided in the above embodiment, in some possible implementations, the software indicator data includes access data of at least one data interface; step S302 may include:

[0092] Step 11: For each data interface, obtain the hardware monitoring trigger rule corresponding to the data interface.

[0093] A data interface refers to a set of regulations or protocols for data exchange and communication, which enables data to be shared between different systems or applications. In an embodiment of the present application, software indicator data can be obtained in units of data interfaces, and whether the access status of each interface is normal can be determined by one or more corresponding hardware objects. As an example, assuming that the software indicator data includes access data of data interface A and data interface B, the hardware objects affecting data interface A can be hardware a and hardware b, and the hardware object affecting data interface B can be hardware c.

[0094] The hardware monitoring trigger rule refers to the condition for triggering hardware monitoring. The hardware monitoring trigger rule is used to indicate the access data abnormality threshold corresponding to the data interface and the reference log collection action.

[0095] The access data abnormality threshold is a pre-set threshold used to measure whether the access to the data interface is abnormal, which may include, for example, an access delay threshold, an access failure rate threshold, etc. For example, the access data includes three types of data: A, B, and C; when the type A data exceeds the value a, it is abnormal, and the access data abnormality threshold is the value a; when the type B data is less than the value b, it is abnormal, and the access data abnormality threshold is the value b; when the type C data is in the interval [c, d], it is abnormal, and the access data abnormality threshold is [c, d].

[0096] The reference log collection action is used to collect log data of hardware objects that affect the access to the data interface. It should be understood that the data interface may correspond to at least one hardware object, and when the access data corresponding to the data interface is abnormal, the reference log collection action may indicate that the log data of at least one hardware object that affects the access to the data interface needs to be collected.

[0097] Access data refers to data generated by accessing a data interface. In some possible implementations, the access data of the data interface may include the number of accesses to the data interface in the current cycle, the number of erroneous accesses, and the access delay.

[0098] The number of accesses means the number of times the data interface is accessed within a cycle. For example, if the data interface is accessed 100 times within 10 minutes, the number of accesses is 100 times; the number of error accesses means the number of times a request is initiated to the data interface within a cycle but the request is not successful. For example, if 1000 download requests are initiated to the data interface in 10 minutes, but 200 downloads fail, the number of error accesses is 200 times; the access delay means the number of times the time difference between the moment a request is sent to the data interface and the moment the data interface responds to the request and returns the request result within a cycle is greater than the preset time. For example, if a request is sent to the data interface within 10 minutes, the request is sent at 9:00 and the request result is received at 9:05, and the preset time is 1 minute, then the access delay is 4 minutes.

[0099] Since the hardware monitoring triggering rules may change according to the actual application of the data interface, some possible implementation methods may also include:

[0100] Step 21: Obtain the latest change time corresponding to the acquired hardware monitoring trigger rule from the rule library for storing the hardware monitoring trigger rule corresponding to the data interface.

[0101] As an example, assuming that the change time of the hardware monitoring trigger rule a corresponding to the data interface A is 9:00 on September 1, 9:00 on October 30, and 9:00 on November 21, the latest change time is 9:00 on November 21.

[0102] Step 22: Based on the latest change time, determine whether the acquired hardware monitoring trigger rule has been updated; if it is determined that the acquired target hardware monitoring trigger rule has been updated, obtain the updated target hardware monitoring trigger rule from the rule library.

[0103] As an example, assuming that the change time of the hardware monitoring trigger rule a obtained by data interface A is 9:00 on October 30, and the latest change time of the hardware monitoring trigger rule a found in the rule base is 9:00 on November 21, it means that the hardware monitoring trigger rule a has changed. At this time, it is necessary to obtain the hardware monitoring trigger rule a corresponding to the most recent change time.

[0104] It should be understood that the hardware monitoring trigger rules are updated in a timely manner through steps 21 and 22, thereby avoiding the problem of hardware monitoring errors caused by inconsistency between the data interface and the hardware monitoring trigger rules, and improving the reliability of hardware monitoring.

[0105] Step 12: For each data interface, determine whether the access data of the data interface meets the preset abnormal conditions based on the access data of the data interface and the access data abnormality threshold indicated by the hardware monitoring trigger rule corresponding to the data interface; if so, determine the reference log collection action indicated by the hardware monitoring trigger rule corresponding to the data interface as the target log collection action.

[0106] The access data abnormality threshold refers to the numerical range in which the access data is abnormal. For example, if the access data is the number of accesses, the access data abnormality threshold is that the number of accesses per minute is less than 1,000.

[0107] Furthermore, assuming that the number of accesses is 200 times per minute, the access data meets the preset abnormal condition of "the number of accesses per minute is less than 1000 times", then the hardware monitoring trigger corresponding to the data interface is determined to be responsible for the corresponding reference log collection action, which is used as the target log collection action to collect log data for the target hardware object corresponding to the access data.

[0108] Specifically, for the access data of each data interface in the acquired software indicator data, it can be determined whether the access data reaches the access data anomaly threshold indicated by the hardware monitoring trigger rule corresponding to the data interface. If so, it can be determined that the access data of the data interface meets the preset anomaly adjustment, and the hardware object that affects the access situation of the data interface can be determined accordingly as the target hardware object, and the target log collection action that needs to be triggered to collect the log data of the target hardware object can be determined.

[0109] In a possible implementation, the access data of the data interface includes the number of accesses, the number of erroneous accesses, and the access delay of the data interface in the current cycle. For the access data of the data interface, a statistical cycle (such as 3 seconds) is usually set, and the access data of the data interface may include the number of accesses, the number of erroneous accesses, and the access delay in the current 3 seconds.

[0110] Wherein, when the access data of the data interface includes the number of accesses, the number of erroneous accesses, and the access delay of the data interface in the current cycle, step 12 may include:

[0111] Step 121: According to the number of accesses and the number of erroneous accesses in the access data, the access failure rate corresponding to the statistical data interface is counted.

[0112] As an example, assuming that the number of accesses is 100 and the number of incorrect accesses is 50, the access failure rate is 50%.

[0113] Step 122: According to the access delay in the access data, the average access delay and the maximum access delay corresponding to the statistical data interface are counted.

[0114] As an example, assuming that the number of accesses is 100 and the number of correct accesses is 50, the average access delay corresponding to the 50 correct accesses is 0.5 seconds and the maximum access delay is 1.5 seconds.

[0115] Step 123: Detect whether the access data meets the preset abnormal condition based on at least one of the relationship between the access failure rate and the failure rate threshold in the access data abnormality threshold, the relationship between the average access delay and the average delay threshold in the access data abnormality threshold, and the relationship between the maximum access delay and the maximum delay threshold in the access data abnormality threshold.

[0116] The access failure rate means the ratio between the number of failed accesses to the data interface and the total number of accesses to the data interface, which is used to indicate the ratio of access failures to the data interface; the average access delay is the ratio between the sum of the number of accesses to the data interface and the number of accesses, which is used to indicate the average delay for each access to the data access interface; the maximum access delay is the longest access delay among the access delays corresponding to each access, which is used to indicate the maximum access delay of the data access interface.

[0117] In an embodiment of the present application, according to the judgment conditions indicated by the hardware monitoring trigger rule, it is possible to determine whether the access data meets the preset abnormal condition based on at least one of the relationship between the access failure rate and the failure rate threshold in the access data abnormality threshold, the relationship between the average access delay and the average delay threshold in the access data abnormality threshold, and the relationship between the maximum access delay and the maximum delay threshold in the access data abnormality threshold.

[0118] It should be understood that in the embodiment of the present application, the software indicator data may include multiple data interfaces, each of which has a corresponding hardware monitoring trigger rule. Therefore, for the data interface that meets the preset abnormal conditions, the corresponding reference log collection action is used as the target log collection action to collect the log of the corresponding hardware object. That is, the embodiment of the present application monitors the hardware in units of data interfaces, and can monitor multiple hardware at the same time, thereby improving the efficiency of hardware monitoring.

[0119] Based on the hardware monitoring method provided in the above embodiment, in some possible implementations, step S303 may include:

[0120] Step 31: Generate a target control instruction according to the target log collection action, and send the target control instruction to the hardware log collection device.

[0121] The target control instruction refers to an instruction for controlling the hardware log collection device to collect logs for the target hardware object.

[0122] As an example, assuming that the target log collection action is an action to collect log data of the target hardware object being a processor, then the generated target control instruction may be a processor log data collection instruction, which is used to control the hardware log collection device to collect the log data of the processor.

[0123] As a possible implementation method, a log collection command set is stored in the hardware log collection device, and the log collection command set includes multiple log collection commands, and the log collection commands are used to collect log data of the corresponding hardware objects. For example, the log collection command set may include log collection commands A, B, and C. Log collection command A is used to collect log data of the processor, log collection command B is used to collect log data of the hard disk, and log collection command C is used to collect log data of the memory.

[0124] The hardware log collection device collects log data of the target hardware object in the following ways:

[0125] Step 41: parse the target control instruction to obtain the target log collection action.

[0126] By analyzing the target control instruction, the action of collecting the target hardware object can be determined, that is, the target log collection action can be determined.

[0127] Step 42: In the log collection command set, search for the target log collection command corresponding to the target log collection action, execute the target log collection command, and obtain the log data of the target hardware object.

[0128] Among them, the target log collection action indicates the action of collecting the target hardware object. Therefore, the corresponding target log collection command can be found in the log collection command set according to the target log collection action, and the target log collection command can be executed to obtain the log data of the target hardware object.

[0129] Step 32: Receive log data of the target hardware object collected by the hardware log collection component in response to the target control instruction.

[0130] The hardware log collection device will respond to the target control instruction, collect the log data of the target hardware object according to the target control instruction, and then send the collected log data of the target hardware object to the sending end of the target control instruction.

[0131] It should be understood that in the embodiments of the present application, the target control instructions are mainly generated through the target log collection action, and the target control instructions are used to control the target log collection device to collect log data for the target hardware object. This is different from the periodic collection of hardware logs in the related art. Since the software indicator data is data that needs to be collected by the related business itself, the present application will not generate additional IO load due to the collection of software indicator data. On this basis, the present application can obtain the collected software indicator data based on a shorter cycle or in real time, and detect whether there is a target hardware object that may have a fault based on this more real-time software indicator data, and trigger the collection of log data of the target hardware object accordingly to further determine whether the target hardware object has a fault. Existing hardware faults can be detected more promptly, reducing hardware fault monitoring delays, and further ensuring the timeliness of fault repair.

[0132] Based on the hardware monitoring method provided in the above embodiment, in a possible implementation manner, the method may further include:

[0133] Step 51: Obtain periodic hardware log data.

[0134] The periodic hardware log data is collected by the hardware collection device when executing the periodic log collection action. The periodic hardware log data includes the log data of each reference type of hardware object.

[0135] Periodic hardware log data refers to log data corresponding to multiple hardware objects in one period.

[0136] In some possible implementations, a log collection command set is stored in the hardware log collection device. The log collection command set includes a variety of log collection commands. The log collection commands are used to collect log data of the corresponding hardware objects. The log collection commands have corresponding execution cycles.

[0137] The hardware log collection device collects periodic hardware log data in the following ways:

[0138] For each log collection command, the log collection command is executed periodically according to the execution period corresponding to the log collection command, and periodic hardware log data of the hardware object corresponding to the log collection command is obtained.

[0139] It should be understood that each hardware object may correspond to a log collection command, and different log collection instructions may have different execution cycles. By periodically executing various log collection commands, periodic hardware log data corresponding to multiple hardware objects may be obtained.

[0140] Step 52: Determine whether a hardware object of each reference type has a fault according to the periodic hardware log data.

[0141] Among them, by analyzing whether the corresponding log data of each hardware object meets the preset abnormal threshold, it is determined whether the hardware object has a fault. For example, the log data corresponding to the data interface is the number of incorrect access times, and the preset abnormal threshold is greater than 10 times per minute. Assuming that the number of incorrect access times is 11 times per minute, it is considered that the data interface has a fault.

[0142] It should be noted that the embodiment of the present application further adds the periodic collection of hardware log data on the basis of hardware fault detection based on software indicator data, so as to periodically detect hardware faults, thereby ensuring that hardware faults are not missed and the reliability of hardware fault monitoring is guaranteed.

[0143] See also Figure 4 , the figure shows a hardware monitoring structure of the related technology, which mainly includes a hardware log data collection module, a hardware monitoring system and a hardware fault repair system. The implementation process of the hardware monitoring is mainly to periodically collect the hardware logs of the server and report them to the hardware monitoring system. The hardware monitoring system analyzes the reported hardware logs and reports the faulty hardware to the hardware fault repair system, which then repairs the hardware.

[0144] However, collecting hardware logs will cause a certain input / output (IO) load on the operating system. In order to prevent the generated IO load from affecting the normal operation of the business, the collection period of hardware logs is usually set to be longer, which will lead to a higher fault monitoring delay, making it difficult for hardware faults to be discovered in time, thereby affecting the timeliness of fault repair.

[0145] In order to solve the above technical problems, see Figure 5 The embodiment of the present application provides a hardware fault monitoring structure, which may include a software indicator data reporting module, a hardware log data collection module, a real-time monitoring module, a hardware monitoring system and a hardware fault repair system.

[0146] Combination Figure 6 As shown, the software indicator data reporting module is deployed on the server and can cooperate with the business module to obtain software indicator data.

[0147] The business module is a program that provides services to external users, including but not limited to http services, download services, login services, etc. The business module will record the number of accesses to the data interface, the number of incorrect accesses, access delays, and error code information, and write them to the log file. The error code information refers to the error type information corresponding to the incorrect access.

[0148] The software indicator data reporting module can read the log file stored in the business module, convert the log file into a json structure, obtain the software indicator data, and report it to the real-time monitoring module. The software indicator data includes access data of at least one data interface, and each data interface access data includes the number of accesses, the number of erroneous accesses, and the access delay.

[0149] Combination Figure 7 As shown, the implementation process of the real-time monitoring module can be:

[0150] A1: Use the object to configure the hardware monitoring trigger rule through the web page, and save the hardware monitoring trigger rule in the configuration system.

[0151] A2: When the server starts, all hardware monitoring trigger rules are read from the configuration system. At the same time, the program will periodically read the latest change time in the rule library. If there is an update, the updated target hardware monitoring trigger rule is obtained from the rule library.

[0152] A3: The real-time monitoring module receives the software indicator data reported by the software indicator data reporting module and writes it into ClickHouse for storage.

[0153] A4: The real-time monitoring module calculates the access failure rate based on the number of accesses and the number of erroneous accesses to the data interface, as well as the average access delay and the maximum access delay based on the data interface in a one-minute cycle; then determines whether the access failure rate, the average access delay and the maximum access delay reach the access data abnormality threshold. If the access data abnormality threshold is reached, the hardware log data collection module deployed on the server is actively triggered to send the target log collection action to the hardware log data collection module.

[0154] Among them, the hardware monitoring trigger rule may include three parts: basic information, preset abnormal conditions and reference log collection action. Among them, the basic information may include the data interface name and rule description; the preset abnormal conditions may be divided into three conditions, namely, the relationship between the access failure rate and the failure rate threshold in the access data abnormal threshold, the relationship between the average access delay and the average delay threshold in the access data abnormal threshold, and the relationship between the maximum access delay and the maximum delay threshold in the access data abnormal threshold, and the relationship between these three conditions is at least one; the reference log collection action is used to collect log data of hardware objects that affect the access to the data interface.

[0155] As an example, the hardware monitoring trigger rule may be:

[0156] Basic information: Data interface: picDownload; Rule description: Picture download interface.

[0157] Access data abnormality thresholds: Access failure rate: greater than 1%; Average access delay: greater than 1000ms; Maximum access delay: greater than 5000ms.

[0158] Refer to the log collection action: Collect smart logs of the image download interface and image upload interface respectively.

[0159] See also Figure 8 , which is a schematic diagram of a hardware log data collection module provided in an embodiment of the present application.

[0160] After receiving the target log collection action sent by the real-time monitoring module, the hardware log data collection module collects the hardware log and operating system log of the corresponding target hardware object according to the target log collection action, and reports them to the hardware monitoring system.

[0161] Among them, the hardware log data collection module can realize triggered collection and timed collection.

[0162] Scheduled collection: The hardware log data collection module sets different log collection commands based on different hardware. By weighing the monitoring sensitivity and the impact on business load, different collection cycles are set, ranging from 2 minutes to 1 hour.

[0163] Trigger collection: After receiving the target control instruction corresponding to the target log collection action of the real-time monitoring module, find the target log collection command corresponding to the target log collection action in the log collection command set, execute the target log collection command, and obtain the log data of the target hardware object.

[0164] Log collection commands can be reused, such as collecting smart data through the smartctl command and collecting dmesg information through the dmesg command. These commands are uniformly encapsulated and can be directly called in scheduled collection and triggered collection.

[0165] See also Fig. 9 , which is a schematic diagram of a hardware monitoring process provided in an embodiment of the present application.

[0166] Combination Fig. 9As shown, the software indicator data reporting module reports to the real-time monitoring module. After detecting the abnormality, the real-time monitoring module sends the target control instruction corresponding to the target log collection action to the hardware log data collection module; the hardware log data collection module collects the log data of the target hardware object according to the target control instruction, and uploads the log data to the hardware monitoring system; after receiving the log data, the hardware monitoring system parses the hardware logs and system logs reported by the massive servers in the whole network in real time according to the pre-configured policy rules, finds the fault of the target hardware object, and creates a ticket in the fault repair system; the hardware fault repair system creates a fault ticket and manages each node of the fault repair process.

[0167] See also Fig.10 , which is a structural schematic diagram of a hardware fault monitoring device provided in an embodiment of the present application.

[0168] Based on the hardware fault monitoring method provided in the above embodiment, the hardware fault monitoring device 1000 provided in the embodiment of the present application may include:

[0169] The software data acquisition module 1001 is used to acquire software indicator data;

[0170] The log collection triggering module 1002 is used to determine whether there is a target log collection action that needs to be triggered according to the software indicator data and the preset abnormal condition; the target log collection action is used to collect the log data of the target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal condition;

[0171] The log collection module 1003 is used to control the hardware log collection device to execute the target log collection action and obtain the log data of the target hardware object when it is determined that there is a target log collection action to be triggered;

[0172] The fault detection module 1004 is used to determine whether the target hardware object has a fault according to the log data of the target hardware object.

[0173] As an example, the software indicator data includes access data of at least one data interface; the log collection trigger module 1002 includes:

[0174] An acquisition unit, for acquiring, for each data interface, a hardware monitoring trigger rule corresponding to the data interface; the hardware monitoring trigger rule is used to indicate an abnormal threshold of access data corresponding to the data interface and a reference log collection action, and the reference log collection action is used to collect log data of hardware objects that affect the access of the data interface;

[0175] The log collection trigger unit is used to determine, for each data interface, whether the access data of the data interface meets the preset abnormal conditions based on the access data of the data interface and the access data abnormality threshold indicated by the hardware monitoring trigger rule corresponding to the data interface; if so, determine the reference log collection action indicated by the hardware monitoring trigger rule corresponding to the data interface as the target log collection action.

[0176] As an example, the access data of the data interface includes the number of accesses, the number of erroneous accesses, and the access delay of the data interface in the current cycle; the log collection trigger unit includes:

[0177] A calculation subunit is used to calculate the access failure rate corresponding to the statistical interface according to the number of accesses and the number of erroneous accesses in the access data; and to calculate the average access delay and the maximum access delay corresponding to the statistical interface according to the access delay in the access data;

[0178] The log collection trigger subunit is used to determine whether the access data meets the preset abnormal condition based on at least one of the relationship between the access failure rate and the failure rate threshold in the access data abnormal threshold, the relationship between the average access delay and the average delay threshold in the access data abnormal threshold, and the relationship between the maximum access delay and the maximum delay threshold in the access data abnormal threshold.

[0179] As an example, the apparatus 1000 further includes:

[0180] An acquisition module, used to acquire the latest change time corresponding to the acquired hardware monitoring trigger rule from a rule library for storing the hardware monitoring trigger rule corresponding to the data interface;

[0181] The update module is used to determine whether the acquired hardware monitoring trigger rule has been updated based on the latest change time; if it is determined that the acquired target hardware monitoring trigger rule has been updated, the updated target hardware monitoring trigger rule is obtained from the rule library.

[0182] As an example, the log collection module 1003 includes:

[0183] A target control instruction generation and sending unit, used to generate a target control instruction according to a target log collection action, and send the target control instruction to a hardware log collection device;

[0184] The receiving unit is used to receive the log data of the target hardware object collected by the hardware log collection device in response to the target control instruction.

[0185] As an example, a log collection command set is stored in the hardware log collection device, and the log collection command set includes multiple log collection commands, and the log collection commands are used to collect log data of the corresponding hardware objects;

[0186] The hardware log collection device collects log data of the target hardware object in the following ways:

[0187] A parsing unit, used to parse the target control instruction and obtain the target log collection action;

[0188] The collection unit is used to search for a target log collection command corresponding to a target log collection action in a log collection command set, execute the target log collection command, and obtain log data of a target hardware object.

[0189] As an example, the apparatus 1000 further includes:

[0190] The periodic acquisition module is used to acquire periodic hardware log data; the periodic hardware log data is collected by the hardware acquisition device performing the periodic log acquisition action, and the periodic hardware log data includes the log data of each reference type of hardware object;

[0191] The periodic fault detection module is used to determine whether a hardware object of each reference type has a fault based on periodic hardware log data.

[0192] As an example, a log collection command set is stored in the hardware log collection device, and the log collection command set includes multiple log collection commands. The log collection commands are used to collect log data of the corresponding hardware objects, and the log collection commands have corresponding execution cycles;

[0193] The hardware log collection device collects periodic hardware log data in the following ways:

[0194] The periodic collection unit is used to periodically execute the log collection command according to the execution period corresponding to each log collection command, and obtain the periodic hardware log data of the hardware object corresponding to the log collection command.

[0195] The hardware fault monitoring device provided in the embodiment of the present application has the same beneficial effects as the hardware fault monitoring method provided in the above embodiment, so they are not described in detail.

[0196] The embodiment of the present application also provides a computer device, which may specifically be a terminal device or a server. The terminal device and the server provided in the embodiment of the present application will be introduced below from the perspective of hardware entity.

[0197] See also Fig.11 , Fig.11 Schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Fig.11For the sake of convenience, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant, a point of sales (POS), a car computer, etc., taking the terminal as a computer as an example:

[0198] Fig.11 FIG. 1 is a block diagram showing a partial structure of a computer related to a terminal provided in an embodiment of the present application. Fig.11 The computer includes: a radio frequency (RF) circuit 1210, a memory 1220, an input unit 1230 (including a touch panel 1231 and other input devices 1232), a display unit 1240 (including a display panel 1241), a sensor 1250, an audio circuit 1260 (which can be connected to a speaker 1261 and a microphone 1262), a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290. Those skilled in the art can understand that Fig.11 The computer structure shown in the figure does not constitute a limitation of the computer, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0199] The memory 1220 can be used to store software programs and modules. The processor 1280 executes various functional applications and data processing of the computer by running the software programs and modules stored in the memory 1220. The memory 1220 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the computer (such as audio data, a phone book, etc.), etc. In addition, the memory 1220 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0200] The processor 1280 is the control center of the computer, which uses various interfaces and lines to connect various parts of the entire computer, and executes various functions of the computer and processes data by running or executing software programs and / or modules stored in the memory 1220, and calling data stored in the memory 1220. Optionally, the processor 1280 may include one or more processing units; preferably, the processor 1280 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1280.

[0201] In the embodiment of the present application, the processor 1280 included in the terminal also has the following functions:

[0202] Get software indicator data;

[0203] According to the software indicator data and the preset abnormal conditions, determine whether there is a target log collection action that needs to be triggered; the target log collection action is used to collect log data of the target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions;

[0204] When it is determined that there is a target log collection action that needs to be triggered, controlling the hardware log collection device to execute the target log collection action to obtain log data of the target hardware object;

[0205] Determine whether the target hardware object has a fault based on the log data of the target hardware object.

[0206] Optionally, the processor 1280 is further used to execute the steps of any implementation method of the hardware fault monitoring method provided in the embodiments of the present application.

[0207] See also Fig.12 , Fig.12A schematic diagram of the structure of a server 1300 provided for an embodiment of the present application. The server 1300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1322 (for example, one or more processors) and a memory 1332, and one or more storage media 1330 (for example, one or more mass storage devices) storing application programs 1342 or data 1344. Among them, the memory 1332 and the storage medium 1330 may be temporary storage or permanent storage. The program stored in the storage medium 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1322 may be configured to communicate with the storage medium 1330 to execute a series of instruction operations in the storage medium 1330 on the server 1300.

[0208] The server 1300 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and / or one or more operating systems, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.

[0209] The steps performed by the server in the above embodiment can be based on the Fig.12 The server structure shown.

[0210] The CPU 1322 is used to execute the following steps:

[0211] Get software indicator data;

[0212] According to the software indicator data and the preset abnormal conditions, determine whether there is a target log collection action that needs to be triggered; the target log collection action is used to collect log data of the target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal conditions;

[0213] When it is determined that there is a target log collection action that needs to be triggered, controlling the hardware log collection device to execute the target log collection action to obtain log data of the target hardware object;

[0214] Determine whether the target hardware object has a fault based on the log data of the target hardware object.

[0215] Optionally, the CPU 1322 may also be used to execute steps of any implementation of the hardware fault monitoring method provided in the embodiments of the present application.

[0216] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program, wherein the computer program is used to execute any one of the implementation methods of a hardware fault monitoring method described in the aforementioned embodiments.

[0217] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes any one of the implementation methods of a hardware fault monitoring method described in the above embodiments.

[0218] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0219] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0220] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0221] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store computer programs.

[0223] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0224] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for monitoring hardware failure, characterized in that: The method comprises: Get software indicator data; According to the software indicator data and the preset abnormal condition, determining whether there is a target log collection action that needs to be triggered; the target log collection action is used to collect log data of a target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal condition; In the case of determining that there is the target log collection action that needs to be triggered, controlling the hardware log collection device to execute the target log collection action to obtain the log data of the target hardware object; Determine whether the target hardware object has a fault according to the log data of the target hardware object.

2. The method according to claim 1, characterized in that The software indicator data includes access data of at least one data interface; The determining, based on the software indicator data and the preset abnormal condition, whether there is a target log collection action that needs to be triggered includes: For each of the data interfaces, obtaining a hardware monitoring trigger rule corresponding to the data interface; the hardware monitoring trigger rule is used to indicate an abnormal threshold of access data corresponding to the data interface and a reference log collection action, and the reference log collection action is used to collect log data of hardware objects that affect access to the data interface; For each of the data interfaces, determine whether the access data of the data interface meets the preset abnormal condition based on the access data of the data interface and the access data abnormality threshold indicated by the hardware monitoring trigger rule corresponding to the data interface; if so, determine the reference log collection action indicated by the hardware monitoring trigger rule corresponding to the data interface as the target log collection action.

3. The method according to claim 2, characterized in that The access data of the data interface includes the number of accesses, the number of erroneous accesses, and the access delay of the data interface in the current cycle; The determining, based on the access data of the data interface and the access data abnormality threshold indicated by the hardware monitoring trigger rule corresponding to the data interface, whether the access data of the data interface meets the preset abnormality condition includes: According to the number of accesses and the number of erroneous accesses in the access data, the access failure rate corresponding to the statistical data interface is calculated; according to the access delay in the access data, the average access delay and the maximum access delay corresponding to the data interface are calculated; Determine whether the access data meets the preset abnormal condition based on at least one of the relationship between the access failure rate and the failure rate threshold in the access data abnormality threshold, the relationship between the average access delay and the average delay threshold in the access data abnormality threshold, and the relationship between the maximum access delay and the maximum delay threshold in the access data abnormality threshold.

4. The method according to claim 2, characterized in that: The method further comprises: Obtaining the latest change time corresponding to the acquired hardware monitoring trigger rule from a rule library for storing the hardware monitoring trigger rule corresponding to the data interface; Based on the latest change time, determine whether the acquired hardware monitoring trigger rule has been updated; if it is determined that the acquired target hardware monitoring trigger rule has been updated, obtain the updated target hardware monitoring trigger rule from the rule library.

5. The method according to claim 1, characterized in that The controlling hardware log collection device to execute the target log collection action to obtain the log data of the target hardware object includes: Generate a target control instruction according to the target log collection action, and send the target control instruction to the hardware log collection device; Receive log data of the target hardware object collected by the hardware log collection device in response to the target control instruction.

6. The method according to claim 5, characterized in that: The hardware log collection device stores a log collection command set, wherein the log collection command set includes a plurality of log collection commands, and the log collection commands are used to collect log data of the corresponding hardware objects; The hardware log collection device collects log data of the target hardware object in the following manner: Parsing the target control instruction to obtain the target log collection action; In the log collection command set, a target log collection command corresponding to the target log collection action is searched, the target log collection command is executed, and the log data of the target hardware object is obtained.

7. The method according to claim 1, characterized in that The method further comprises: Acquire periodic hardware log data; the periodic hardware log data is collected by the hardware collection device performing a periodic log collection action, and the periodic hardware log data includes log data of each reference type of hardware object; According to the periodic hardware log data, it is determined whether each hardware object of the reference type has a fault.

8. The method according to claim 7, characterized in that The hardware log collection device stores a log collection command set, wherein the log collection command set includes a plurality of log collection commands, wherein the log collection commands are used to collect log data of the corresponding hardware objects, and the log collection commands have corresponding execution cycles; The hardware log collection device collects the periodic hardware log data in the following manner: For each of the log collection commands, the log collection command is periodically executed according to the execution period corresponding to the log collection command to obtain periodic hardware log data of the hardware object corresponding to the log collection command.

9. A hardware fault monitoring device, characterized in that: The device comprises: A software data acquisition module is used to acquire software indicator data; A log collection triggering module is used to determine whether there is a target log collection action that needs to be triggered according to the software indicator data and the preset abnormal condition; the target log collection action is used to collect log data of a target hardware object, and the target hardware object is a hardware object that can make the software indicator data meet the preset abnormal condition; A log collection module, configured to control the hardware log collection device to execute the target log collection action to obtain the log data of the target hardware object when it is determined that there is the target log collection action to be triggered; The fault detection module is used to determine whether the target hardware object has a fault according to the log data of the target hardware object.

10. A computer device, characterized in that: The computer device includes a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the hardware fault monitoring method according to any one of claims 1 to 8 according to the computer program.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the hardware fault monitoring method according to any one of claims 1 to 8.