Fault monitoring method and device, equipment and storage medium
By using fault identification models and fault probability judgment methods in large-scale training clusters, the faults in training clusters are accurately identified and monitored, and the problem of frequent failures and insufficient monitoring capabilities in training clusters is solved, and high-accurate fault monitoring is achieved.
Patent Information
- Application Number
- CN202510078324.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
In large-scale training clusters, how to accurately identify and monitor hardware and software failures, especially when there are many nodes, the frequency of failures is high.
By obtaining logs of nodes, software and training processes in the training cluster, these logs are used as inputs to the preset fault identification model, identifying fault information, and determining the fault probability based on the fault information of the fault object in the preset time window, and then determining the fault occurrence.
A more comprehensive fault monitoring is achieved, the accuracy and capability of fault monitoring is improved, and fault missed and missed faults are prevented.
Smart Images

Figure CN119987943A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of fault monitoring, and in particular to a fault monitoring method, device, equipment and storage medium. Background Art
[0002] Currently, large-scale training clusters are required for training large models. During large-scale model training, it is necessary to monitor the faults (including hardware faults, software faults, and training process working status) generated in the training cluster to ensure the availability of the training cluster. However, in actual large-scale training, hardware and software faults in the training cluster are very frequent. The more nodes in the training cluster, the higher the frequency of faults. Therefore, in this case, how to accurately identify the faults that occur in the training cluster is a technical problem that needs to be solved. Summary of the invention
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a fault monitoring method, device, equipment and storage medium.
[0004] In a first aspect, an embodiment of the present disclosure provides a fault monitoring method, the method comprising:
[0005] Obtaining a first log of a node included in the training cluster, a second log of the software, and a training log of the training process;
[0006] Using the first log, the second log, and the training log as inputs of a preset fault identification model, and obtaining fault information of the training cluster based on the fault identification model, wherein the fault information includes information of a fault object;
[0007] For any fault object, determining the fault probability of the fault object based on the fault information of the fault object within a preset time window;
[0008] In response to the failure probability of the fault object being higher than a first preset threshold, it is determined that the fault object fails.
[0009] Optionally, for any fault object, determining the fault probability of the fault object based on fault information of the fault object within a preset time window includes:
[0010] For at least one fault occurring in the fault object within the preset time window, fault information of a target fault is filtered out from the fault information of the at least one fault, wherein the target fault refers to a fault occurring at a frequency less than a preset frequency within the preset time window.
[0011] Optionally, for any fault object, determining the fault probability of the fault object based on fault information of the fault object within a preset time window includes:
[0012] The fault information of the fault object within a preset time window is input into a preset decision model, and the fault probability of the fault object is determined based on the decision model.
[0013] Optionally, the method of the first aspect further includes:
[0014] In response to the failure probability of the fault object being lower than a second preset threshold, it is determined that the fault object has not failed, and the second preset threshold is lower than the first preset threshold.
[0015] Optionally, the method of the first aspect further includes:
[0016] In response to the fault probability of the fault object being greater than the second preset threshold and less than the first preset threshold, inputting the fault information of the fault object within the preset time window into the fault identification model, and determining the query object and the query command based on the fault identification model;
[0017] Querying information from the query object based on the query command, where the queried information is used to assist in determining whether the fault object has a fault;
[0018] The queried information is stored in the log of the fault object.
[0019] Optionally, the method of the first aspect further includes:
[0020] Acquire a training sample, wherein the training sample includes fault information of a target fault object within the preset time window and a fault label of the target fault object, wherein the target fault object is any node, software or training process in any training cluster;
[0021] Based on the fault information and fault label of the target fault object, a preset first model is trained into the decision model.
[0022] Optionally, the method of the first aspect further includes:
[0023] Training a preset second model based on an instruction manual of query commands in a training cluster, technical information related to the training cluster, and log information, so that the second model learns knowledge related to the training cluster, wherein the instruction manual is used to explain and illustrate the query commands in the training cluster;
[0024] Based on the sample logs and training labels of at least one training cluster, training the second model that has learned the knowledge to obtain the fault identification model;
[0025] The sample log of the at least one training cluster includes logs of nodes, software and training processes in the at least one training cluster;
[0026] The training label includes real fault information of the at least one training cluster and a query operation performed when the real fault information is identified, wherein the query operation includes a query object and a query command.
[0027] In a second aspect, an embodiment of the present disclosure provides a fault monitoring device, the device comprising:
[0028] A log collection module, used to obtain a first log of a node included in the training cluster, a second log of the software, and a training log of the training process;
[0029] a fault identification module, configured to use the first log, the second log, and the training log as inputs of a preset fault identification model, and obtain fault information of the training cluster based on the fault identification model, wherein the fault information includes information of a fault object;
[0030] A fault decision module, for determining, for any fault object, the fault probability of the fault object based on the fault information of the fault object within a preset time window;
[0031] The fault decision module is further configured to determine that a fault occurs to the fault object in response to a fault probability of the fault object being higher than a first preset threshold.
[0032] Optionally, the fault decision module is used to:
[0033] For at least one fault occurring in the fault object within the preset time window, fault information of a target fault is filtered out from the fault information of the at least one fault, wherein the target fault refers to a fault occurring at a frequency less than a preset frequency within the preset time window.
[0034] Optionally, the fault decision module is used to:
[0035] The fault information of the fault object within a preset time window is input into a preset decision model, and the fault probability of the fault object is determined based on the decision model.
[0036] Optionally, the fault decision module is used to: determine that the fault object has not failed in response to the fault probability of the fault object being lower than a second preset threshold, and the second preset threshold is smaller than the first preset threshold.
[0037] Optionally, the fault monitoring device may further include:
[0038] a supplementary query module, configured to: in response to a fault probability of the fault object being greater than the second preset threshold and less than the first preset threshold, input the fault information of the fault object within the preset time window into the fault identification model, and determine a query object and a query command based on the fault identification model;
[0039] A query execution module is used to query information from the query object based on the query command, the queried information is used to assist in determining whether the fault object has a fault; and store the queried information in the log of the fault object.
[0040] Optionally, the fault monitoring device may further include:
[0041] The first training module is used to:
[0042] Acquire a training sample, wherein the training sample includes fault information of a target fault object within the preset time window and a fault label of the target fault object, wherein the target fault object is any node, software or training process in any training cluster;
[0043] Based on the fault information and fault label of the target fault object, a preset first model is trained into the decision model.
[0044] Optionally, the fault monitoring device may further include:
[0045] The second training module is used to:
[0046] Training a preset second model based on an instruction manual of query commands in a training cluster, technical information related to the training cluster, and log information, so that the second model learns knowledge related to the training cluster, wherein the instruction manual is used to explain and illustrate the query commands in the training cluster;
[0047] Based on the sample logs and training labels of at least one training cluster, training the second model that has learned the knowledge to obtain the fault identification model;
[0048] The sample log of the at least one training cluster includes logs of nodes, software and training processes in the at least one training cluster;
[0049] The training label includes real fault information of the at least one training cluster and a query operation performed when the real fault information is identified, wherein the query operation includes a query object and a query command.
[0050] In a third aspect, an embodiment of the present disclosure provides a computer device, the computer device comprising:
[0051] Memory;
[0052] Processor; and
[0053] Computer programs;
[0054] The computer program is stored in the memory and is configured to be executed by the processor to implement any method as described in the first aspect.
[0055] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement any of the methods described in the first aspect.
[0056] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including computer program instructions, which can implement any method described in the first aspect when executed by a processor.
[0057] The fault monitoring method, apparatus, device and storage medium provided by the embodiments of the present disclosure obtain the first log of the nodes included in the training cluster, the second log of the software and the training log of the training process, use the first log, the second log and the training log as the input of the preset fault identification model, and obtain the fault information of the training cluster through the fault identification model. Compared with the string matching method of the related technology, more comprehensive fault monitoring can be achieved, the fault monitoring capability is improved, and missed faults can be prevented. For any fault object, the fault probability of the fault object is determined based on the fault information of the fault object within a preset time window, and when the fault probability of the fault object is higher than the first preset threshold, it is determined that the fault object has a fault, which can prevent false detection and improve the accuracy of fault monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0060] Figure 1 is a flow chart of a model training method provided by an embodiment of the present disclosure;
[0061] Figure 2 is a flow chart of another model training method provided by an embodiment of the present disclosure;
[0062] Figure 3 is a flowchart of a fault monitoring method provided by an embodiment of the present disclosure;
[0063] Figure 4 is a flow chart of a method for determining a failure probability of a fault object provided by an embodiment of the present disclosure;
[0064] Figure 5 is a flow chart of an information query method provided by an embodiment of the present disclosure;
[0065] Figure 6 is a structural schematic diagram of a fault monitoring device provided by an embodiment of the present disclosure;
[0066] Figure 7 A schematic diagram of the structure of a computer device embodiment provided in the present disclosure. DETAILED DESCRIPTION
[0067] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0068] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0069] At present, related technologies generally use supporting fault monitoring software to monitor faults in training clusters.
[0070] However, there are the following problems in actual operation:
[0071] 1. Poor fault monitoring capability: The fault monitoring software used in related technologies can only match the pre-summarized fault information from the logs of the nodes, the logs of the software, the training logs of the training process and other information contained in the training cluster based on the string hard matching method. Faults that have not been summarized and covered in advance cannot be monitored. However, in the actual large model training process, unsummarized and uncovered faults often occur. In fact, it is also difficult to cover all kinds of hardware components, commands (for example, shell commands, but not limited to shell commands), software and other faults in the training cluster with text rules. Therefore, it is impossible to monitor all faults in the training cluster through the string hard matching method, and the fault monitoring capability is poor.
[0072] 2. Low accuracy of fault monitoring: In some scenarios, the accuracy of fault information identified based on logs is low. For example, when a node in the training cluster cannot be logged in through the Secure Shell Protocol (SSH), it may be that the node is faulty or that it is temporarily unable to log in due to temporary network reasons. After the network is restored, the node can still be logged in, and the node is not faulty at this time. In this case, the relevant technology will often judge it as a node failure.
[0073] 3. Lack of fault information decision-making capabilities: Related technologies can only monitor fault information by hard matching of strings. It is unable to comprehensively analyze the fault information within a time window to determine the true fault information, cannot eliminate misjudgments, and cannot determine the fault probability of the fault object based on all the fault information of the fault object within a time window. At the same time, it is also unable to make further query operations based on the fault information, nor can it assist fault monitoring through further query operations and query results, and the accuracy of fault monitoring is low.
[0074] In view of the above problems, the embodiments of the present disclosure provide a fault monitoring method, device, equipment and storage medium. The solutions provided by the embodiments of the present disclosure are described below in conjunction with exemplary embodiments.
[0075] For example, Figure 1 is a flow chart of a model training method provided by an embodiment of the present disclosure. The method can be exemplarily performed by a computer device, which can be any device with computing and processing capabilities. Figure 1 As shown, in some implementations, the model training method provided by the embodiments of the present disclosure may include steps 101 and 102.
[0076] Step 101: Train a preset second model based on the instruction manual of the query command in the training cluster, technical information related to the training cluster, and log information, so that the second model learns knowledge related to the training cluster.
[0077] The instruction manual of the query command (such as shell command, but not limited to shell command) is used to explain the query command involved in the training cluster, such as the specific form, function and corresponding output of the query command.
[0078] The technical information related to the training cluster can be understood as all technical information related to the training cluster collected through various channels. For example, in some embodiments, the technical information related to the training cluster may include technical documents, scientific literature, files and codes of open source projects related to the training cluster, and faults that occur in open source projects. The collection channels of the above technical information include technical question and answer websites, platforms for providing technical literature, and sharing platforms for open source projects. Of course, this is only an example of the collection channel of technical information and is not the only limitation.
[0079] The log information includes logs of nodes in the existing training cluster, logs of software (such as task scheduling software, but not limited to task scheduling software), training logs of the training process, etc. These logs include fault information in the existing training cluster.
[0080] The second model referred to in the embodiments of the present disclosure may be any artificial intelligence model. For example, in one example, the second model may be a large language model (such as llama3) but is not limited to a large language model.
[0081] The disclosed embodiment may exemplarily adopt a self-supervised learning method to train the second model, so that the second model learns knowledge related to the training cluster of the large model from the instruction manual of the query command, technical information related to the training cluster, and log information. The specific training method can refer to the self-supervised learning method provided by the relevant technology, which will not be repeated here.
[0082] Step 102: Based on the sample logs and training labels of at least one training cluster, the second model that has learned the knowledge is trained to obtain a fault recognition model.
[0083] The sample log of at least one training cluster referred to in the embodiment of the present disclosure may include logs of nodes, software and training processes in the training clusters. The sample log of at least one training cluster referred to in the embodiment of the present disclosure may be obtained from an existing training cluster.
[0084] The training label referred to in the embodiment of the present disclosure may include the real fault information of at least one training cluster mentioned above, and when the real fault information is identified, a further query operation is taken to accurately determine the cause and location of the fault, and the query operation at least includes: a query object and a query command. In some embodiments, the query operation may also include a logical explanation for taking the query operation. Among them, the query object may be a node, a software or a process in the training cluster. Among them, the content of the query for different query objects may be different, and the query commands used may also be different. However, the query commands used are all included in the instruction manual of the query command.
[0085] In some implementations, the embodiments of the present disclosure may use a supervised training method to further train the second model that has learned the knowledge related to the training cluster, so that the trained model (i.e., the fault identification model referred to in the embodiments of the present disclosure) has the following capabilities, wherein the supervised training method can refer to the relevant technology and will not be described in detail here:
[0086] 1. Ability to identify fault information based on the logs of nodes, software, and training processes in the training cluster. The fault information may include the fault object (such as the specific node, software, or process that has failed), fault content (such as being unable to log in to a node via SSH, but not limited to the fault content listed here), the IP address of the fault object, and the time when the fault occurred.
[0087] 2. Based on the fault information, decide whether additional query information is needed, what information to query, and provide the query object and query command capabilities.
[0088] 3. After deciding the query object and query command, give a logical explanation of the above decision.
[0089] In the disclosed embodiment, a preset second model is trained based on the instruction manual of the query command in the training cluster, the technical information related to the training cluster, and the log information, so that the second model learns the knowledge related to the training cluster, and the second model is further trained based on the sample log and training labels of at least one training cluster to obtain a fault identification model, and then more comprehensive fault information can be identified through the fault identification model, especially faults that cannot be summarized by text rules, and further query objects and query commands can be given, and further query operations can be used to assist in the identification of faults occurring in the training cluster, thereby improving the accuracy of fault monitoring.
[0090] Figure 2 FIG. 1 is a flow chart of another model training method provided by an embodiment of the present disclosure. Figure 2 As shown, in some implementations, the model training method provided by the embodiments of the present disclosure may include steps 201-202.
[0091] Step 201: Acquire training samples, where the training samples include fault information of a target fault object within a preset time window and a fault label of the target fault object, wherein the target fault object is any node, software or training process in any training cluster.
[0092] Among them, the size of the preset time window can be set as needed, and is not limited to a specific window size. For ease of understanding, the preset time window can be exemplarily understood as 15 minutes in the embodiment of the present disclosure. That is, in an example of the embodiment of the present disclosure, the fault information and fault label of the target fault object within every 15 minutes can be used as a training sample. Among them, the fault label of the target fault object is used to indicate whether the target fault object has a fault and what fault has occurred.
[0093] It should be noted that the training samples referred to in the embodiments of the present disclosure may include training samples of multiple target fault objects within multiple 15-minute periods.
[0094] Step 202: Based on the fault information and fault label of the target fault object, a preset first model is trained into a decision model.
[0095] The first model may be any artificial intelligence model, such as a large language model, a neural network model, a self-attention model (Transformer), etc., but is not limited to the models listed here.
[0096] The embodiment of the present disclosure may use a supervised model training method to train the first model, so that the trained model (i.e., the decision model referred to in the embodiment of the present disclosure) can identify the probability of the fault object actually failing based on the fault information of the fault object within a preset time window. The process of training the first model based on the supervised model training method can refer to the relevant technology, and will not be described in detail here.
[0097] In the embodiment of the present disclosure, training samples are obtained, and the training samples include fault information of the target fault object within a preset time window and the fault label of the target fault object. Based on the fault information and fault label of the target fault object, a preset first model is trained into a decision model. The probability of a real fault of the fault object is determined by the decision model, which can prevent misjudgment and improve the accuracy of fault monitoring.
[0098] Figure 3 : is a flowchart of a fault monitoring method provided by an embodiment of the present disclosure. The method can be exemplarily performed by a fault monitoring device. In some embodiments, the fault monitoring device can be integrated as a functional module on a node (such as a management node, but not limited to a management node) of the training cluster. Or in other embodiments, the fault monitoring device can also be specifically configured as a separate hardware device in the training cluster, and the hardware device is used to monitor faults in the training cluster. Figure 3 As shown, in some implementations, the fault monitoring method provided by the embodiment of the present disclosure may include steps 301 to 304.
[0099] Step 301: Obtain a first log of a node included in a training cluster, a second log of software, and a training log of a training process.
[0100] The training cluster may include at least the following nodes: computing nodes (e.g., graphics processing unit (GPU) computing nodes, but not limited to GPU computing nodes), storage nodes, switch nodes, and management nodes. The computing nodes are used to execute training tasks. The storage nodes store the data involved in the large model training process. The management nodes are used to manage the nodes in the training cluster and to schedule and allocate training tasks.
[0101] The software in the training cluster includes but is not limited to task scheduling software. The task scheduling software is used to schedule and allocate large model training tasks in the training cluster. The task scheduling software can be installed in the management node of the training cluster. The computing nodes of the training cluster are equipped with an agent of the task scheduling software to record the tasks assigned to the computing nodes and the task execution status.
[0102] A training process refers to a process in a training cluster that is used to execute specific training tasks.
[0103] The naming of the first log and the second log in the embodiment of the present disclosure is only used to distinguish the log of the node from the diary of the software and has no other meaning.
[0104] The first log of a node contains information about the tasks performed by the node, fault information of the node (such as fault name, fault content, fault time, etc.), query operations performed after the node fails, and the state information of the node at each time. The query operation includes the query object and the query command. The state information of the node at each time includes information such as resource utilization, processor power, and network card traffic.
[0105] The second log of the software includes information on tasks executed by the software (such as information on task scheduling, but not limited to information on task scheduling), software failure information, and query operations executed after a software failure.
[0106] The training log of the training process includes information about the training tasks executed by the training process, fault information, and query operations performed after a training process fault occurs.
[0107] In some implementations, the first log of the nodes in the training cluster, the second log of the software, and the training log of the training process can be stored in a preset log system. The disclosed embodiment can obtain the first log, the second log, and the training log of the training cluster from the log system at preset intervals.
[0108] Step 302: Use the first log, the second log, and the training log as inputs of a preset fault identification model, and obtain fault information of the training cluster based on the fault identification model, where the fault information includes information of the fault object.
[0109] In some implementations, the fault identification model referred to in the embodiments of the present disclosure may be a fault identification model using the above Figure 1 The model obtained by training the model training method of the embodiment.
[0110] In the disclosed embodiment, the input of the fault identification model includes the first log, the second log and the training log, and the output includes the fault information of the training cluster. The fault information includes at least the information of the fault object (e.g., the node, software or process where the fault occurs). In some implementations, it may also include the fault name, fault content, fault time and other information.
[0111] It should be noted that the fault information of one or more fault objects can be identified based on the first log, the second log and the training log. For any fault object, the fault information of different faults occurring at the same time and / or the fault information of one or more identical and / or different faults occurring at different times can be identified based on the first log, the second log and the training log. For example, if the fault object has fault a and fault b at the first time, and fault b and fault c at the second time, then the fault information that can be identified based on the first log, the second log and the training log includes the fault information of fault a and fault b at the first time, and the fault information of fault b and fault c at the second time.
[0112] Step 303: for any fault object, determine the fault probability of the fault object based on the fault information of the fault object within a preset time window.
[0113] In practice, a faulty object may have multiple faults within a preset time window, and these faults may include one or more identical faults or multiple different faults.
[0114] In some implementations, statistics can be collected for the same fault of the fault object within a preset time window to determine the frequency of the same fault. For example, within a 15-minute time window, the fault object has a certain fault in the first minute, the second minute, and the third minute, respectively. The frequency of occurrence of the fault can be calculated as 3 divided by 15, which equals 20%. If only one fault occurs to the fault object within the preset time window, the frequency of occurrence of the fault is used as the fault probability of the fault object. If multiple faults occur to the fault object within the preset time window, the frequency of occurrence of the fault with the highest frequency can be used as the fault probability of the fault object.
[0115] In other embodiments, the fault information of the fault object within a preset time window can also be input into a preset decision model, and the failure probability of the fault object can be output through the decision model. Figure 2 The model obtained by training the method of the embodiment.
[0116] Step 304: In response to the failure probability of the fault object being higher than a first preset threshold, determine that the fault object has failed.
[0117] If the fault probability of the fault object is lower than the second preset threshold, it can be determined that the fault object has not failed. The second preset threshold is lower than the first preset threshold. The values of the second preset threshold and the first preset threshold can be set as needed.
[0118] If the failure probability of the fault object is greater than the second preset threshold and less than the first preset threshold, it is determined that the fault object is suspected to have failed.
[0119] In the disclosed embodiment, a first log of a node included in a training cluster, a second log of software, and a training log of a training process are obtained, and the first log, the second log, and the training log are used as inputs of a preset fault identification model. The fault information of the training cluster is obtained through identification of the fault identification model. Compared with the string matching method of the related art, more comprehensive fault monitoring can be achieved, the fault monitoring capability can be improved, and missed faults can be prevented. For any fault object, the fault probability of the fault object is determined based on the fault information of the fault object within a preset time window, and when the fault probability of the fault object is higher than a first preset threshold, it is determined that the fault object has a fault, which can prevent false detection and improve the accuracy of fault monitoring.
[0120] Figure 4 is a flow chart of a method for determining the failure probability of a fault object provided by an embodiment of the present disclosure. Figure 4 As shown, in some implementations, the failure probability of the faulty object can be determined by the following methods of step 401 and step 402.
[0121] Step 401: for at least one fault occurring in a fault object within a preset time window, filter out fault information of a target fault from the fault information of the at least one fault, where the target fault refers to a fault occurring with a frequency less than a preset frequency within the preset time window.
[0122] In some implementations, the disclosed embodiments can perform time domain average filtering on the faults that occur in the fault object within a preset time window, and filter out the faults whose occurrence frequency is less than the preset frequency from the fault information of the fault object. For example, the fault object has a card drop fault and a disk space shortage fault within the preset time window. If the card drop fault has a frequency greater than the preset frequency within the preset time window, and the disk space shortage fault has a frequency less than the preset frequency within the preset time window, then the fault information of the disk space shortage fault is filtered out.
[0123] The time domain average filtering process can filter out occasional false alarms for the fault object and improve the accuracy of fault monitoring.
[0124] Step 402: input the remaining fault information of the fault object within the preset time window into a preset decision model, and determine the fault probability of the fault object based on the decision model.
[0125] The fault information may include, for example, the fault object, fault name, fault content, fault time and other information.
[0126] In some implementations, the remaining fault information of the faulty object within a preset time window may be directly input into a preset decision model, and the fault probability of the faulty object may be determined based on the decision model.
[0127] In other embodiments, the remaining fault information of the faulty object within a preset time window and the status information of the faulty object at each moment within the preset time window, such as processor utilization, network card traffic, resource utilization, etc., may be input into a decision model, and the failure probability of the faulty object may be output through the decision model.
[0128] Among them, before the remaining fault information of the fault object within the preset time window and the state information of the fault object at each moment within the preset time window are input into the decision model, the fault information and the state information can also be encoded by a preset encoding method (for example, one-hot, but not limited to the one-hot encoding method), and then the encoded information is input into the decision model. For example, the preset time window is 15 minutes. In some embodiments, the fault information (if a fault occurs, it is encoded into a preset vector) and the state information of the fault object in each minute can be encoded respectively, and the fault information and the state information of each minute can be encoded into a vector to form 15 vectors, and then the 15 vectors are input into the decision model to determine the failure probability of the fault object.
[0129] The disclosed embodiment performs time-domain average filtering on the fault information of the fault object within a preset time window, and inputs the remaining fault information after filtering and the state information of the fault object at each moment into a decision model to determine the fault probability of the fault object, thereby improving the accuracy of determining the fault probability.
[0130] Figure 5 is a flow chart of an information query method provided by an embodiment of the present disclosure. Figure 5 As shown, in some implementations, when it is determined that the failure probability of the fault object is greater than the second preset threshold and less than the first preset threshold, the embodiment of the present disclosure may further include steps 501 to 503.
[0131] Step 501: input fault information of a fault object within a preset time window into a fault identification model, and determine a query object and a query command based on the fault identification model.
[0132] The input of the fault information of the fault object within the preset time window into the fault identification model can be understood as inputting all the fault information of the fault object within the preset time window into the fault identification model. Alternatively, it can also be understood as inputting the remaining fault information after the time domain average filtering process into the fault identification model. The fault identification model in this embodiment can be understood as the above Figure 1 The model obtained by training the method of the embodiment.
[0133] It should be noted that for different faults, the query object and query command output by the fault identification model are different. For example, for a fault in which the storage node disk space is insufficient, the query object output by the fault identification model may be the storage node, and the query command may be a command for querying the remaining disk space. For another example, for a fault in which the computing node is inaccessible, the query object output by the fault identification model may be a switch connected to the computing node, and the query command may be a command for querying the working status of the switch. Of course, this is only an example and not a sole limitation.
[0134] Step 502: query information from the query object based on the query command, and the queried information is used to assist in determining whether a fault occurs in the fault object.
[0135] For example, for a failure of insufficient disk space on a storage node, a query request can be sent to the storage node based on a query command, and the information obtained from the query is the remaining storage space of the storage node. For another example, for a failure of inaccessibility of a computing node, a query command can be sent to the corresponding switch, and the information obtained from the query command is the working status of the switch. Of course, this is only an example and not a sole limitation.
[0136] Step 503: Store the queried information in the log of the fault object.
[0137] Continuing with the above example, for the above storage node disk space failure, the remaining space information of the storage node that can be queried can be written into the log of the storage node. For the above computing node inaccessible failure, the working status of the above switch can be written into the log of the computing node. Of course, this is only an example and not a sole limitation.
[0138] By inputting the fault information of the fault object within a preset time window into the fault identification model, determining the query object and query command based on the fault identification model, querying information from the query object based on the query command, and storing the queried information in the log of the fault object, it can assist in determining whether the fault object is actually faulty, thereby improving the accuracy of fault monitoring.
[0139] Figure 6 is a structural diagram of a fault monitoring device provided by an embodiment of the present disclosure, such as Figure 6 As shown, the fault monitoring device 60 may include:
[0140] The log collection module 61 is used to obtain the first log of the node included in the training cluster, the second log of the software, and the training log of the training process;
[0141] A fault identification module 62, configured to use the first log, the second log and the training log as inputs of a preset fault identification model, and obtain fault information of the training cluster based on the fault identification model, wherein the fault information includes information of a fault object;
[0142] A fault decision module 63 is used to determine the fault probability of any fault object based on the fault information of the fault object within a preset time window;
[0143] The fault decision module 63 is further configured to determine that a fault occurs to the fault object in response to the fault probability of the fault object being higher than a first preset threshold.
[0144] Optionally, the fault decision module 63 is used to:
[0145] For at least one fault occurring in the fault object within the preset time window, fault information of a target fault is filtered out from the fault information of the at least one fault, wherein the target fault refers to a fault occurring at a frequency less than a preset frequency within the preset time window.
[0146] Optionally, the fault decision module 63 is used to:
[0147] The fault information of the fault object within a preset time window is input into a preset decision model, and the fault probability of the fault object is determined based on the decision model.
[0148] Optionally, the fault decision module is used to: determine that the fault object has not failed in response to the fault probability of the fault object being lower than a second preset threshold, and the second preset threshold is smaller than the first preset threshold.
[0149] Optionally, the fault monitoring device may further include:
[0150] a supplementary query module, configured to: in response to a fault probability of the fault object being greater than the second preset threshold and less than the first preset threshold, input the fault information of the fault object within the preset time window into the fault identification model, and determine a query object and a query command based on the fault identification model;
[0151] A query execution module is used to query information from the query object based on the query command, the queried information is used to assist in determining whether the fault object has a fault; and store the queried information in the log of the fault object.
[0152] Optionally, the fault monitoring device may further include:
[0153] The first training module is used to:
[0154] Acquire a training sample, wherein the training sample includes fault information of a target fault object within the preset time window and a fault label of the target fault object, wherein the target fault object is any node, software or training process in any training cluster;
[0155] Based on the fault information and fault label of the target fault object, a preset first model is trained into the decision model.
[0156] Optionally, the fault monitoring device may further include:
[0157] The second training module is used to:
[0158] Training a preset second model based on an instruction manual of query commands in a training cluster, technical information related to the training cluster, and log information, so that the second model learns knowledge related to the training cluster, wherein the instruction manual is used to explain the query commands in the training cluster;
[0159] Based on the sample logs and training labels of at least one training cluster, training the second model that has learned the knowledge to obtain the fault identification model;
[0160] The sample log of the at least one training cluster includes logs of nodes, software and training processes in the at least one training cluster;
[0161] The training label includes real fault information of the at least one training cluster and a query operation performed when the real fault information is identified, wherein the query operation includes a query object and a query command.
[0162] The fault monitoring device provided in the embodiment of the present disclosure can execute the method of any of the above method embodiments, and its execution method and beneficial effects are similar, which will not be repeated here.
[0163] It should also be noted that the division of modules in the above-mentioned fault monitoring device in the embodiment of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0164] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present disclosure.
[0165] Figure 7 The schematic diagram of the structure of the computer device embodiment provided by the present disclosure embodiment. Figure 7 As shown, the computer device includes a memory 121 and a processor 122 .
[0166] The memory 121 is used to store programs. In addition to the above programs, the memory 121 can also be configured to store various other data to support operations on the computer device. Examples of such data include instructions for any application or method operating on the computer device, contact member data, phone book member data, messages, pictures, videos, etc.
[0167] The memory 121 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0168] In some implementations, the processor 122 is coupled to the memory 121 to execute a program stored in the memory 121 to perform the method of any of the above method embodiments.
[0169] Further, if Figure 7 As shown, the computer device may also include: a communication component 123, a power component 124, an audio component 125, a display 126 and other components. Figure 7 Only some components are shown schematically, and it does not mean that the computer equipment only includes Figure 7 Components shown.
[0170] The communication component 123 is configured to facilitate wired or wireless communication between the computer device and other devices. The computer device can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 123 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 123 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared member data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0171] The power supply component 124 provides power to various components of the computer device. The power supply component 124 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the computer device.
[0172] The audio component 125 is configured to output and / or input audio signals. For example, the audio component 125 includes a microphone (MIC), and when the computer device is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 121 or sent via the communication component 123. In some embodiments, the audio component 125 also includes a speaker for outputting audio signals.
[0173] The display 126 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0174] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the method described in any of the above method embodiments.
[0175] In the embodiments of the present disclosure, the above-mentioned computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO)), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drives (SSDs)), etc.
[0176] The embodiment of the present disclosure also provides a service testing system, which includes a first device and a second device. The first device can execute the above Figure 1-Figure 2 The method of any embodiment of the present invention, the second device may execute the above Figure 3 The method of the embodiment has similar execution mode and beneficial effects, which will not be described in detail here.
[0177] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0178] An embodiment of the present disclosure provides a computer program product, including computer program instructions, which can implement the method described in any of the above method embodiments when executed by a processor.
[0179] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0180] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A fault monitoring method, characterized in that: The method comprises: Obtaining a first log of a node included in the training cluster, a second log of the software, and a training log of the training process; Using the first log, the second log, and the training log as inputs of a preset fault identification model, and obtaining fault information of the training cluster based on the fault identification model, wherein the fault information includes information of a fault object; For any fault object, determining the fault probability of the fault object based on the fault information of the fault object within a preset time window; In response to the failure probability of the fault object being higher than a first preset threshold, it is determined that the fault object fails.
2. The method according to claim 1, characterized in that The determining, for any fault object, the fault probability of the fault object based on the fault information of the fault object within a preset time window includes: For at least one fault occurring in the fault object within the preset time window, fault information of a target fault is filtered out from the fault information of the at least one fault, wherein the target fault refers to a fault occurring at a frequency less than a preset frequency within the preset time window.
3. The method according to claim 1, characterized in that The determining, for any fault object, the fault probability of the fault object based on the fault information of the fault object within a preset time window includes: The fault information of the fault object within a preset time window is input into a preset decision model, and the fault probability of the fault object is determined based on the decision model.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: In response to the failure probability of the fault object being lower than a second preset threshold, it is determined that the fault object has not failed, and the second preset threshold is lower than the first preset threshold.
5. The method according to claim 4, characterized in that The method further comprises: In response to the fault probability of the fault object being greater than the second preset threshold and less than the first preset threshold, inputting the fault information of the fault object within the preset time window into the fault identification model, and determining the query object and the query command based on the fault identification model; Querying information from the query object based on the query command, where the queried information is used to assist in determining whether the fault object has a fault; The queried information is stored in the log of the fault object.
6. The method according to claim 3, characterized in that: The method further comprises: Acquire a training sample, wherein the training sample includes fault information of a target fault object within the preset time window and a fault label of the target fault object, wherein the target fault object is any node, software or training process in any training cluster; Based on the fault information and fault label of the target fault object, a preset first model is trained into the decision model.
7. The method according to claim 1, characterized in that The method further comprises: Training a preset second model based on an instruction manual of query commands in a training cluster, technical information related to the training cluster, and log information, so that the second model learns knowledge related to the training cluster, wherein the instruction manual is used to explain and illustrate the query commands in the training cluster; Based on the sample logs and training labels of at least one training cluster, training the second model that has learned the knowledge to obtain the fault identification model; The sample log of the at least one training cluster includes logs of nodes, software and training processes in the at least one training cluster; The training label includes real fault information of the at least one training cluster and a query operation performed when the real fault information is identified, wherein the query operation includes a query object and a query command.
8. A fault monitoring device, characterized in that: include: A log collection module, used to obtain a first log of a node included in the training cluster, a second log of the software, and a training log of the training process; a fault identification module, configured to use the first log, the second log, and the training log as inputs of a preset fault identification model, and obtain fault information of the training cluster based on the fault identification model, wherein the fault information includes information of a fault object; A fault decision module, for determining, for any fault object, the fault probability of the fault object based on the fault information of the fault object within a preset time window; The fault decision module is further configured to determine that a fault occurs to the fault object in response to a fault probability of the fault object being higher than a first preset threshold.
9. A computer device, characterized in that: The computer device comprises: Memory; Processor; and Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.