Fault detection method and device of server component, storage medium and electronic equipment

By training the target recognition model to identify the association relationship of server component logs, the problem of low server component failure detection efficiency is solved, and the effect of quickly and accurately locate the source of the fault is achieved.

CN120336056APending Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510398402.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, server component fault detection efficiency is low, and it is difficult to quickly and accurately locate the cause of the fault, mainly because the log system lacks globality and accuracy, and manual log screening is time-consuming and labor-intensive.

Method used

The initial recognition model is trained using log samples with entity feature labels and semantic feature labels to generate a target recognition model. Through the association relationship between entity feature and semantic feature recognition logs, the association logs are selected to locate the source of the fault.

Benefits of technology

It improves the efficiency and accuracy of server component failure detection, reduces the tedious process of manually screening logs, and quickly and accurately locates the source of the fault.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336056A_ABST
    Figure CN120336056A_ABST
Patent Text Reader

Abstract

The invention discloses a fault detection method and device for a server component, a storage medium and electronic equipment, and relates to the field of computers.The fault detection method comprises the steps that a log sample with an entity feature tag and a semantic feature tag is used for training a model, therefore, the model can identify the association relationship of the component parameters of the server components of the input log record and the association relationship in the log semantics. When a server fails, entity features and semantic features of a target log and logs currently recorded by the server are recognized through a model, associated parameters are generated, then associated logs related to the target log are screened out from the logs currently recorded by the server, the association degree between the associated logs and the target log is analyzed, a fault source is accurately positioned, and the fault diagnosis accuracy is improved. It is not needed to manually check the logs one by one from a log system to screen and analyze the associated logs, the technical problem that the fault detection efficiency of the server components is low is solved, and the technical effect of improving the fault detection efficiency of the server components is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computers, and more particularly, to a method and device for detecting faults of server components, a storage medium, and an electronic device. Background Art

[0002] When a server fails, it is necessary to analyze the cause of the failure in a timely manner and process the faulty components in a timely manner to reduce the impact of the failure on the overall operation of the server. In the related art, operation and maintenance personnel need to view the logs one by one in the log system, or search for relevant log information through keywords, and then analyze the selected logs one by one to locate the root cause of the failure. However, with the growth of the log data volume, the method of manually screening logs is not only time-consuming but also inefficient. On the one hand, the functional collaboration in the server system is complex, and a function usually requires multiple server components to cooperate to complete. When analyzing the logs, it is necessary to query the log information of multiple related server components. However, the log system usually collects the log information of each server component separately, ignoring the connection between the server components, resulting in the lack of globality and accuracy of the analysis results. On the other hand, the log system contains a large amount of redundant information, which further increases the difficulty of screening useful logs. Therefore, it is difficult for operation and maintenance personnel to quickly find the log information related to the failure from the log system, and it is even more difficult to quickly locate the location where the failure occurs.

[0003] In view of the technical problems such as low efficiency of fault detection of server components in the related art, no effective solution has been proposed yet. Summary of the Invention

[0004] The embodiments of the present application provide a method and device for detecting faults of server components, a storage medium, and an electronic device, so as to at least solve the technical problems such as low efficiency of fault detection of server components in the related art.

[0005] According to an embodiment of the present application, a method for detecting faults of server components is provided, and the method includes:

[0006] Inputting a reference log and a target log into a target recognition model, and obtaining entity features and semantic features output by the target recognition model, where the target log is a log indicating that a target server component is in a target fault state, the reference log is the log currently recorded by the server, the entity features are used to characterize the association relationship between the reference log and the target log in terms of component parameters of the recorded server components, the semantic features are used to characterize the association relationship between the reference log and the target log in terms of log semantics, and the target recognition model is obtained by training an initial recognition model using a log sample group labeled with entity feature labels and semantic feature labels;

[0007] Generate the association parameters of the reference log according to the entity features and semantic features, where the association parameters are used to indicate the degree of association between the reference log and the target log;

[0008] Filter out the associated log corresponding to the target log from the reference log according to the association parameters;

[0009] Detect the cause of the failure of the target server component in the target failure state according to the associated log.

[0010] Optionally, the target recognition model includes: an entity recognition model and a semantic recognition model. Input the reference log and the target log into the target recognition model, and obtain the entity features and semantic features output by the target recognition model, including:

[0011] Input the reference log and the target log into the entity recognition model to obtain the entity features output by the entity recognition model;

[0012] Input the reference log and the target log into the semantic recognition model to obtain the semantic features output by the semantic recognition model.

[0013] Optionally, the entity recognition model includes: a convergence log partitioning layer, an entity extraction layer, an entity splicing layer, and an entity feature extraction layer. The convergence log partitioning layer is connected to the entity extraction layer, and the entity splicing layer is connected between the entity extraction layer and the entity feature extraction layer. Input the reference log and the target log into the entity recognition model to obtain the entity features output by the entity recognition model, including:

[0014] Call the convergence log partitioning layer to partition the reference log into multiple segments of reference log information, and partition the target log into multiple segments of target log information, and transmit the reference log information and the target log information to the entity extraction layer, where each segment of log information is labeled with the information type;

[0015] Call the entity extraction layer to extract the reference log entity corresponding to the reference log from the multiple segments of reference log information, and extract the target log entity corresponding to the target log from the multiple segments of target log information, and input the reference log entity and the target log entity into the entity splicing layer, where the reference log entity is the information for recording the component parameters of the server component in the reference log, and the target log entity is the information for recording the component parameters of the server component in the target log;

[0016] Call the entity splicing layer to splice the reference log entity and the target log entity to obtain a spliced log entity, and input the spliced log entity into the entity feature extraction layer;

[0017] Call the entity feature extraction layer to extract features from the spliced log entity to obtain entity features.

[0018] Optionally, before calling the convergence log partitioning layer to partition the reference log into multiple segments of reference log information and the target log into multiple segments of target log information, the method further includes:

[0019] Input the log samples labeled with log information tags, information type tags, and information position tags into the initial log partitioning layer to obtain the predicted log information, predicted information type, and predicted information position output by the initial log partitioning layer. Among them, the log information tag is a lexical segment in the log sample, the information type tag is used to indicate the information type of the lexical segment, the information position tag is used to indicate the position of the lexical segment in the log sample, the predicted log information is obtained by the initial log partitioning layer predicting the log information in the log sample, the predicted information type is obtained by the initial log partitioning layer identifying the information type of the lexical segment indicated by the predicted log information, and the predicted information position is obtained by the initial log partitioning layer identifying the position of the lexical segment indicated by the predicted log information in the log sample;

[0020] Generate a first loss value according to the predicted log information and the log information tag, generate a second loss value according to the predicted information type and the information type tag, and generate a third loss value according to the predicted information position and the information position tag;

[0021] Adjust the model parameters of the initial log partitioning layer according to the first loss value, the second loss value, the third loss value, and the preset convergence condition until the model parameters converge to obtain the convergence log partitioning layer.

[0022] Optionally, the semantic recognition model includes: a convergence semantic feature extraction layer, a similarity calculation layer, and a feature fusion layer. The similarity calculation layer is connected between the convergence semantic feature extraction layer and the feature fusion layer. Input the reference log and the target log into the semantic recognition model to obtain the semantic features output by the semantic recognition model, including:

[0023] Call the convergence semantic feature extraction layer to extract reference semantic features from the reference log and target semantic features from the target log, and transmit the reference semantic features and the target semantic features to the similarity calculation layer;

[0024] Call the similarity calculation layer to calculate the feature similarity between the reference semantic features and the target semantic features, and transmit the reference semantic features, the target semantic features, and the feature similarity to the feature fusion layer. Among them, the correlation relationship between the reference log and the target log in terms of log semantics is positively correlated with the feature similarity;

[0025] Call the feature fusion layer to perform feature fusion on the reference semantic features and the target semantic features according to the feature similarity to obtain the semantic features.

[0026] Optionally, before calling the convergence semantic feature extraction layer to extract reference semantic features from the reference log and target semantic features from the target log, the method further includes:

[0027] Obtain log sample pairs labeled with feature similarity labels, where a log sample pair includes two log samples, and the feature similarity label is used to indicate the strength of the association relationship in log semantics between the two log samples included in the log sample pair;

[0028] Input the two log samples of the log sample pair into the initial semantic feature extraction layer to obtain two sample semantic features corresponding to the two log samples output by the initial semantic feature extraction layer;

[0029] Generate the actual feature similarity of the two sample semantic features;

[0030] Adjust the model parameters of the initial semantic feature extraction layer according to the feature similarity label and the actual feature similarity to obtain the convergence semantic feature extraction layer.

[0031] Optionally, generating the association parameter of the reference log according to the entity feature and the semantic feature includes:

[0032] Concatenate the entity feature and the semantic feature to obtain a fusion feature, where the fusion feature is used to characterize the association relationship between the reference log and the target log in the component parameters of the server component recorded, and the association relationship between the reference log and the target log in log semantics;

[0033] Input the fusion feature into the target classification model to obtain the association parameter output by the target classification model, where the target classification model is obtained by training the initial classification model with fusion feature samples labeled with association parameter labels, the fusion feature sample is obtained by concatenating the entity feature sample and the semantic feature sample, and the association parameter label is used to indicate the association parameter between the reference log and the target log corresponding to the entity feature sample and the semantic feature sample corresponding to the fusion feature sample.

[0034] According to another embodiment of the embodiments of the present application, there is also provided a fault detection device for a server component, including:

[0035] A first input module, configured to input a reference log and a target log into a target recognition model, and obtain entity features and semantic features output by the target recognition model. The target log is a log indicating that a target server component is in a target failure state, and the reference log is the log currently recorded by the server. The entity features are used to characterize the association relationship between the reference log and the target log in terms of the component parameters of the server components recorded in the logs, and the semantic features are used to characterize the association relationship between the reference log and the target log in terms of the log semantics. The target recognition model is obtained by training an initial recognition model using a log sample group labeled with entity feature labels and semantic feature labels;

[0036] A first generation module, configured to generate association parameters of the reference log according to the entity features and the semantic features, where the association parameters are used to indicate the degree of association between the reference log and the target log;

[0037] A screening module, configured to screen out the associated log corresponding to the target log from the reference logs according to the association parameters;

[0038] A detection module, configured to detect the cause of the failure that the target server component is in the target failure state according to the associated logs.

[0039] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above-mentioned server component failure detection methods when executing the computer program.

[0040] Through the present application, since the initial recognition model is trained using log samples with entity feature labels and semantic feature labels, a target recognition model is obtained. This enables the target recognition model to recognize the association relationship between the component parameters of the server components recorded in the input logs and the association relationship in the log semantics. Therefore, when an exception occurs in the server, through the target recognition model, the association relationship between the target log and the log currently recorded by the server in terms of the component parameters of the server components recorded in the logs and the association relationship in the log semantics can be recognized, and then the entity features and semantic features can be obtained, and the association parameters are generated through these features. Then, according to the association parameters, the logs related to the target log are screened out from the logs currently recorded by the server, that is, the associated logs. Finally, by analyzing the degree of association between the associated logs and the target log, the fault source is accurately located, avoiding the cumbersome process of manually viewing logs one by one in the log system to screen and analyze the log information related to specific situations, significantly improving the efficiency of screening associated logs from the log system, and improving the accuracy of locating the fault source. Therefore, the technical problems of low efficiency in detecting server component failures in the related art can be solved, and the technical effect of improving the efficiency of detecting server component failures is achieved. Description of the Drawings

[0041] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] Figure 1 is a hardware structure block diagram of a computer device for a method of detecting faults in server components according to an embodiment of the present application;

[0043] Figure 2 is a flowchart of a method of detecting faults in server components according to an embodiment of the present application;

[0044] Figure 3 is a structural diagram of an entity recognition model according to an embodiment of the present application;

[0045] Figure 4 is a structural diagram of a convergence log partitioning layer according to an embodiment of the present application;

[0046] Figure 5 is a structural diagram of a semantic recognition model according to an embodiment of the present application;

[0047] Figure 6 is a schematic diagram of a process for performing fault detection using a method of detecting faults in server components according to an embodiment of the present application;

[0048] Figure 7 is a flowchart of extracting semantic features from log pairs according to an embodiment of the present application;

[0049] Figure 8 is a structural block diagram of a device for detecting faults in server components according to an embodiment of the present application;

[0050] Figure 9 is a schematic diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0052] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second" etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0053] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0054] The method embodiments provided in the embodiments of this application can be executed on a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is a hardware structure block diagram of a computer device for a method of detecting faults in a server component according to an embodiment of this application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a configuration different from Figure 1 shown in the figure.

[0055] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the method of detecting faults in a server component in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0056] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0057] The nouns involved in the embodiments of the present application are explained as follows:

[0058] BERT: Bidirectional Encoder Representations from Transformers, a pre-trained natural language processing model based on Transformers;

[0059] BMC: Baseboard Management Controller, a server management unit;

[0060] CPU: Central Processing Unit;

[0061] NER: Named Entity Recognition;

[0062] ResNet: Residual Neural Network.

[0063] In this embodiment, a method for detecting faults in server components is provided. Figure 2 is a flowchart of a method for detecting faults in server components according to an embodiment of the present application, as Figure 2 shown, and the process includes the following steps:

[0064] Step S12: Input the reference log and the target log into the target recognition model to obtain the entity features and semantic features output by the target recognition model. Among them, the target log is a log indicating that the target server component is in a target fault state, the reference log is the log currently recorded by the server, the entity features are used to characterize the correlation relationship between the reference log and the target log in terms of the component parameters of the recorded server components, the semantic features are used to characterize the correlation relationship between the reference log and the target log in terms of the log semantics, and the target recognition model is obtained by training an initial recognition model using a log sample group labeled with entity feature labels and semantic feature labels.

[0065] Optionally, in this embodiment, it is assumed that the server has an abnormal situation of frequent restart during operation, and the cause of the abnormality needs to be found. First, obtain the target log. The target log can be, but is not limited to, the abnormal information carried in the fault signal output when the server has an abnormality, which records the time of the abnormality, the server status, events, etc. The fault signal can be the error code output by the server. Then, collect the reference log from the server log system. The reference log can be, but is not limited to, all the logs included in the current server log system. The reference log records various events, operations, and status changes that occur in each module during the operation of the server, including information related to the abnormality, which can help analyze the operation of the server in the past period of time. Therefore, the cause of the abnormality can be determined by viewing the information related to the abnormality in the reference log. For example, if it is found in the reference log that the CPU temperature is too high and the cooling fan fails, it can be analyzed that it may be the failure of the cooling fan that causes the heat to not be dissipated in time, resulting in the continuous rise of the CPU temperature, and finally triggering the overheat protection mechanism to forcibly cut off the power supply and restart to prevent the CPU from being damaged due to long-term high-temperature operation. After the above analysis, the root cause of the frequent restart abnormality of the server can be located as the failure of the cooling fan.

[0066] Optionally, in this embodiment, the target recognition model can be, but is not limited to, an AI (Artificial Intelligence) model that has the ability to analyze and process log text information after pre-training and fine-tuning, such as a pre-trained language model that has the ability to recognize key entities included in the input log and the correlation relationship of the input log in terms of log semantics. The initial recognition model can be, but is not limited to, an AI model that has not been trained and does not have the ability to analyze and process log text information.

[0067] Optionally, in this embodiment, the entity feature label may, but is not limited to, be the true entity feature of the log sample group, such as the core vocabulary or phrases describing the server status (such as descriptions related to CPU temperature, memory status, disk status, or network status, etc.), events (such as descriptions related to error events, warning events, or exception events, etc.), or operations (such as descriptions related to startup operations, stop operations, or configuration operations, etc.) in the log, such as "CPU temperature too high", "power failure", etc. The semantic feature label may, but is not limited to, be the true semantic feature of the log sample group, such as the logical relationship and deep semantic information between logs, such as the association relationship between "CPU temperature too high" and "abnormal fan speed". The training process of the pre-trained language model may include two stages: pre-training and fine-tuning. In the pre-training stage, the model is pre-trained using a large amount of general log data to learn general feature representations. In the fine-tuning stage, the model is fine-tuned using task-specific log data to adjust the model parameters to adapt to the specific task.

[0068] Optionally, in this embodiment, the training process of obtaining the target recognition model may, but is not limited to, include: First, input the log sample group labeled with entity feature labels and semantic feature labels into the initial recognition model, and the initial recognition model outputs the predicted entity features and predicted semantic features of the log sample group based on the initial model parameters. Then, aiming to narrow the difference between the predicted features (predicted entity features and predicted entity features) and the true features (predicted semantic features and predicted semantic features), continuously adjust the initial model parameters until the model loss function converges or the training iteration times are completed. At this time, the initial recognition model has learned the ability to recognize entity features and semantic features in the log sample group through training, and the target recognition model is obtained.

[0069] Optionally, in this embodiment, after obtaining the reference log and the target log, input the reference log and the target log into the target recognition model. After analyzing and processing the text information of the reference log and the target log, the target recognition model can recognize the key entities contained in the reference log and the target log and output entity features; at the same time, the target recognition model can also recognize the association relationship in the log semantics between the reference log and the target log and output semantic features.

[0070] Step S14, generate the association parameter of the reference log according to the entity feature and the semantic feature, where the association parameter is used to indicate the degree of association between the reference log and the target log;

[0071] Optionally, in this embodiment, since the entity feature reflects the specific hardware status, event, or operation mentioned in the log, and the semantic feature reflects the logical connection and deep semantic information between logs. Therefore, by combining the entity feature and the semantic feature, the degree of association between the reference log and the target log can be analyzed more comprehensively.

[0072] Optionally, in this embodiment, the correlation parameter may, but is not limited to, be a quantization index calculated based on entity features and semantic features, and is used to measure the correlation degree between the reference log and the target log. The larger the value of the correlation parameter, the stronger the correlation degree between the reference log corresponding to the correlation parameter and the target log.

[0073] Step S16: Screen out the correlation logs corresponding to the target log from the reference logs according to the correlation parameter;

[0074] Optionally, in this embodiment, the method of screening correlation logs from the reference logs according to the correlation parameter may, but is not limited to, include: determining the reference logs corresponding to the correlation parameters that meet the preset conditions as correlation logs according to a preset threshold or ranking rule, that is, the logs with a higher correlation degree with the target log.

[0075] Step S18: Detect the cause of the failure of the target server component in the target failure state according to the correlation logs.

[0076] Optionally, in this embodiment, the method of detecting the cause of the failure may, but is not limited to, include: comprehensively analyzing the information in the correlation logs to determine the specific cause of the failure of the target server component in the target failure state. For example, if "CPU temperature is too high" and "fan speed is abnormal" are mentioned in multiple correlation logs, it can be inferred that the heat dissipation module may be the root cause of the abnormality.

[0077] Through this application, since the initial recognition model is trained using log samples with entity feature tags and semantic feature tags, a target recognition model is obtained. This enables the target recognition model to identify the association relationship of the component parameters of the server components recorded in the input log, as well as the association relationship in the log semantics. Therefore, when the server has an abnormality, the target recognition model can identify the association relationship of the component parameters of the server components recorded in the target log and the log currently recorded by the server, as well as the association relationship in the log semantics, and then obtain entity features and semantic features, and generate a correlation parameter through these features. Then, according to the correlation parameter, the logs related to the target log are screened out from the logs currently recorded by the server, that is, the correlation logs. Finally, by analyzing the correlation degree between the correlation logs and the target log, the fault source is accurately located, thereby achieving the technical effect of improving the fault detection efficiency of server components, and further solving the technical problem of low fault detection efficiency of server components.

[0078] As an optional solution, the target recognition model includes: an entity recognition model and a semantic recognition model. Inputting the reference log and the target log into the target recognition model, the target recognition model outputs entity features and semantic features, including:

[0079] S21. Input the reference log and the target log into the entity recognition model to obtain the entity features output by the entity recognition model;

[0080] S22. Input the reference log and the target log into the semantic recognition model to obtain the semantic features output by the semantic recognition model.

[0081] Optionally, in this embodiment, the target recognition model may include, but is not limited to, two pre-trained language models: an entity recognition model and a semantic recognition model. The entity recognition model may be used to recognize, but is not limited to, the key entities contained in the input log, such as the core words or phrases in the log that describe the server status, events, or operations. The semantic recognition model may be used to recognize, but is not limited to, the association relationships in the semantics of the input log, such as the logical relationships and deep semantic information between the logs.

[0082] As an optional solution, the entity recognition model includes: a convergence log partitioning layer, an entity extraction layer, an entity splicing layer, and an entity feature extraction layer. The convergence log partitioning layer is connected to the entity extraction layer, and the entity splicing layer is connected between the entity extraction layer and the entity feature extraction layer. Input the reference log and the target log into the entity recognition model to obtain the entity features output by the entity recognition model, including:

[0083] S31. Call the convergence log partitioning layer to partition the reference log into multiple segments of reference log information, and partition the target log into multiple segments of target log information, and transmit the reference log information and the target log information to the entity extraction layer, where each segment of log information is labeled with the information type;

[0084] S32. Call the entity extraction layer to extract the reference log entities corresponding to the reference log from the multiple segments of reference log information, and extract the target log entities corresponding to the target log from the multiple segments of target log information, and input the reference log entities and the target log entities into the entity splicing layer, where the reference log entities are the information for recording the component parameters of the server components in the reference log, and the target log entities are the information for recording the component parameters of the server components in the target log;

[0085] S33. Call the entity splicing layer to splice the reference log entities and the target log entities to obtain the spliced log entities, and input the spliced log entities into the entity feature extraction layer;

[0086] S34. Call the entity feature extraction layer to extract features from the spliced log entities to obtain the entity features.

[0087] Optionally, in this embodiment, Figure 3 is a structural diagram of an entity recognition model according to an embodiment of the present application, as shown in Figure 3As shown in the figure, the entity recognition model includes: a convergence log partitioning layer, an entity extraction layer, an entity splicing layer, and an entity feature extraction layer. The convergence log partitioning layer can be, but is not limited to, a pre-trained language model that can recognize the log information in the reference log and the target log. For example, the already trained BERT model has the function of recognizing entities and entity types in the log text. The entity type can be, but is not limited to, the information type of the entity in the text. For example, the entity types in the log text include: Server Status, Event, Operation, Module Status, Time Information, System Status, etc. The format of the log information output by the BERT model can be a structured list or dictionary, containing information such as the entities and their categories recognized from the log text. For example, for the following log:

[0088] Log A: "2023-04-01 12:00:00, CPU temperature is too high, reaching 85°C, system status: warning";

[0089] The log information output by the BERT model may be as follows:

[0090] {"time_information": "2023-04-01 12:00:00";

[0091] "server_status": "CPU temperature is too high";

[0092] "event": "85°C";

[0093] "system_status": "warning";}.

[0094] Optionally, in this embodiment, Figure 4 is a structural diagram of a convergence log partitioning layer according to an embodiment of the present application. As Figure 4 shown, the BERT model (convergence log partitioning layer) can recognize the log information of the reference log, divide the reference log into multiple segments of reference log information according to the entity types in the log information, and the log information of the target log, divide the target log into multiple segments of target log information according to the entity types in the log information, and transmit the reference log information and the target log information to the entity extraction layer.

[0095] Optionally, in this embodiment, as Figure 3As shown, the entity extraction layer can be, but is not limited to, a module with a filtering function. According to the specific entity types required by the task, entities that meet the specific entity types are extracted from the input log information, such as server status and system status. After receiving the reference log information and the target log information transmitted by the BERT model (convergent log partitioning layer), the entity extraction layer can, according to the specific entity types required by the task, such as the entity type that records the component parameters of server components, extract the reference log entities that meet the specific entity types from the reference log information, extract the target log entities that meet the specific entity types from the target log information, and transmit the reference log entities and the target log entities to the entity splicing layer.

[0096] Optionally, in this embodiment, as Figure 3 shown, the entity splicing layer can be, but is not limited to, a module with an information splicing function, such as a pooling layer, which can splice the reference log entities extracted from the reference log and the target log entities extracted from the target log to obtain a spliced log entity containing the key entities in the reference log and the key entities in the target log, and transmit the spliced log entity to the entity feature extraction layer.

[0097] Optionally, in this embodiment, as Figure 3 shown, the entity feature extraction layer can be, but is not limited to, a module that extracts features from the spliced entity information. For example, methods such as the Bag of Words model and TF-IDF (Term Frequency-Inverse Document Frequency) can be used to extract features from the spliced log entity to obtain entity features containing the key entity features in the reference log and the key entity features in the target log.

[0098] As an optional solution, before calling the convergent log partitioning layer to partition the reference log into multiple segments of reference log information and partition the target log into multiple segments of target log information, the method further includes:

[0099] S41. Input the log samples labeled with log information tags, information type tags, and information location tags into the initial log partitioning layer to obtain the predicted log information, predicted information type, and predicted information location output by the initial log partitioning layer. Here, the log information tag is a lexical segment in the log sample, the information type tag is used to indicate the information type of the lexical segment, and the information location tag is used to indicate the position of the lexical segment in the log sample. The predicted log information is obtained by the initial log partitioning layer predicting the log information in the log sample. The predicted information type is identified by the initial log partitioning layer for the information type of the lexical segment indicated by the predicted log information. The predicted information location is identified by the initial log partitioning layer for the position of the lexical segment indicated by the predicted log information in the log sample.

[0100] S42. Generate a first loss value based on the predicted log information and the log information tag, generate a second loss value based on the predicted information type and the information type tag, and generate a third loss value based on the predicted information location and the information location tag.

[0101] S43. Adjust the model parameters of the initial log partitioning layer according to the first loss value, the second loss value, the third loss value, and the preset convergence condition until the model parameters converge to obtain a converged log partitioning layer.

[0102] Optionally, in this embodiment, before calling the BERT model (converged log partitioning layer) to identify entities and entity types in the log text, the BERT model needs to be trained to enable it to have the ability to identify entities and entity types in the log text.

[0103] Optionally, in this embodiment, the method for training the initial BERT model (initial converged log partitioning layer) may include, but is not limited to: First, preprocess the log samples labeled with log information tags, information type tags, and information location tags and convert them into a format acceptable to the initial BERT model. For example, use a tokenizer to tokenize the log text. Here, the log information tag can be the real entity (lexical segment) in the log sample, the information type tag can be the entity type corresponding to the lexical segment, and the information location tag can be the position of the lexical segment in the log sample, usually including the start position and end position of the lexical segment in the log sample. For example, the following is log A labeled with log information tags, information type tags, and information location tags:

[0104] Log A: "2023-04-01 12:00:00, CPU temperature is too high, reaching 85°C, system status: warning";

[0105] Log information tags: 2023-04-01 12:00:00, CPU temperature is too high, 85°C, warning;

[0106] Information type tags: time information, server status, events, system status;

[0107] Information location tags: [1,18] (location of "2023-04-01 12:00:00"), [19,26] (location of "CPU temperature is too high"), [29,33] (location of "85°C"), [46,48] (location of "Warning").

[0108] Then, the preprocessed data is input into the initial BERT model for training. The inputs of the initial BERT model include: input ID (vocabulary index corresponding to the tokenized log text), attention mask (used to indicate which positions are actual inputs and which are padding), entity label (log information and information type label), and location label (information location label). During the training process, the initial BERT model predicts the log information for the vocabulary segments in the log sample, and simultaneously obtains the information type (predicted information type) of the vocabulary segments indicated by the predicted log information and the location (predicted information location) of the vocabulary segments indicated by the predicted log information in the log sample. Then, calculate the loss between the predicted value and the true value, such as cross-entropy loss. Obtain the first loss value according to the predicted log information and the log information label, the second loss value according to the predicted information type and the information type label, and the third loss value according to the predicted information location and the information location label. Then, update the model parameters of the initial BERT model by backpropagation according to the first loss value, the second loss value, and the third loss value, and optimize the performance of the initial BERT model until the preset convergence condition is met, such as the convergence of model parameters, to obtain the trained BERT model, that is, the convergence log partitioning layer. The trained BERT model can be used to extract key entities and their types and locations from new logs.

[0109] As an optional solution, the semantic recognition model includes: a convergence semantic feature extraction layer, a similarity calculation layer, and a feature fusion layer. The similarity calculation layer is connected between the convergence semantic feature extraction layer and the feature fusion layer. Input the reference log and the target log into the semantic recognition model to obtain the semantic features output by the semantic recognition model, including:

[0110] S51, call the convergence semantic feature extraction layer to extract reference semantic features from the reference log and target semantic features from the target log, and transmit the reference semantic features and target semantic features to the similarity calculation layer;

[0111] S52, call the similarity calculation layer to calculate the feature similarity between the reference semantic features and the target semantic features, and transmit the reference semantic features, target semantic features, and feature similarity to the feature fusion layer. Among them, the correlation relationship between the reference log and the target log in terms of log semantics is positively correlated with the feature similarity;

[0112] S53, call the feature fusion layer to perform feature fusion on the reference semantic feature and the target semantic feature according to the feature similarity to obtain a semantic feature.

[0113] Optionally, in this embodiment, Figure 5 is a structural diagram of a semantic recognition model according to an embodiment of the present application, as Figure 5 shown, the semantic recognition model includes: a convergent semantic feature extraction layer, a similarity calculation layer, and a feature fusion layer. The convergent semantic feature extraction layer can be but is not limited to a pre-trained language model capable of recognizing the correlation relationship between the reference log and the target log in terms of log semantics, such as the pre-trained RoBERTa model, which has the function of recognizing the logical relationship and deep semantic information between logs. For example, for the following log pair:

[0114] Log A: "2023-04-01 12:00:00, CPU temperature is too high, reaching 85°C, system status: warning";

[0115] Log B: "2023-04-01 12:05:00, fan speed is abnormal, system status: warning".

[0116] The steps for the RoBERTa model to extract semantic features can be but are not limited to including:

[0117] Step 1: Split the input log text into words and add special tokens at the beginning and end of the log text. For example, add [CLS] at the beginning and [SEP] at the end to get:

[0118] Log A: [CLS]2023-04-01 12:10:00CPU temperature is too high, system status: warning[SEP];

[0119] Log B: [CLS]2023-04-01 12:05:00fan speed is abnormal, system status: warning[SEP].

[0120] Step 2: Embedding layer. It includes word embedding (Token Embedding) and position embedding. In the word embedding stage, each word is mapped into a predefined high-dimensional vector space, and these vectors are learned through the pre-training process and contain the basic semantic information and context information of the vocabulary. Then, position embedding is added to each word to represent its position in the sentence. For example, for the word "CPU", its word embedding vector v CPU is learned in the pre-training stage and represents the basic semantic information of CPU. The position embedding vector p CPU represents the position information of the word "CPU" in the sentence. Based on the results of word embedding and position embedding, the output vector V of the final embedding layer of the word "CPU" is obtainedCPU , where V CPU = v CPU + p CPU .

[0121] Step 3: Transformer Encoder. Through the multi-head self-attention mechanism, the RoBERTa model focuses on different parts of the input sequence from multiple perspectives to capture the relationships between words: Head 1: may focus on the relationship between "CPU" and "warning"; Head 2: may focus on the relationship between "CPU" and "abnormal fan speed"; Head 3: may focus on the relationship between "system status" and "warning". Then, further process the output of the self-attention mechanism through a Feed-Forward Neural Network to extract higher-level semantic features.

[0122] Step 4: Output semantic vectors. After being processed by the Transformer encoder, RoBERTa outputs the semantic vectors of each word. For example, the semantic vector h of the word "CPU" CPU , which contains the semantic information of the word "CPU" itself, the position information of "CPU" in the sentence, and the relationship information between "CPU" and other words such as "too high temperature" and "warning". Based on the semantic vectors of each word, obtain the semantic features of Log A and Log B:

[0123] V A = [V CLS , V 2023 ,..., V CPU ,..., V 警告 , V SEP ;

[0124] V B = [V CLS , V 2023 ,..., V 风扇 ,..., V 警告 , V SEP .

[0125] The Roberta model has the function of identifying the logical relationship and deep semantic information between logs. Therefore, the output semantic features can capture the semantic similarity between Log A and Log B, including context information and semantic association information. For example, both logs contain context information such as "warning" and "system status". RoBERTa will encode this information into the output vector, making the two vectors corresponding to "warning" and "system status" show similarity. At the same time, although "CPU temperature too high" and "abnormal fan speed" are different entities, they both belong to the category of hardware failures and have a logical relationship. RoBERTa learned this semantic association during the pre-training process. Therefore, the output semantic features will map them to nearby positions in the semantic space.

[0126] Optionally, in this embodiment, RoBERTa (convergent semantic feature extraction layer) can identify the logical relationship and deep semantic information between the reference log and the target log, and extract the reference semantic features from the reference log and the target semantic features from the target log based on the identified information, and transmit the reference semantic features and the target semantic features to the similarity calculation layer.

[0127] Optionally, in this embodiment, a cross-attention mechanism module is introduced into the semantic recognition model. Before transmitting the reference semantic features and the target semantic features to the similarity calculation layer, it further includes: inputting the reference semantic features and the target semantic features into the cross-attention mechanism module, and the cross-attention mechanism module allows the model to dynamically allocate attention weights according to the semantic relationship between the input logs. For example: the attention weight between "CPU temperature too high" in Log A and "abnormal fan speed" in Log B; the attention weight between "warning" in Log X and "warning" in Log B.

[0128] Optionally, in this embodiment, as Figure 5 shown, the similarity calculation layer can be but is not limited to a module with the function of calculating the similarity of feature vectors, such as an attention calculation layer. After receiving the reference semantic features and the target semantic features transmitted by the RoBERTa model (convergent semantic feature extraction layer), the similarity calculation layer calculates the attention weight (feature similarity) between the reference semantic features and the target semantic features, and transmits the reference semantic features, the target semantic features, and the feature similarity to the feature fusion layer.

[0129] Optionally, in this embodiment, the manner in which the similarity calculation layer calculates the attention weights between the reference semantic features and the target semantic features may include, but is not limited to: calculating the similarity value between the reference semantic features and the target semantic features by using methods such as dot product, cosine similarity, or Euclidean distance. Then, the similarity value is normalized through the Softmax function to obtain the attention weights between the reference semantic features and the target semantic features. The attention weights represent the correlation between the logs. For example, if the attention weight between log A and log B is large, it indicates that the association relationship between log A and log B in terms of log semantics is strong.

[0130] Optionally, in this embodiment, as Figure 5 shown, the feature fusion layer may be, but is not limited to, a module with the ability to fuse features, such as a pooling layer, which can perform weighted fusion (feature fusion) on the reference semantic features and the target semantic features according to the attention weights to obtain semantic features.

[0131] As an alternative solution, before calling the convergent semantic feature extraction layer to extract the reference semantic features from the reference logs and the target semantic features from the target logs, the method further includes:

[0132] S61, obtaining log sample pairs labeled with feature similarity labels, where a log sample pair includes two log samples, and the feature similarity label is used to indicate the strength of the association relationship between the two log samples included in the log sample pair in terms of log semantics;

[0133] S62, inputting the two log samples of the log sample pair into the initial semantic feature extraction layer to obtain two sample semantic features corresponding to the two log samples output by the initial semantic feature extraction layer;

[0134] S63, generating the actual feature similarity of the two sample semantic features;

[0135] S64, adjusting the model parameters of the initial semantic feature extraction layer according to the feature similarity label and the actual feature similarity to obtain the convergent semantic feature extraction layer.

[0136] Optionally, in this embodiment, before calling the RoBERTa model (convergent semantic feature extraction layer) to extract the reference semantic features and the target semantic features, the RoBERTa model needs to be trained to enable it to recognize the logical relationships and deep semantic information between the logs.

[0137] Optionally, in this embodiment, the method for training the initial RoBERTa model (convergent semantic feature extraction layer) may include, but is not limited to: First, preprocess the log sample pairs labeled with feature similarity labels into a format acceptable to the initial RoBERTa model. For example, use a tokenizer to tokenize the log text. Here, a log sample pair includes two log samples, and the feature similarity label indicates the strength of the correlation between the two log samples in terms of log semantics. For example, if the feature similarity label of log pair 1 is 0.8, it means that the two log samples in log pair 1 have a strong semantic correlation; if the feature similarity label of log pair 2 is 0.2, it means that the two log samples in log pair 2 have a weak semantic correlation. During training, the initial RoBERTa model predicts the strength of the correlation between the two log samples in the log sample pair in terms of log semantics, outputs the two sample semantic features corresponding to the two log samples according to the predicted correlation, and then generates the actual feature similarity of the two sample semantic features by calculating the attention weights. Then compare the actual feature similarity with the feature similarity label, calculate the loss between the actual feature similarity and the feature similarity label, such as using cross-entropy loss, and obtain the fourth loss value based on the actual feature similarity and the feature similarity label. Then update the model parameters of the initial RoBERTa model by backpropagation according to the fourth loss value, optimize the performance of the initial RoBERTa model until the preset convergence condition is met, such as the convergence of model parameters, to obtain the trained RoBERTa model, that is, the convergent semantic feature extraction layer. The trained RoBERTa model can be used to extract two semantic features with semantic relevance from new log pairs, where semantic relevance represents the correlation between the two logs in terms of log semantics.

[0138] As an alternative solution, generating the association parameter of the reference log according to the entity feature and the semantic feature includes:

[0139] S71, concatenate the entity feature and the semantic feature to obtain a fused feature, where the fused feature is used to represent the association between the reference log and the target log in terms of the component parameters of the server components recorded, and the association between the reference log and the target log in terms of log semantics;

[0140] S72, input the fused feature into the target classification model to obtain the association parameter output by the target classification model. Here, the target classification model is obtained by training the initial classification model with the fused feature samples labeled with association parameter labels. The fused feature samples are obtained by concatenating the entity feature samples and the semantic feature samples, and the association parameter label is used to indicate the association parameter between the reference log and the target log corresponding to the entity feature sample and the semantic feature sample corresponding to the fused feature sample.

[0141] Optionally, in this embodiment, the fused feature can be, but is not limited to, obtained by performing a vector concatenation operation on the entity feature and the semantic feature. The fused feature carries the information of the key entities included in the reference log and the target log and the information of the correlation relationship between the reference log and the target log in terms of log semantics.

[0142] Optionally, in this embodiment, the target classification model can be, but is not limited to, an AI model that has the ability to generate the correlation parameters of the input features after training, such as the ReaNet model. This model can generate the correlation parameters of the log pair indicated by the entity feature and the semantic feature in the fused feature. The correlation parameters can indicate the degree of correlation between the reference log corresponding to the fused feature and the target log.

[0143] Optionally, in this embodiment, the training process of obtaining the target classification model can be, but is not limited to, including: First, input the fused feature samples labeled with the correlation parameter labels into the initial classification model. The initial classification model outputs the predicted correlation parameters of the fused feature samples based on the initial classification model parameters. Among them, the fused feature samples are obtained by concatenating the entity feature samples and the semantic feature samples. Then, aiming to narrow the difference between the predicted correlation parameters and the correlation parameter labels, continuously adjust the initial classification model parameters until the model loss function converges or the training iteration times are completed. At this time, the initial classification model has learned the ability to generate the correlation parameters of the log pair indicated by the entity feature and the semantic feature in the fused feature through training, and the target classification model is obtained.

[0144] Optionally, in this embodiment, after inputting the fused feature into the target classification model, the target classification model can generate the correlation parameters of the log pair indicated by the entity feature and the semantic feature after analyzing and processing the entity feature and the semantic feature carried in the fused feature.

[0145] Optionally, in this embodiment, the method of screening the associated logs corresponding to the target log from the reference logs can be, but is not limited to, including: First, obtain a preset correlation parameter threshold; then, screen out the logs whose correlation parameters are greater than or equal to the correlation parameter threshold from the reference logs, which are the associated logs. Among them, the degree of correlation between the reference log whose correlation parameter is greater than or equal to the correlation parameter threshold and the target log is greater than or equal to the target correlation degree.

[0146] Optionally, in this embodiment, before screening associated logs from the reference logs, the target classification model can also perform redundancy processing on the logs to be screened according to the similarity between the logs to be screened and the already screened associated logs, including: calculating the similarity between the logs to be screened and the associated logs, and then deleting or ignoring the logs to be screened with a similarity higher than the preset similarity threshold according to the preset similarity threshold. For example, the preset similarity threshold is 0.85, and the associated logs already screened by the target classification model are:

[0147] Log A: "2023-04-01 12:00:00, CPU temperature is too high, reaching 85°C, system status: warning";

[0148] Log B: "2023-04-01 12:05:00, fan speed is abnormal, system status: warning";

[0149] The logs to be screened in the reference logs are:

[0150] Log C: "2023-04-01 12:15:00, CPU temperature is too high, reaching 87°C, system status: warning";

[0151] By calculation, the similarity situation between the logs is as follows: the similarity between Log A and Log C is 0.95, and the similarity between Log B and Log C is 0.70. According to the preset similarity threshold of 0.85, delete or ignore the logs with a similarity higher than 0.85, that is, delete or ignore C because Log C is highly similar to Log A. By performing redundancy processing on the logs to be screened in the reference logs, the data volume of the associated logs is reduced, and the efficiency and accuracy of subsequent analysis of the associated logs are improved.

[0152] Optionally, in this embodiment, by analyzing the associated logs, the root cause of the anomaly can be detected. For example, when the target log indicates that the anomaly is "the server restarts frequently", based on Log A and Log B in the associated logs, it can be analyzed that it may be due to the abnormal fan speed that the heat cannot be dissipated in time, resulting in the continuous rise of the CPU temperature. Then, the root cause of "the server restarts frequently" can be located as a fan failure. The operation and maintenance personnel only need to repair the fan to solve the anomaly of "the server restarts frequently", achieving the technical effect of accurately locating the fault source and improving the fault detection efficiency of server components.

[0153] Optionally, in this embodiment, in order to better understand the process of performing fault detection for the above-mentioned server component fault detection method, the following further describes the process of performing fault detection for the above-mentioned server component fault detection method in combination with an optional embodiment, but it is not used to limit the technical solutions of the embodiments of the present application.

[0154] In this embodiment, a method for detecting faults in server components is provided. Figure 6 It is a schematic diagram of the process of detecting faults according to the method for detecting faults in server components of an embodiment of the present application. As Figure 6 shown, it mainly includes the following steps:

[0155] First, obtain the target log carried in the fault signal output when the server has an exception and all source logs (reference logs) included in the current server log system, and input the source logs and the target log into the target recognition model. The target recognition model includes an entity extractor (entity recognition model) and a meaning extraction module (semantic recognition model). Among them, the entity extractor uses the named entity recognition (NER) technology based on the pre-trained language model BERT model, which can identify entities with specific meanings from unstructured text and classify them into predefined categories. The steps for the entity extractor to obtain the entity features of the source logs and the target log are as follows:

[0156] Step 1: BERT model. Input the source logs and the target log into the pre-trained language model (BERT model). The BERT model can identify the entities and entity types in the source logs and the target log, and output the key entities corresponding to the source logs and the key entities corresponding to the target log. By extracting these key entities, the functional information of the logs can be initially understood;

[0157] Step 2: Screening layer. According to the categories of the entities output by the BERT model, screen out the functional information (entities) related to the server state;

[0158] Step 3: Pooling layer. Input the key entities of the source logs and the target log into the pooling layer for splicing to obtain the spliced log entities;

[0159] Step 4: Feature extraction layer. Then, use methods such as the bag-of-words model and TF-IDF to extract features from the spliced log entities to obtain the entity features of the source logs and the target log.

[0160] Figure 7 It is a flowchart for extracting semantic features in the log pair according to an embodiment of the present application. As Figure 7 shown, the meaning extraction module includes a RoBERTa model and a cross-attention mechanism module. Among them, the RoBERTa model can identify the deep semantic relationship between the source logs and the target log (the correlation relationship in the log semantics), and the cross-attention mechanism module allows the model to be able to focus on the relevant information in another sequence when processing one sequence. The meaning extraction module can obtain the semantic correlation between the logs, thereby establishing the functional connection between the logs. The steps for the meaning extraction module to obtain the semantic features of the source logs and the target log are as follows:

[0161] Step 1: The RoBERTa model extracts two semantic features carrying semantic relevance from the source log and the target log based on the semantic association relationship between the two logs, and encodes the log information into vector representations;

[0162] Step 2: The attention calculation layer inputs the two obtained semantic features into the cross-attention mechanism module. The cross-attention mechanism module dynamically assigns attention weights according to the semantic relationship between the source log and the target log, then calculates the similarity value between the two semantic features, and then normalizes the similarity value through the Softmax function to obtain the attention weights between the two semantic features.

[0163] Step 3: The pooling layer then performs weighted fusion of the two semantic features based on the attention weights and the two semantic features to obtain a fused vector (semantic feature), which contains the semantic relevance information between the source log and the target log;

[0164] Step 4: The output layer takes the fused vector representation as the input for subsequent classification or clustering tasks.

[0165] Then, the obtained entity feature and semantic feature are concatenated to obtain a fused feature and input it into the ResNet model (target classification model). The ResNet model can generate the association parameters of the log pairs indicated by the entity feature and semantic feature in the fused feature according to the entity feature and semantic feature carried in the fused feature. The specific steps are as follows:

[0166] Step 1: The convolutional layer extracts log features through multiple convolutional kernels. The convolutional layer performs convolutional operations on the feature vector in a sliding window manner to extract local features.

[0167] Step 2: Residual connection. A residual connection is introduced between the convolutional layers, and the input is directly added to the output to alleviate the problems of gradient disappearance and gradient explosion.

[0168] Step 3: The fully connected layer connects the output of the convolutional layer to the fully connected layer for classification. The fully connected layer converts the output of the convolutional layer into a classification result through a weight matrix.

[0169] Step 4: The output layer outputs the classification result. The softmax function can be used to convert the classification result into a probability distribution, so as to obtain the probability value (association parameter) of each category.

[0170] Finally, the logs with association parameters greater than or equal to the association parameter threshold are filtered out from the reference logs to obtain associated logs. By analyzing the associated logs, the fault source is accurately located, thus achieving the technical effect of improving the fault detection efficiency of server components, and further solving the technical problem of low fault detection efficiency of server components.

[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0172] Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0173] In this embodiment, a fault detection device for server components is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0174] Figure 8 is a structural block diagram of a fault detection device for server components according to an embodiment of the present application; as Figure 8 shown, it includes:

[0175] A first input module 802, configured to input a reference log and a target log into a target recognition model, and obtain entity features and semantic features output by the target recognition model. Among them, the target log is a log used to indicate that a target server component is in a target fault state, the reference log is the log currently recorded by the server, the entity features are used to characterize the association relationship between the reference log and the target log in terms of the component parameters of the server component recorded, the semantic features are used to characterize the association relationship between the reference log and the target log in terms of the log semantics, and the target recognition model is obtained by training an initial recognition model using a log sample group labeled with entity feature labels and semantic feature labels;

[0176] A first generation module 804, configured to generate association parameters of the reference log according to the entity features and the semantic features, where the association parameters are used to indicate the degree of association between the reference log and the target log;

[0177] A screening module 806, configured to screen out the associated logs corresponding to the target log from the reference logs according to the association parameters;

[0178] A detection module 808, configured to detect the fault cause of the target server component being in the target fault state according to the associated logs.

[0179] In an exemplary embodiment, the target recognition model includes: an entity recognition model and a semantic recognition model. The first input module includes:

[0180] A first input unit, configured to input a reference log and a target log into the entity recognition model to obtain entity features output by the entity recognition model;

[0181] A second input unit, configured to input a reference log and a target log into the semantic recognition model to obtain semantic features output by the semantic recognition model.

[0182] In an exemplary embodiment, the entity recognition model includes: a convergence log division layer, an entity extraction layer, an entity splicing layer, and an entity feature extraction layer. The convergence log division layer is connected to the entity extraction layer, and the entity splicing layer is connected between the entity extraction layer and the entity feature extraction layer. The first input unit is further configured to:

[0183] Call the convergence log division layer to divide the reference log into multiple segments of reference log information and divide the target log into multiple segments of target log information, and transmit the reference log information and the target log information to the entity extraction layer, where each segment of log information is labeled with an information type;

[0184] Call the entity extraction layer to extract reference log entities corresponding to the reference log from the multiple segments of reference log information, and extract target log entities corresponding to the target log from the multiple segments of target log information, and input the reference log entities and the target log entities into the entity splicing layer, where the reference log entities are information for recording component parameters of server components in the reference log, and the target log entities are information for recording component parameters of server components in the target log;

[0185] Call the entity splicing layer to splice the reference log entities and the target log entities to obtain spliced log entities, and input the spliced log entities into the entity feature extraction layer;

[0186] Call the entity feature extraction layer to extract features from the spliced log entities to obtain entity features.

[0187] In an exemplary embodiment, the apparatus further includes:

[0188] A second input module, configured to input a log sample labeled with a log information tag, an information type tag, and an information position tag into an initial log partitioning layer before invoking a convergence log partitioning layer to partition a reference log into multiple segments of reference log information and partition a target log into multiple segments of target log information, so as to obtain predicted log information, a predicted information type, and a predicted information position output by the initial log partitioning layer, where the log information tag is a lexical segment in the log sample, the information type tag is used to indicate the information type of the lexical segment, the information position tag is used to indicate the position of the lexical segment in the log sample, the predicted log information is obtained by the initial log partitioning layer predicting the log information in the log sample, the predicted information type is obtained by the initial log partitioning layer identifying the information type of the lexical segment indicated by the predicted log information, and the predicted information position is obtained by the initial log partitioning layer identifying the position of the lexical segment indicated by the predicted log information in the log sample;

[0189] A first generation module, configured to generate a first loss value according to the predicted log information and the log information tag, generate a second loss value according to the predicted information type and the information type tag, and generate a third loss value according to the predicted information position and the information position tag;

[0190] A first adjustment module, configured to adjust the model parameters of the initial log partitioning layer according to the first loss value, the second loss value, the third loss value, and a preset convergence condition until the model parameters converge, so as to obtain a convergence log partitioning layer.

[0191] In an exemplary embodiment, the semantic recognition model includes: a convergence semantic feature extraction layer, a similarity calculation layer, and a feature fusion layer. The similarity calculation layer is connected between the convergence semantic feature extraction layer and the feature fusion layer. The second input unit is further configured to:

[0192] Invoke the convergence semantic feature extraction layer to extract reference semantic features from the reference log and target semantic features from the target log, and transmit the reference semantic features and the target semantic features to the similarity calculation layer;

[0193] Invoke the similarity calculation layer to calculate the feature similarity between the reference semantic features and the target semantic features, and transmit the reference semantic features, the target semantic features, and the feature similarity to the feature fusion layer, where the correlation relationship between the reference log and the target log in terms of log semantics is positively correlated with the feature similarity;

[0194] Invoke the feature fusion layer to perform feature fusion on the reference semantic features and the target semantic features according to the feature similarity, so as to obtain semantic features.

[0195] In an exemplary embodiment, the apparatus further includes:

[0196] An acquisition module, configured to obtain log sample pairs labeled with feature similarity tags before calling a convergence semantic feature extraction layer to extract reference semantic features from reference logs and target semantic features from target logs, where a log sample pair includes two log samples, and the feature similarity tag is used to indicate the strength of the association relationship in log semantics between the two log samples included in the log sample pair;

[0197] A third input module, configured to input the two log samples of the log sample pair into an initial semantic feature extraction layer to obtain two sample semantic features corresponding to the two log samples output by the initial semantic feature extraction layer;

[0198] A second generation module, configured to generate the actual feature similarity of the two sample semantic features;

[0199] A second adjustment module, configured to adjust the model parameters of the initial semantic feature extraction layer according to the feature similarity tag and the actual feature similarity to obtain a convergence semantic feature extraction layer.

[0200] In an exemplary embodiment, the first generation module includes:

[0201] A splicing unit, configured to splice the entity feature and the semantic feature to obtain a fusion feature, where the fusion feature is used to characterize the association relationship between the reference log and the target log in terms of component parameters of the server components recorded, and the association relationship between the reference log and the target log in log semantics;

[0202] A third input unit, configured to input the fusion feature into a target classification model to obtain an association parameter output by the target classification model, where the target classification model is obtained by training an initial classification model with a fusion feature sample pair labeled with an association parameter tag, the fusion feature sample is obtained by splicing an entity feature sample and a semantic feature sample, and the association parameter tag is used to indicate the association parameter between the reference log and the target log corresponding to the entity feature sample and the semantic feature sample corresponding to the fusion feature sample.

[0203] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are separately located in different processors in any combination form.

[0204] For the description of the features in the corresponding embodiment of the fault detection device of the server component, reference can be made to the relevant description in the corresponding embodiment of the fault detection method of the server component, which will not be elaborated here one by one.

[0205] An embodiment of the present application further provides an electronic device, Figure 9 is a schematic diagram of the electronic device according to the embodiment of the present application, asFigure 9 As shown in Figure 9 , the electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the embodiment of the fault detection method for any one of the above server components.

[0206] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0207] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary embodiments, and details will not be repeated here.

[0208] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (abbreviated as ROM), random access memory (abbreviated as RAM), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.

[0209] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (abbreviated as ROM), random access memory (abbreviated as RAM), mobile hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.

[0210] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the methods in the various embodiments of the present application; the computer program product further includes a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores the computer program. When the computer program is executed by the processor, it implements the steps of the fault detection method for server components in the various embodiments of the present application.

[0211] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0212] The above has introduced in detail a method for detecting faults in server components provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for detecting faults in server components, characterized in that, The method includes: Inputting the reference log and the target log into the target recognition model to obtain the entity features and semantic features output by the target recognition model. Among them, the target log is a log indicating that the target server component is in a target failure state, the reference log is the log currently recorded by the server, the entity features are used to characterize the association relationship between the reference log and the target log in terms of the component parameters of the recorded server components, the semantic features are used to characterize the association relationship between the reference log and the target log in terms of log semantics, and the target recognition model is obtained by training an initial recognition model using a log sample group labeled with entity feature labels and semantic feature labels; Generating an association parameter for the reference log according to the entity features and the semantic features, where the association parameter is used to indicate the degree of association between the reference log and the target log; Screening out the associated log corresponding to the target log from the reference log according to the association parameter; Detecting the cause of the target server component being in the target failure state according to the associated log.

2. The method according to claim 1, wherein: The target recognition model includes: an entity recognition model and a semantic recognition model. The inputting the reference log and the target log into the target recognition model to obtain the entity features and semantic features output by the target recognition model includes: Inputting the reference log and the target log into the entity recognition model to obtain the entity features output by the entity recognition model; Inputting the reference log and the target log into the semantic recognition model to obtain the semantic features output by the semantic recognition model.

3. The method according to claim 2, wherein: The entity recognition model includes: A convergence log division layer, an entity extraction layer, an entity splicing layer, and an entity feature extraction layer. The convergence log division layer is connected to the entity extraction layer, and the entity splicing layer is connected between the entity extraction layer and the entity feature extraction layer. The inputting the reference log and the target log into the entity recognition model to obtain the entity features output by the entity recognition model includes: Invoking the convergence log division layer to divide the reference log into multiple segments of reference log information, and dividing the target log into multiple segments of target log information, and transmitting the reference log information and the target log information to the entity extraction layer, where each segment of log information is labeled with an information type; Invoking the entity extraction layer to extract the reference log entity corresponding to the reference log from multiple segments of the reference log information, and extract the target log entity corresponding to the target log from multiple segments of the target log information, and input the reference log entity and the target log entity into the entity splicing layer, where the reference log entity is the information for recording the component parameters of the server component in the reference log, and the target log entity is the information for recording the component parameters of the server component in the target log; Call the entity splicing layer to splice the reference log entity and the target log entity to obtain a spliced log entity, and input the spliced log entity into the entity feature extraction layer; Call the entity feature extraction layer to extract features from the spliced log entity to obtain the entity features.

4. The method according to claim 3, wherein, Before the step of calling the convergent log division layer to divide the reference log into multiple segments of reference log information and divide the target log into multiple segments of target log information, the method further includes: Input a log sample labeled with a log information label, an information type label, and an information position label into an initial log division layer to obtain predicted log information, predicted information type, and predicted information position output by the initial log division layer, where the log information label is a lexical segment in the log sample, the information type label is used to indicate the information type of the lexical segment, the information position label is used to indicate the position of the lexical segment in the log sample, the predicted log information is obtained by the initial log division layer predicting the log information in the log sample, the predicted information type is obtained by the initial log division layer identifying the information type of the lexical segment indicated by the predicted log information, and the predicted information position is obtained by the initial log division layer identifying the position of the lexical segment indicated by the predicted log information in the log sample; Generate a first loss value according to the predicted log information and the log information label, generate a second loss value according to the predicted information type and the information type label, and generate a third loss value according to the predicted information position and the information position label; Adjust the model parameters of the initial log division layer according to the first loss value, the second loss value, the third loss value, and a preset convergence condition until the model parameters converge to obtain the convergent log division layer.

5. The method according to claim 2, wherein, The semantic recognition model includes: a convergent semantic feature extraction layer, a similarity calculation layer, and a feature fusion layer. The similarity calculation layer is connected between the convergent semantic feature extraction layer and the feature fusion layer. The step of inputting the reference log and the target log into the semantic recognition model to obtain the semantic features output by the semantic recognition model includes: Call the convergent semantic feature extraction layer to extract reference semantic features from the reference log and target semantic features from the target log, and transmit the reference semantic features and the target semantic features to the similarity calculation layer; Call the similarity calculation layer to calculate the feature similarity between the reference semantic features and the target semantic features, and transmit the reference semantic features, the target semantic features, and the feature similarity to the feature fusion layer, where the correlation between the association relationship of the reference log and the target log in log semantics and the feature similarity is positive; Call the feature fusion layer to perform feature fusion on the reference semantic feature and the target semantic feature according to the feature similarity to obtain the semantic feature.

6. The method according to claim 5, wherein Before calling the convergent semantic feature extraction layer to extract the reference semantic feature from the reference log and the target semantic feature from the target log, the method further includes: Obtain a log sample pair labeled with a feature similarity label, where the log sample pair includes two log samples, and the feature similarity label is used to indicate the strength of the association relationship in log semantics between the two log samples included in the log sample pair; Input the two log samples of the log sample pair into the initial semantic feature extraction layer to obtain two sample semantic features corresponding to the two log samples output by the initial semantic feature extraction layer; Generate the actual feature similarity of the two sample semantic features; Adjust the model parameters of the initial semantic feature extraction layer according to the feature similarity label and the actual feature similarity to obtain the convergent semantic feature extraction layer.

7. The method according to claim 1, wherein Generating the association parameter of the reference log according to the entity feature and the semantic feature includes: Concatenate the entity feature and the semantic feature to obtain a fusion feature, where the fusion feature is used to characterize the association relationship between the reference log and the target log in terms of the component parameters of the server component recorded, and the association relationship between the reference log and the target log in log semantics; Input the fusion feature into a target classification model to obtain the association parameter output by the target classification model, where the target classification model is obtained by training an initial classification model using a fusion feature sample pair labeled with an association parameter label, the fusion feature sample is obtained by concatenating an entity feature sample and a semantic feature sample, and the association parameter label is used to indicate the association parameter between the reference log and the target log corresponding to the entity feature sample and the semantic feature sample corresponding to the fusion feature sample.

8. A failure detection device for a server component, characterized in that, The apparatus includes: A first input module, configured to input a reference log and a target log into a target recognition model to obtain entity features and semantic features output by the target recognition model, where the target log is a log indicating that a target server component is in a target fault state, the reference log is the currently recorded log of the server, the entity feature is used to characterize the association relationship between the reference log and the target log in terms of the component parameters of the server component recorded, the semantic feature is used to characterize the association relationship between the reference log and the target log in log semantics, and the target recognition model is obtained by training an initial recognition model using a log sample group labeled with entity feature labels and semantic feature labels; A first generation module, configured to generate an association parameter of the reference log according to the entity feature and the semantic feature, where the association parameter is used to indicate the degree of association between the reference log and the target log. A screening module, configured to screen out the associated logs corresponding to the target logs from the reference logs according to the associated parameters; A detection module, configured to detect the cause of the failure of the target server component in the target failure state according to the associated logs.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the failure detection method of the server component according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the failure detection method of the server component according to any one of claims 1 to 7 when executed by a processor.