Server anomaly detection method, training method and electronic equipment
Through the collaborative detection method of deep learning model and pre-trained language model, the problem of insufficient accuracy of server fault detection is solved, efficient identification and root cause analysis of complex fault patterns are achieved, and the accuracy and efficiency of server abnormal detection are improved.
Patent Information
- Application Number
- CN202510837789.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The prior art has low accuracy in server fault detection, especially when facing complex and diverse fault modes, the detection accuracy is insufficient.
The deep learning model and the pre-trained language model collaborative detection method are adopted. By embedding the position coding and attention mechanism fusion features in the log sequence to be detected, and logical reasoning and causal analysis are combined with the pre-trained language model to be performed to generate more accurate anomaly detection results.
It improves the accuracy of server abnormal detection, reduces the interference of redundant data on detection results, can better identify complex failure modes and analyze their root causes, and improves detection efficiency and accuracy.
Smart Images

Figure CN120336980A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to the fields of deep learning and anomaly detection technology, and more specifically to a server anomaly detection method, a training method, and an electronic device. Background Art
[0002] With the rapid expansion of the scale of data centers, the number of servers has increased exponentially, and the complexity of the operating environment has increased significantly, increasing the frequency of server failures.
[0003] In order to reduce the frequency of server failures, it is necessary to perform anomaly detection on servers. However, due to the complex and diverse types of server failures, such as: memory corruption, hard disk failure, operating system failure, driver conflict, etc., the detection accuracy is relatively low. Summary of the Invention
[0004] In view of the above problems, this application provides a server anomaly detection method, a training method, and an electronic device.
[0005] According to the first aspect of this application, a server anomaly detection method is provided, including: in response to determining that the initial detection result of the target server indicates an anomaly, using the target model, based on the timestamps of the logs to be detected in the log sequence to be detected, embedding positional encodings in the log sequence to be detected, and generating features of the logs to be detected; using the target model, based on the attention mechanism, fusing the features of any log to be detected with the features of at least two adjacent logs to be detected in the log sequence to be detected, and generating a first detection result; where at least two logs to be detected indicate different log event types; and in response to determining that the first detection result includes at least two anomaly types with a correlation degree less than a predetermined threshold, using a pre-trained language model to analyze the log sequence to be detected, and generating a second detection result; where the second detection result indicates the server anomaly type and the cause of the server anomaly.
[0006] The second aspect of this application provides a training method for a target model, including: using an initial model, based on the timestamps of the sample logs in the sample log sequence, embedding sample positional encodings in the sample log sequence, and generating features of the sample logs; using the initial model, based on the attention mechanism, fusing the features of any sample log in the sample log sequence with the features of at least two sample logs in the adjacent sample log sequence, and generating a sample detection result; at least two sample logs indicate different log event types; training the initial model based on the target loss function, the sample detection result, and the sample label to generate the target model involved in the server anomaly detection method described above.
[0007] The third aspect of the present application provides a server anomaly detection device, including: a first encoding module, a first detection module, and an analysis module.
[0008] The first encoding module is configured to, in response to determining that the initial detection result of the target server indicates an anomaly, use the target model to embed position encoding in the to-be-detected log sequence based on the timestamps of the to-be-detected logs in the to-be-detected log sequence, and generate features of the to-be-detected logs.
[0009] The first detection module is configured to use the target model to fuse the features of any to-be-detected log with the features of at least two adjacent to-be-detected logs respectively based on the attention mechanism, and generate a first detection result; wherein, the log event types indicated by the at least two to-be-detected logs are different.
[0010] The first analysis module is configured to, in response to determining that the first detection result includes at least two anomaly types with a correlation degree less than a predetermined threshold, use a pre-trained language model to analyze the to-be-detected log sequence, and generate a second detection result; wherein, the second detection result indicates the server anomaly type and the cause of the server anomaly.
[0011] The fourth aspect of the present application provides a training device for a target model, including: a second encoding module, a second detection module, and a training module.
[0012] The second encoding module is configured to use the initial model to embed sample position encoding in the sample log sequence based on the timestamps of the sample logs in the sample log sequence, and generate features of the sample logs.
[0013] The second detection module is configured to use the initial model to fuse the features of any sample log in the sample log sequence with the features of at least two sample logs in the respective adjacent sample log sequences based on the attention mechanism, and generate a sample detection result; the log event types indicated by the at least two sample logs are different;
[0014] The training module is configured to train the initial model based on the target loss function, the sample detection result, and the sample label to generate a target model. Wherein, the target loss function includes parameters for dynamically invoking sample class weights and sample recognition difficulty weights; the sample label indicates the anomaly category of the sample log sequence.
[0015] The fifth aspect of the present application provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein, the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0016] The sixth aspect of the present application further provides a computer-readable storage medium, on which a computer program or instructions are stored, and when the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0017] The seventh aspect of the present application further provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0018] According to the embodiments of the present application, when the initial detection result of the log sequence to be detected for the target server indicates an anomaly, the target model is called for detection. Since the target model integrates the features of any log to be detected with those of at least two adjacent logs to be detected, it at least reduces the interference of redundant data on the detection result during the feature fusion process, and further improves the accuracy of the detection result. When the detection result of the target model includes at least two unrelated anomaly types, a pre-trained language model is called for analysis. By leveraging the semantic understanding ability and logical reasoning ability of the pre-trained language model, it forms an organic whole for collaborative detection with the deep learning model, at least solving the problem that the target model has a low detection accuracy for complex fault patterns due to its low generalization ability, and achieving the technical effect of improving the anomaly detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Through the following description of the embodiments of the present application with reference to the drawings, the above content and other objectives, features, and advantages of the present application will become clearer. In the drawings:
[0020] Figure 1 Shows an application scenario diagram of the server anomaly detection method according to the embodiments of the present application;
[0021] Figure 2 Shows a flowchart of the server anomaly detection method according to the embodiments of the present application;
[0022] Figure 3 Shows a schematic diagram of the server anomaly detection method according to the embodiments of the present application;
[0023] Figure 4 Shows a schematic diagram of calling a pre-trained language model to decompress a log file and parse log data to generate a log sequence to be detected according to the embodiments of the present application;
[0024] Figure 5 Shows a schematic diagram of using the target model to detect the log sequence to be detected according to the embodiments of the present application;
[0025] Figure 6 Shows a schematic diagram of calling a pre-trained language model to perform collaborative analysis on the log sequence to be detected according to the embodiments of the present application;
[0026] Figure 7 shows a flowchart of a target model training method according to an embodiment of the present application;
[0027] Figure 8 shows a schematic diagram of a target model training method according to an embodiment of the present application;
[0028] Figure 9 shows a structural block diagram of a server anomaly detection device according to an embodiment of the present application;
[0029] Figure 10 shows a structural block diagram of a target model training device according to an embodiment of the present application;
[0030] Figure 11 shows a block diagram of an electronic device suitable for implementing a server anomaly detection method or a target model training method according to an embodiment of the present application. Detailed implementation manners
[0031] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present application. However, obviously, one or more embodiments may be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0032] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0035] As the scale of the data center expands, the operating environment of the server becomes increasingly complex, and the types of server anomalies or faults are also complex and diverse. The methods for detecting server anomalies in related examples all have their respective drawbacks. For example, the anomaly detection method based on a statistical model relies on assumptions about the data distribution and deviates from the actual server operating scenario. The anomaly detection method based on expert rules relies on subjective expert experience and is difficult to flexibly adapt to the complex and changeable operating environment. The anomaly detection method based on a deep learning model depends on the quality of the sample data for detection accuracy.
[0036] In view of this, an embodiment of the present application provides a server anomaly detection method, which forms an organic whole for collaborative detection with a deep learning model, and at least solves the problem that the target model has a low detection accuracy for complex fault modes due to its low generalization ability, achieving the technical effect of improving the anomaly detection accuracy.
[0037] Figure 1 The application scenario diagram of the server anomaly detection method according to an embodiment of the present application is shown.
[0038] As Figure 1 shown, the application scenario 100 according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for the communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0039] Users can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, data analysis applications, etc. (only as examples).
[0040] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.
[0041] The server 103 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the terminal device 101 (only as an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0042] It should be noted that the server anomaly detection method provided by the embodiments of the present application can generally be executed by the server 103. The server 103 can be configured with a target model and a pre-trained language model for executing the server anomaly detection method described above. Correspondingly, the server anomaly detection device or the training device of the target model provided by the embodiments of the present application can generally be set in the server 103.
[0043] For example: The log sequence to be detected is input into the terminal device 101, and the terminal device 101 sends the log sequence to be detected to the server 103. The server 103 uses the target model and the pre-trained language model (LLM) to execute the server anomaly detection method provided by the embodiments of the present application to generate the detection result 111. Finally, the detection result 111 is fed back to the terminal device 101.
[0044] The server anomaly detection method or the training method of the target model provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103. Correspondingly, the server anomaly detection device or the training device of the target model provided by the embodiments of the present application can also be set in a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103.
[0045] For example: The terminal device 101 can be loaded with a target model and a pre-trained language model. When the log sequence to be detected is input into the terminal device 101, the terminal device 101 can call the target model and the pre-trained language model to execute the server anomaly detection method provided by the embodiments of the present application to generate the detection result 111.
[0046] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in
[0047] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the scenario described below Figures 2 to 6 the server anomaly detection method of the embodiment will be described in detail through
[0048] Figure 2 shows a flowchart of the server anomaly detection method according to the embodiments of the present application.
[0049] As Figure 2 shown, the server anomaly detection method 200 of this embodiment includes operations S210 to S230.
[0050] In operation S210, in response to determining that the initial detection result of the target server indicates an anomaly, the target model is utilized to embed positional encoding in the sequence of logs to be detected based on the timestamps of the logs to be detected in the sequence of logs to be detected, generating features of the logs to be detected.
[0051] In operation S220, the target model is utilized to fuse the features of any log to be detected with the features of at least two adjacent logs to be detected respectively based on the attention mechanism, generating a first detection result.
[0052] In operation S230, in response to determining that the first detection result includes at least two anomaly types with a correlation degree less than a predetermined threshold with each other, the pre-trained language model is utilized to analyze the sequence of logs to be detected, generating a second detection result.
[0053] In some embodiments, the initial detection result can be obtained by detecting the sequence of logs to be detected based on a rule engine. Various empirical rules or statistical rules for detecting server anomalies can be configured in the rule engine. For example: the empirical rule can be that when the temperature of hardware Q continuously rises and exceeds a predetermined temperature threshold, it can be determined that hardware Q in the server has a fault S1. When the temperature of hardware Q recorded in the sequence of logs to be detected continuously rises and exceeds the predetermined temperature threshold, it can be determined that it matches this empirical rule, and the initial detection result is determined to be that hardware Q in the server has a fault S1.
[0054] However, statistical rules often rely on assumptions about data distribution, and the data distribution in actual application scenarios is complex and variable, making it difficult to meet the assumptions relied on by statistical rules. Therefore, it may lead to the initial detection result detected based on statistical rules indicating an anomaly. For example: the initial detection result is empty.
[0055] Similarly, empirical rules are usually summarized based on historical experience. As the complexity of the operating environment of the data center increases, the types of anomaly types that occur in the server become more and more complex and variable. Therefore, it may lead to the initial detection result detected based on empirical rules indicating an anomaly. For example: the initial detection result includes multiple anomaly types that conflict with each other.
[0056] Therefore, in the embodiments of the present application, when it is determined that the initial detection result indicates an anomaly, the target model is utilized to embed positional encoding in the sequence of logs to be detected based on the timestamps of the logs to be detected in the sequence of logs to be detected, generating features of the logs to be detected.
[0057] In some embodiments, the target model can be a trained Transformer model. The Transformer model is a neural network architecture based on the attention mechanism, which can dynamically calculate the association weights of different positions in the input sequence based on the attention mechanism, thereby capturing the dependencies between long-distance features in the input sequence.
[0058] Embed positional encoding in the log sequence to be detected based on the timestamps of the logs to be detected in the log sequence to be detected, so that the model can perceive the actual time when the logs occur during the detection process, rather than the relative position, thereby perceiving the density of the log events represented in the logs to be detected, and further improving the accuracy of the detection results.
[0059] Since the server will automatically adjust for certain anomalies during operation to reduce the probability of failures. For example, when the power consumption of the central processing unit (CPU) is too high, the temperature of the CPU will continue to rise. At this time, the fan starts to assist the CPU in heat dissipation. During this process, the temperature of the CPU may fluctuate, first decreasing, then increasing, and then decreasing again. For this normal fluctuation phenomenon, if it is identified as an anomaly of the CPU or the fan, it will lead to false detection.
[0060] Therefore, in the embodiments of the present application, use the target model to fuse the features of any log to be detected with the features of at least two adjacent logs to be detected based on the attention mechanism to generate a first detection result. Thereby reducing the repeated perception of the model for the same type of log events occurring at adjacent times, and reducing the false detection probability for the normal fluctuation phenomenon during the server's self-regulation process.
[0061] Since the detection accuracy of the target model depends highly on the sample data but has low generalization, the detection accuracy for complex fault patterns is poor. The pre-trained language model is obtained by pre-training based on deep learning technology through a large amount of text data, and has strong semantic understanding ability and logical reasoning ability.
[0062] Therefore, when the first detection result output by the target model includes at least two anomaly types with a correlation less than a predetermined threshold, the pre-trained language model can be used to analyze the log sequence to be detected. For example, after comprehensively analyzing the detection results of different types of logs to be detected, through causal reasoning of the detection results, determine whether there is a logical conflict between the detection results. If not, the detection results can be fused to generate a second detection result.
[0063] According to an embodiment of the present application, when the initial detection result of the log sequence to be detected for the target server indicates an anomaly, the target model is called for detection. Since the target model fuses the features of any log to be detected with the features of at least two adjacent logs to be detected, it at least reduces the interference of redundant data on the detection result during the feature fusion process, and further improves the accuracy of the detection result. When the detection result of the target model includes at least two unrelated anomaly types, a pre-trained language model is called for analysis. By utilizing the semantic understanding ability and logical reasoning ability of the pre-trained language model, it at least solves the problem that the target model has a low detection accuracy for complex fault patterns due to its low generalization ability, achieving the technical effect of improving the anomaly detection accuracy.
[0064] In an actual application scenario, the log files of the servers in the data center are usually stored in a compressed file format. The log files are usually generated by different systems and have inconsistent formats. Therefore, the compressed files need to be preprocessed before the log sequences to be detected can be obtained. The operations of the preprocessing include, but are not limited to, decompressing the log files and data parsing, etc.
[0065] In addition, based on the rule engine, simple anomaly types can be quickly identified. For example, by performing rule matching on a single log message, it can be quickly determined whether there is an anomaly in the server. Thereby, the detection efficiency is improved.
[0066] The following Figure 3 details the collaborative detection method that integrates the rule engine, the target model, and the pre-trained language model.
[0067] Figure 3 shows a schematic diagram of the server anomaly detection method according to an embodiment of the present application.
[0068] As Figure 3 shown, in this embodiment 300, first, the compressed file 301 of the log to be detected is preprocessed to obtain the log sequence 302 to be detected.
[0069] Then, the rule engine 310 is called to detect the log sequence 302 to be detected, and the initial detection result 303 is obtained.
[0070] Next, operation S310 is executed to determine whether the initial detection result 303 indicates an anomaly conflict. If so, operation S320 is executed to call the target model for detection to obtain the first detection result. If not, the initial detection result 303 is determined as the final detection result, and the process ends.
[0071] Next, perform operation S321 to determine whether the correlation degree between the abnormal types in the first detection result is less than a predetermined threshold. If so, perform operation S330 to call a pre-trained language model for detection and generate a second detection result 304. If not, determine the first detection result as the final detection result and end the process.
[0072] Next, combine Figures 4 to 6 with Figure 3 the server anomaly detection method shown in
[0073] Figure 4 shows a schematic diagram of decompressing a log file and parsing log data by calling a pre-trained language model according to an embodiment of the present application to generate a log sequence to be detected.
[0074] Preprocessing the compressed file of the log to be detected to generate a log sequence to be detected may include the following operations: obtaining the compressed file of the log to be detected; determining the type of the compressed file based on the content of the compressed file; based on the type of the compressed file, using a pre-trained language model to call a first tool to decompress the compressed file to generate multiple log files; using a pre-trained language model to call a second tool to parse the multiple log files respectively to generate a log sequence to be detected.
[0075] In some embodiments, the type of the compressed file may be determined based on the extension name of the compressed file, such as: zip, rar, etc. The type of the compressed file may also be determined by identifying the file header of the compressed file. Therefore, the content of the compressed file includes but is not limited to the file header, extension name, and file information included in the compressed file.
[0076] Then, based on the type of the compressed file, a pre-trained language model may be used to call a first tool to decompress the compressed file.
[0077] In some embodiments, the mapping relationship between various types of compressed files and decompression tools and parsing tools may be pre-configured, and the call interfaces of the decompression tools and parsing tools may be configured.
[0078] Then, the pre-trained language model may determine the first tool for decompressing the compressed file of this type and the interface information for calling the first tool by identifying the type of the compressed file, so as to call the first tool based on the interface information of the first tool to decompress the compressed file.
[0079] For example: for a zip-type compressed file, the pre-trained language model may call a Python tool for decompression.
[0080] Then, the pre-trained language model can determine the interface information of the second tool to call the second tool to parse the decompressed multiple logs, for example: data deduplication, data analysis, conversion of unstructured log data into structured log data, etc., to generate a log sequence to be detected.
[0081] In the embodiment of the present application, identifying the type of the compressed file based on the content of the compressed file further improves the accuracy of identifying the type of log compressed files with different storage formats from different sources. Using the pre-trained language model can dynamically select the decompression tool and the parsing tool based on the type of the compressed file, which further improves the execution efficiency of the preprocessing operation on the compressed file of the log to be detected.
[0082] In some embodiments, in addition to pre-configuring the interface information of each parsing tool and decompression tool, the generation ability of the pre-trained language model can also be directly utilized to directly generate parameters for calling the decompression tool and the parsing tool that match the type of the compressed file, which further improves the data processing efficiency and also further improves the flexibility of calling tools.
[0083] For example: Based on the type of the compressed file, use the pre-trained language model to call the first tool to decompress the compressed file to generate multiple log files, which may include the following operations: determining the parameter information of the first tool based on the type of the compressed file; using the pre-trained language model, based on the parameter information of the first tool, generating a first parameter for calling the first tool; and calling the first tool to decompress the compressed file based on the first parameter to generate multiple log files.
[0084] In the embodiment of the present application, the parameter information of the first tool includes but is not limited to information such as the format and content of the input parameters required when calling the first tool.
[0085] As Figure 4 shown, the compressed file 301 of the log to be detected can be input into the LLM1410, and the first parameter is output.
[0086] For example: The type of the compressed file and the compressed file can be input as a Prompt (hint) into the pre-trained language model. The pre-trained language model determines the parameter information of the first tool by identifying the type of the compressed file and outputs the first parameter. It is also possible to first determine the parameter information of the first tool corresponding to the type of the pre-compressed file, and then input the parameter information of the first tool and the compressed file as a Prompt into the pre-trained language model to generate the first parameter.
[0087] Finally, use the first parameter to call the decompression tool 411 to decompress the compressed file of the log to be detected to generate the log file 401.
[0088] Similarly, the pre-trained language model is used to call the second tool to parse multiple log files respectively, generating the log sequences to be detected, including the following operations: determining the parameter information of the second tool based on the parsing requirements of the log to be detected; using the pre-trained language model to generate the second parameter for calling the second tool based on the parameter information of the second tool; and calling the second tool to parse multiple log files based on the second parameter, generating the log sequences to be detected.
[0089] In the embodiments of the present application, the parameter information of the second tool includes but is not limited to information such as the format and content of the input parameters required when calling the second tool.
[0090] The parsing requirements of the log to be detected can be determined according to the log source, including but not limited to: removing redundant data, converting unstructured log data into structured log data, etc. Removing redundant data can reduce the overhead of data storage and processing. The removal of redundant data can usually be achieved by finding duplicate records. The parsed structured log data can be used by subsequent rule engines, target models, etc. to extract keyword fields, format log timestamps, or identify the types of log events.
[0091] As Figure 4 shown, the log file 401 can be input into the LLM2420, and the second parameter is output.
[0092] For example: The parsing requirements and the log file can be input as a Prompt into the pre-trained language model. The pre-trained language model determines the parameter information of the second tool by identifying the parsing requirements and outputs the second parameter. It is also possible to first determine the parameter information of the second tool corresponding to the pre-compressed file type, and then input the parameter information of the second tool and the log file as a Prompt into the pre-trained language model to generate the second parameter.
[0093] Finally, the parsing tool 421 is called using the second parameter to perform data parsing on the log file 401, generating the log sequence 302 to be detected.
[0094] Using the generation ability of the pre-trained model to dynamically generate the parameters for the decompression tool used to decompress the log file and the parsing tool for parsing the log file further improves the flexibility of the preprocessing operation for the compressed file of the log to be detected and adapts to the increasingly complex and changeable operating environment of the server.
[0095] In actual application scenarios, there are still some types of anomalies that can be detected based on the rule engine. The method of using the rule engine to detect server anomalies has the characteristic of fast data processing speed and is suitable for detection scenarios with high requirements for timeliness.
[0096] However, as the operating environment of the servers in the data center becomes increasingly complex and changeable, the rules for detecting server anomalies are also becoming more and more complex. There may also be correlation relationships between different rules. The reproduction rate of the anomaly logs adapted to such complex determination rules is relatively low.
[0097] Therefore, when there are a large number of correlation rules in the historical rule library, the processing speed of the rule engine is instead reduced.
[0098] In view of this, the embodiments of the present application extract uncorrelated rule sets from the historical rule library through the following operations to obtain a predetermined rule library.
[0099] For example: The correlation strength between rules can be defined as follows:
[0100] (1)
[0101] In the embodiments of the present application, two rules with a correlation degree greater than a predetermined correlation threshold can be determined to have a correlation relationship with each other.
[0102] Then, the rules in the historical rule library are scored according to the following scoring rules:
[0103] (2)
[0104] Next, based on the scoring results obtained from formula (2), the rules in the historical rule library are sorted, and the top N rules can be selected to construct a predetermined rule library. N can be a positive integer greater than 1, and the specific value can be configured according to actual needs. For example: N = 100.
[0105] Finally, based on the predetermined rule library, rule matching is performed on the log sequence to be detected to generate an initial detection result. The predetermined rule library includes multiple predetermined rules for detecting server anomalies; the correlation degrees between the multiple predetermined rules are less than the predetermined correlation threshold.
[0106] Since there is no correlation between the multiple predetermined rules for detecting server anomalies included in the predetermined rule library, the number of rules to be matched and the matching difficulty during the invocation of the rule engine are reduced, further improving the detection efficiency in the initial detection stage.
[0107] Next, in combination with Figure 5 embodiments of using the target model to detect the log sequence to be detected will be described in detail.
[0108] Figure 5 shows a schematic diagram of using the target model to detect the log sequence to be detected according to the embodiments of the present application.
[0109] In the server anomaly detection scenario, the actual time when the anomaly event occurs is often more important than the relative time when the anomaly occurs.
[0110] Therefore, in the embodiments of the present application, when performing position encoding on the log sequence to be detected, the timestamp of the log to be detected is introduced to truly represent the real event interval between different log events.
[0111] As Figure 5 shown, first, the time interval 520 is introduced to perform position encoding on the log sequence 302 to be detected, and the features F1, F2,..., Fn 521 of the log to be detected are generated.
[0112] In some embodiments, using the target model, based on the timestamp of the log to be detected in the log sequence to be detected, embedding position encoding in the log sequence to be detected to generate the features of the log to be detected may include the following operations: obtaining a first time interval between the logs to be detected that are adjacent in position in the log sequence to be detected based on the timestamp of the log to be detected in the log sequence to be detected; generating position encoding based on the timestamp of the log to be detected and the first time interval; and embedding the position encoding in the log sequence to be detected to generate the features of the log to be detected.
[0113] For example: The position encoding can be generated according to formula (3) :
[0114] (3)
[0115] where k = 1, 2,..., d model , representing the dimension in the log sequence to be detected; t represents the timestamp of the log; represents the first time interval between the logs to be detected that are adjacent in position in the log sequence to be detected; represents the position encoding of the absolute time; represents the position encoding of the relative time interval.
[0116] In the embodiments of the present application, the position encoding of the absolute time is used to retain the global timing information of the log events. For example: the specific moment of a certain fault in the server operation cycle. The position encoding of the relative time interval is used to capture the local timing relationship between adjacent log events. For example: whether the interval between two abnormal logs is "suddenly continuously occurring".
[0117] Based on the time interval when the log event occurs, the position encoding is dynamically adjusted, so that while the target model perceives the global timing information of the log events, it can perceive the density of the occurrence of the log events, and further improve the accuracy of server anomaly detection.
[0118] Since each log sequence to be detected includes multiple log messages, feature fusion can be performed hierarchically. For example: sentence-level attention mechanism fusion can be performed first, that is, each log message is regarded as a sentence, and self-attention mechanism fusion can be performed. Then, the fused features are further fused at the passage level, that is, the features of the entire log sequence to be detected are fused through an interactive attention mechanism to generate fused feature 523. Then, the fused feature 523 is detected to generate the first detection result 524.
[0119] In some embodiments, using a target model, based on the attention mechanism, fusing the features of any log to be detected with the features of at least two adjacent logs to be detected to generate a first detection result may include the following operations: based on the attention mechanism, fusing the features of the log to be detected with the features of at least two adjacent logs to be detected to generate a fused feature; and detecting the fused feature to generate a first detection result.
[0120] Since the event interval when the log event actually occurs is introduced during the position encoding process, the model will pay more attention to the log events that occur within a shorter time interval during the feature fusion process. However, due to the server having an automatic regulation mechanism, during the automatic regulation process, it may cause the same type of log events to be repeatedly triggered within a short period of time. The repeated triggering of the same type of log events within a short period of time is actually not an abnormal operation of the server, but the automatic regulation mechanism of the server is activated. Therefore, when performing feature fusion based on the attention mechanism, the model can be made to pay more attention to different types of log events that occur at adjacent moments by introducing a penalty parameter 522.
[0121] In some embodiments, based on the attention mechanism, fusing the features of the log to be detected with the features of at least two adjacent logs to be detected to generate a fused feature may include the following operations: obtaining a penalty parameter according to the timestamps and target information of at least two logs to be detected; based on the attention mechanism and the penalty parameter, fusing the features of the log to be detected with the features of at least two adjacent logs to be detected to generate a fused feature.
[0122] In the embodiments of the present application, the target information is used to indicate whether the types of the log events indicated by at least two logs to be detected are the same.
[0123] For example: the fused feature can be generated based on the attention mechanism and the penalty parameter according to formula (4):
[0124] (4)
[0125] Where Q, K, and V respectively represent the query matrix, key matrix, and value matrix. represents the dimension of the key vector, and the scaling factor Used to prevent the gradient from vanishing due to an excessive dot product quantity. λ ∈ [0, 1] represents the penalty coefficient, which is used to control the intensity of duplicate suppression. R ∈ represents the penalty matrix, and the dimension of the penalty matrix is the same as that of the Q, K, and V matrices.
[0126] In some embodiments, the penalty coefficient and the penalty matrix can be pre-configured to reduce the attention to log events of the same type within adjacent time instants.
[0127] Due to the repeated occurrence of log events within a short time period triggered by the server's automatic regulation mechanism, and the false detection events that may be caused by the model's attention to the intensive occurrence of repeated log events, the introduction of penalty parameters in the attention mechanism reduces the probability of false detection events.
[0128] In addition to pre-configuring the penalty coefficient and the penalty matrix, in some embodiments, the penalty parameter can also be obtained according to the timestamps and target information of at least two logs to be detected, so as to dynamically adapt the penalty intensity according to the changes in the actual detection scenario and further improve the detection accuracy.
[0129] In some embodiments, the second time interval between at least two logs to be detected can be obtained according to the timestamps of at least two logs to be detected; and the penalty parameter is generated based on the second time interval and the target information.
[0130] For example: The element R in the penalty matrix can be defined ij as follows:
[0131] (5)
[0132] where, t i represents the timestamp of log event i, t j represents the timestamp of log event j, and τ represents the event decay constant, which is used to control the decay speed. For example: When τ = 60, it means that for every 60s increase in the time interval between log event i and log event j, the value of the penalty parameter decays to 1 / e of the original value. represents the type matching function, which is used to impose a penalty when log event i and log event j are of the same type.
[0133] Therefore, the type matching function is defined as follows:
[0134] (6)
[0135] In the embodiments of the present application, the target information is calculated using formula (6).
[0136] It is understandable that the elements in the penalty matrix are used to measure the temporal proximity between log event i and log event j. When the time interval between log event i and log event j is smaller, the value of the penalty parameter is larger and closer to 1, then a strong penalty is imposed. As the time interval increases, the value of the penalty parameter decays exponentially.
[0137] By using the elements in the penalty matrix to measure the temporal proximity between different log events, the value of the penalty parameter can be dynamically determined, further improving the matching degree of the penalty parameter with the requirements of the actual detection scenario and further enhancing the detection accuracy.
[0138] Although introducing the time interval between log events into the position encoding and introducing a penalty factor in the attention calculation process can improve the detection accuracy of the target model. However, since the detection accuracy of the target model is highly dependent on the training samples, when facing a relatively complex comprehensive fault mode, multiple abnormal types that are completely unrelated to each other may be recognized, and it is difficult to perform root cause analysis of server anomalies.
[0139] In view of this, the embodiments of this application introduce a pre-trained language model and utilize the causal reasoning ability of the pre-trained language model to perform in-depth root cause analysis on the log sequence to be detected, further improving the accuracy of the detection results.
[0140] The following combines Figure 6 Embodiments for analyzing the log sequence to be detected using a pre-trained language model are described in detail.
[0141] In some embodiments, the pre-trained language model includes: a first pre-trained language model for performing conflict identification and multiple second pre-trained language models for performing logical analysis.
[0142] In the embodiments of this application, in response to determining that the first detection result includes at least two abnormal types with a correlation degree less than a predetermined threshold, using the pre-trained language model to analyze the log sequence to be detected and generate a second detection result may include the following operations: Based on the types of the logs to be detected in the log sequence to be detected, use the corresponding second pre-trained language models for each type to perform logical analysis on the logs to be detected respectively, generating multiple initial analysis results; use the first pre-trained language model to perform conflict identification on the multiple initial analysis results; in the case of determining that there is no conflict between the multiple initial analysis results, generate a second detection result for fusing the multiple initial analysis results.
[0143] In the embodiments of the present application, the log sequence to be detected may include multiple types of logs. For example: system logs responsible for analyzing the operating system and kernel, network logs for analyzing network devices and communication data, storage logs for monitoring the running status of the storage system, etc.
[0144] Since calling a pre-trained language model requires a large amount of computing resources, in order to reduce the redundant data input to the pre-trained language model, data screening rules can be configured first to filter out logs without abnormal fields and normal logs for periodic health checks or heartbeat signals, thereby reducing the resource consumption of the pre-trained language model and further improving the detection speed.
[0145] Figure 6 The figure shows a schematic diagram of collaborative analysis of the log sequence to be detected by calling a pre-trained language model according to the embodiments of the present application.
[0146] As Figure 6 shown, the log sequence 302 to be detected may include system logs 3021, network logs 3022, and storage logs 3033. Then, according to the type of the log to be detected, the system logs 3021, network logs 3022, and storage logs 3033 can be respectively assigned to the corresponding second pre-trained language models for log analysis.
[0147] For example: The system logs 3021 are assigned to LLM3 610 for log analysis to generate an initial analysis result of the system logs. The network logs 3022 are assigned to LLM4 620 for log analysis to generate an initial analysis result of the network logs. The storage logs 3033 are assigned to LLM5 630 for log analysis to generate an initial analysis result of the storage logs.
[0148] Then, the initial analysis results of the system logs, the initial analysis results of the network logs, and the initial analysis results of the storage logs are input into LLM0 640 for conflict identification and result fusion.
[0149] For example: The Prompt for inputting into LLM0 640 may include role information, instruction information, the initial analysis results of the system logs, the initial analysis results of the network logs, and the initial analysis results of the storage logs. The role information can be used to indicate the identity information held by the first pre-trained language model during conflict identification and result fusion. For example: Suppose you are a data analyst with rich experience in logical conflict analysis. The instruction information can be used to indicate the operations required to be performed by the first pre-trained language model: conflict identification, information induction, correction of the analysis direction, etc.
[0150] In operation S3301, if the first pre-trained language model's recognition result for multiple initial analysis results indicates no conflict, information induction is performed on the multiple initial analysis results to generate a second detection result 303.
[0151] By leveraging the causal reasoning ability of the pre-trained language model, through identifying the initial analysis results of various types of logs and then performing in-depth root cause analysis on the multiple initial analysis results, the accuracy of the detection result is further improved.
[0152] In some embodiments, using the first pre-trained language model to perform conflict recognition on multiple initial analysis results may include the following operations: obtaining association information associated with the multiple initial analysis results based on a predetermined knowledge graph; and using the first pre-trained language model to perform conflict recognition on the multiple initial analysis results based on the association information.
[0153] In the embodiments of the present application, the predetermined knowledge graph indicates the topological relationship and fault mode between the components configured in the server. When performing conflict recognition, the first pre-trained language model can first query the information associated with the initial analysis result from the predetermined knowledge graph. For example, when the initial analysis result output by the second pre-trained language model for detecting system logs indicates an abnormal increase in CPU usage, the association information queried through the predetermined knowledge graph may include processes, services, and hardware failures that cause high CPU load. Then, based on the queried association information, conflict recognition of the multiple initial analysis results can be performed.
[0154] When the initial analysis result of other types of logs to be detected includes a hardware failure that causes high CPU, it can be determined that there is no conflict among the multiple initial analysis results.
[0155] However, if the hardware failure included in the initial analysis result of other types of logs to be detected does not cause high CPU, it can be determined that there may be a conflict among the multiple initial analysis results.
[0156] Using the predetermined knowledge graph, background information or context information such as the hardware topology structure or fault mode associated with the abnormal log can be queried, so as to perform in-depth analysis between the multiple initial analysis results in combination with the background information or context information to determine whether there is a conflict, further improving the accuracy of the analysis result.
[0157] In addition to the predetermined knowledge graph, the logical reasoning ability of the pre-trained language model can also be used to perform conflict recognition on multiple initial analysis results.
[0158] In some embodiments, using a first pre-trained language model to perform conflict identification on multiple initial analysis results based on association information may include the following operations: using the first pre-trained language model to perform causal reasoning on the multiple initial analysis results based on the association information to generate at least two reasons for the server anomaly; in response to determining that there is an association between the at least two reasons, determining that there is no conflict between the multiple initial analysis results; and in response to determining that there is no association between the at least two reasons, determining that there is a conflict between the multiple initial analysis results.
[0159] In the server anomaly detection scenario, there is not only a correlation between different anomaly types in terms of hardware or modes, but also a causal relationship between different log events.
[0160] Therefore, causal reasoning can also be performed on multiple initial analysis results based on the association information to generate the reasons for the server anomaly. For example: Is the reason for the server anomaly a configuration error or a service performance degradation caused by a hardware failure.
[0161] Since in the server anomaly detection scenario, the reasons for the server anomaly obtained through causal reasoning belong to the deep root causes, it can be understood that what the log events describe are only surface features. For the same period of collected logs to be detected, the root causes obtained should be the same or at least correlated.
[0162] Therefore, when there is an association between at least two analyzed reasons, it can be determined that there is no conflict between the multiple initial analysis results. When there is no association between at least two analyzed reasons, it can be determined that there is a conflict between the multiple initial analysis results.
[0163] Utilize the causal reasoning ability of the pre-trained language model to mine the root causes of abnormal logs, thereby more accurately judging intermittent or complex faults during the server operation.
[0164] The first pre-trained language model can not only allocate different types of logs to be detected to different second pre-trained language models to perform logical analysis operations through task scheduling, but also perform conflict identification on multiple initial analysis results output by the multiple second pre-trained language models, and timely adjust the analysis direction to guide each second pre-trained language model to perform log analysis in the correct analysis direction, reducing the excessive consumption of computing resources by repeatedly invoking the pre-trained language model.
[0165] In view of this, when it is determined that there are conflicts among multiple initial analysis results in the embodiments of the present application, the following operations may further be included: generating a direction to be corrected corresponding to the cause of the anomaly; determining a target initial analysis result associated with the direction to be corrected from among the multiple initial analysis results; using a second pre-trained language model for outputting the target initial analysis result to correct the target initial analysis result based on the direction to be corrected to generate an intermediate analysis result; updating the target initial analysis result among the multiple initial analysis results to the intermediate analysis result to obtain multiple intermediate analysis results; using the first pre-trained language model to perform conflict identification on the multiple intermediate analysis results, and when it is determined that there are no anomalies among the multiple intermediate analysis results, generating a second detection result for fusing the multiple intermediate analysis results.
[0166] As Figure 6 shown, in operation S3301, when it is determined that there are conflicts among multiple initial detection results, the analysis direction is corrected, and multiple second pre-trained language models are rescheduled to perform the analysis task according to the corrected analysis direction. Such an iterative loop is performed until there are no conflicts among the generated analysis results, and a second detection result is generated.
[0167] In some embodiments, when initially performing detection, the matching degree between the direction to be corrected corresponding to the cause of the anomaly generated by the first pre-trained language model and the user intention may be relatively low. Therefore, the number of iterative loops will be relatively high.
[0168] To reduce the number of iterative loops of the pre-trained language model, feedback information for the second detection result may be received; based on the feedback information, the processing strategy of the pre-trained language model is adjusted.
[0169] Since the second detection result information includes the cause of the server anomaly, it may further include the evidence chain information generated by the pre-trained language model to support the above cause. Therefore, relevant personnel can provide feedback by querying the evidence chain information.
[0170] For example: the feedback information may include: whether the detection result is correct, whether key information is missing, whether the causal reasoning process is reasonable, etc.
[0171] Based on the feedback information, the processing strategy of the pre-trained language model in each stage can be adjusted.
[0172] For example: the processing strategy may include at least one of the following: conflict identification strategy, logical analysis strategy, cause analysis strategy, correction strategy.
[0173] The conflict recognition strategy can be a strategy for determining whether there is a conflict between multiple initial detection results. The logical analysis strategy can be a strategy for the second pre-trained language models to perform log analysis. The cause analysis strategy can be an analysis strategy for the first pre-trained language model to perform causal reasoning using a knowledge graph. The correction strategy can be a strategy generated by the first pre-trained language model for correcting the logical analysis direction of the second pre-trained language models.
[0174] Adjust the processing strategy of the pre-trained model based on the feedback information. As the number of detections increases and the usage time of the model grows, the matching degree of these processing strategies with the user's detection intention will also become higher and higher, thereby continuously improving the accuracy of the model for detecting complex faults.
[0175] Figure 7 The flowchart of the target model training method according to an embodiment of the present application is shown.
[0176] As Figure 7 shown, the training method 700 may include operation S710 to operation S730.
[0177] In operation S710, using the initial model, based on the timestamps of the sample logs in the sample log sequence, embed sample position encoding in the sample log sequence to generate the features of the sample logs.
[0178] In operation S720, using the initial model, based on the attention mechanism, fuse the features of any sample log in the sample log sequence with the features of at least two sample logs in the sample log sequences at their respective adjacent moments to generate sample detection results.
[0179] In operation S730, train the initial model based on the target loss function, sample detection results, and sample labels to generate the target model.
[0180] Due to the complex server operating environment, numerous types of anomalies, and the uneven distribution of the number of different types of anomalies or faults, when training a model with a Transformer architecture, it is easy for the imbalance of training samples in terms of categories to cause the imbalance of gradient propagation during the training process, resulting in a relatively low prediction accuracy of the trained model.
[0181] In view of this, the target loss function in the embodiment of the present application includes parameters for dynamically invoking sample class weights and sample recognition difficulty weights.
[0182] For example: The target loss function L can be expressed as follows:
[0183] (7)
[0184] Among them, represents the category imbalance compensation coefficient; , is the number of samples of class c; is the focusing parameter, which can take values from 2 to 5; is the regularization coefficient; is the L2 loss function; represents the predicted probability.
[0185] The L2 loss function can be expressed as follows:
[0186] (8)
[0187] where, represents the sample label of the i-th sample log; represents the sample detection result of the i-th sample log; m represents the number of samples.
[0188] In the embodiment of the present application, when training the initial model using the target loss function, the loss weight of the majority class can be automatically reduced through ; if = 1 indicates that the current sample is a rare fault, then = 1 / (1 + 1) = 0.5. Therefore, the weight of the current sample is significantly increased, enabling the model to focus on the features of the minority class samples.
[0189] Secondly, when training the initial model using the target loss function, by using the (1 - )γ term, the model can focus on the features of difficult samples for learning during the training process. Difficult samples can be, for example, sample logs in a mixed fault mode composed of multiple abnormal types.
[0190] Thirdly, when training the initial model using the target loss function, when the predicted probability → 1, (1 - )γ → 0, and the loss weight decreases. When → 0, (1 - )γ → 1, and the loss weight remains high, thereby enabling the model to continuously optimize the samples near the classification boundary, thus solving the problem of unbalanced gradient propagation caused by uneven sample classes.
[0191] Figure 8 shows a schematic diagram of the target model training method according to the embodiment of the present application.
[0192] As Figure 8 shown, in this embodiment 800, the initial model 810 can be a neural network model with a Transformer architecture. The sample log sequence 801 has the same definition range as the log sequence to be detected in the server anomaly detection method described above, and will not be elaborated here.
[0193] First, input the sample log sequence 801 into the initial model 810, and output the sample detection result 802. The sample detection result 802 represents the predicted anomaly category of the server associated with the sample log sequence. The sample label 803 indicates the anomaly category of the sample log sequence.
[0194] In the embodiment of this application, during the process of using the initial model to process the sample log sequence to generate the sample detection result and the process of using the target model to process the log sequence to be detected to generate the first detection result in the server anomaly detection method described above, the involved position encoding and feature fusion operations are the same, and will not be elaborated here.
[0195] In the embodiment of this application, using the initial model, based on the timestamps of the sample logs in the sample log sequence, embedding position encoding in the sample log sequence to generate the features of the sample logs may include the following operations: obtaining the first sample time interval between adjacent sample logs in the sample log sequence based on the timestamps of the sample logs in the sample log sequence; generating sample position encoding based on the timestamps of the sample logs and the first sample time interval; and embedding the sample position encoding in the sample log sequence to generate the features of the sample logs.
[0196] For example: The sample position encoding can be calculated based on formula (3) described above, and the sample position encoding is embedded in the sample log sequence to generate the features of the sample logs.
[0197] Dynamically adjust the position encoding based on the time interval of the occurrence of the sample log events, so that while the initial model learns the global timing information of the log events during the training process, it also learns the density of the occurrence of the sample log events, thereby improving the target model's perception ability of the actual occurrence time and density of the log events.
[0198] In the embodiment of this application, using the initial model, based on the attention mechanism, fusing the features of any sample log in the sample log sequence with the features of at least two sample logs in the adjacent sample log sequences to generate the sample detection result may include the following operations: fusing the features of the sample log with the features of at least two adjacent sample logs based on the attention mechanism to generate sample fusion features; and detecting the sample fusion features to generate the sample detection result.
[0199] In some embodiments, based on the attention mechanism, the features of the sample logs are fused with the features of at least two adjacent sample logs to generate sample fusion features, which may include the following operations: obtaining a sample penalty parameter according to the timestamps and sample information of at least two sample logs; where the sample information is used to indicate whether the types of log events indicated by at least two sample logs are the same; and based on the attention mechanism and the sample penalty parameter, fusing the features of the sample logs with the features of at least two adjacent sample logs to generate sample fusion features.
[0200] For example: Feature fusion can be performed based on formula (4) in the server anomaly detection method described above to generate sample fusion features, which will not be elaborated here.
[0201] Due to the repeated occurrence of log events within a short period caused by the triggering of the server automatic regulation mechanism, and the model being aware that the dense occurrence of repeated log events may cause misdetection events, by introducing a penalty parameter into the attention mechanism, the model can learn the characteristics of the repeated occurrence of log events within a short period caused by the server automatic regulation mechanism in the actual application scenario, further improving the model's perception ability of repeated events.
[0202] In the embodiments of the present application, obtaining the sample penalty parameter according to the timestamps and sample information of at least two sample logs may include the following operations: using the initial model, obtaining the second time interval between at least two sample logs according to the timestamps of at least two sample logs; and using the initial model, generating a sample penalty parameter based on the second time interval and the sample information.
[0203] For example: The sample penalty parameter can be calculated based on formula (5) in the server anomaly detection method described above, which will not be elaborated here.
[0204] By using the elements in the penalty matrix to measure the temporal proximity between different sample log events, the model can learn how to dynamically adjust to the penalty parameter that matches the requirements of the sample detection scenario, enabling the trained model to dynamically determine the penalty parameter that matches the actual detection scenario in the actual detection scenario, further improving the model's perception ability of the detection scenario.
[0205] Then, based on the objective loss function shown in formula (7), the loss value 804 can be calculated.
[0206] Next, perform operation S810 to determine whether the loss value 804 converges. If not, adjust the model parameters and continue training. If so, obtain the target model 720 for performing the server anomaly detection method described above.
[0207] In the embodiments of the present application, the loss value convergence condition includes, but is not limited to, the loss value being less than a predetermined loss threshold or reaching the maximum number of training times.
[0208] Based on the above server anomaly detection method, embodiments of the present application further provide a server anomaly detection device, which will be described in detail below in conjunction with Figure 9 for detailed description.
[0209] Figure 9 shows a structural block diagram of a server anomaly detection device according to an embodiment of the present application.
[0210] As Figure 9 shown, the server anomaly detection device 900 may include: a first encoding module 910, a first detection module 920, and an analysis module 930.
[0211] The first encoding module 910 is configured to, in response to determining that the initial detection result of the target server indicates an anomaly, use the target model to embed position encoding in the sequence of logs to be detected based on the timestamps of the logs to be detected in the sequence of logs to be detected, and generate features of the logs to be detected.
[0212] The first detection module 920 is configured to use the target model to fuse the features of any log to be detected with the features of at least two adjacent logs to be detected respectively based on the attention mechanism, and generate a first detection result; wherein, the log event types indicated by the at least two logs to be detected are different.
[0213] The analysis module 930 is configured to, in response to determining that the first detection result includes at least two anomaly types with a correlation degree less than a predetermined threshold, use a pre-trained language model to analyze the sequence of logs to be detected, and generate a second detection result; wherein, the second detection result indicates the server anomaly type and the cause of the server anomaly.
[0214] According to an embodiment of the present application, the first encoding module 910 includes: a first obtaining sub-module, a first generating sub-module, and an embedding sub-module.
[0215] The first obtaining sub-module is configured to obtain a first time interval between adjacent logs to be detected in the sequence of logs to be detected based on the timestamps of the logs to be detected in the sequence of logs to be detected.
[0216] The first generating sub-module is configured to generate position encoding based on the timestamps of the logs to be detected and the first time interval.
[0217] The embedding sub-module is configured to embed the position encoding in the sequence of logs to be detected to generate features of the logs to be detected.
[0218] According to an embodiment of the present application, the first detection module 920 may include: a first fusion sub-module and a first detection sub-module.
[0219] The first fusion sub-module is configured to fuse the features of the log to be detected with the features of at least two adjacent logs to be detected based on the attention mechanism, and generate fused features.
[0220] The first detection sub-module is configured to detect the fused features and generate a first detection result.
[0221] According to an embodiment of the present application, the first fusion sub-module includes: a first acquisition unit and a first fusion unit.
[0222] The first acquisition unit is configured to obtain a penalty parameter according to the timestamps and target information of at least two logs to be detected; wherein, the target information is used to indicate whether the types of the log events indicated by the at least two logs to be detected are the same.
[0223] The first fusion unit is configured to fuse the features of the log to be detected with the features of at least two adjacent logs to be detected based on the attention mechanism and the penalty parameter, and generate fused features.
[0224] According to an embodiment of the present application, the first acquisition unit includes: a first acquisition sub-unit and a first generation sub-unit.
[0225] The first acquisition sub-unit is configured to obtain a second time interval between at least two logs to be detected according to the timestamps of the at least two logs to be detected.
[0226] The first generation sub-unit is configured to generate a penalty parameter based on the second time interval and the target information.
[0227] According to an embodiment of the present application, the pre-trained language model includes: a first pre-trained language model for performing conflict recognition and multiple second pre-trained language models for performing logical analysis; the analysis module 930 may include: a logical analysis sub-module, a conflict recognition sub-module, and a result fusion sub-module.
[0228] The logical analysis sub-module is configured to perform logical analysis on the logs to be detected respectively by using the second pre-trained language models corresponding to the respective types based on the types of the logs to be detected in the log sequence to be detected, and generate multiple initial analysis results.
[0229] The conflict recognition sub-module is configured to perform conflict recognition on the multiple initial analysis results by using the first pre-trained language model.
[0230] The first result fusion sub-module is configured to generate a second detection result for fusing the multiple initial analysis results in the case that no conflict exists between the multiple initial analysis results.
[0231] According to an embodiment of the present application, the analysis module 930 further includes: a direction generation sub-module, a determination sub-module, a correction sub-module, an update sub-module, and a second result fusion sub-module.
[0232] The direction generation sub-module is configured to generate a direction to be corrected corresponding to the cause of the abnormality in the case where a conflict exists between multiple initial analysis results.
[0233] The determination sub-module is configured to determine a target initial analysis result associated with the direction to be corrected from multiple initial analysis results.
[0234] The correction sub-module is configured to use a second pre-trained language model for outputting the target initial analysis result to correct the target initial analysis result based on the direction to be corrected, and generate an intermediate analysis result.
[0235] The update sub-module is configured to update the target initial analysis result among multiple initial analysis results to the intermediate analysis result to obtain multiple intermediate analysis results.
[0236] The second result fusion sub-module is configured to use a first pre-trained language model to perform conflict recognition on multiple intermediate analysis results, and generate a second detection result for fusing multiple intermediate analysis results in the case where no conflict exists between multiple intermediate analysis results.
[0237] According to an embodiment of the present application, the conflict recognition sub-module includes: a first acquisition unit and a first recognition unit.
[0238] The first acquisition unit is configured to obtain association information associated with multiple initial analysis results based on a predetermined knowledge graph; wherein, the predetermined knowledge graph indicates the topological relationship and fault mode between each component configured in the server.
[0239] The first recognition unit is configured to use a first pre-trained language model to perform conflict recognition on multiple initial analysis results based on the association information.
[0240] According to an embodiment of the present application, the first recognition unit includes: an inference sub-unit and a determination sub-unit.
[0241] The inference sub-unit is configured to use a first pre-trained language model to perform causal inference on multiple initial analysis results based on the association information, and generate at least two reasons for causing the server abnormality.
[0242] The determination sub-unit is configured to determine that no conflict exists between multiple initial analysis results in response to determining that there is an association between at least two reasons; and determine that a conflict exists between multiple initial analysis results in response to determining that there is no association between at least two reasons.
[0243] According to an embodiment of the present application, the above device may further include: a receiving module and an adjustment module.
[0244] The receiving module is configured to receive feedback information for the second detection result.
[0245] The adjustment module is configured to adjust the processing strategy of the pre-trained language model based on the feedback information; wherein, the processing strategy includes at least one of the following: a conflict recognition strategy, a logical analysis strategy, a cause analysis strategy, and a correction strategy.
[0246] According to an embodiment of the present application, the above device may further include an acquisition module, a determination module, a decompression module, and an analysis module.
[0247] The acquisition module is configured to acquire a compressed file of the log to be detected.
[0248] The determination module is configured to determine the type of the compressed file based on the content of the compressed file.
[0249] The decompression module is configured to decompress the compressed file by using a first tool called by the pre-trained language model based on the type of the compressed file to generate a plurality of log files.
[0250] The analysis module is configured to analyze the plurality of log files respectively by using a second tool called by the pre-trained language model to generate a log sequence to be detected.
[0251] According to an embodiment of the present application, the decompression module includes: a first determination sub-module, a first generation sub-module, and a decompression sub-module.
[0252] The first determination sub-module is configured to determine the parameter information of the first tool based on the type of the compressed file.
[0253] The first generation sub-module is configured to generate a first parameter for calling the first tool based on the parameter information of the first tool by using the pre-trained language model.
[0254] The decompression sub-module is configured to decompress the compressed file by calling the first tool based on the first parameter to generate a plurality of log files.
[0255] According to an embodiment of the present application, the analysis module includes: a second determination sub-module, a second generation sub-module, and an analysis sub-module.
[0256] According to an embodiment of the present application, the above device further includes: a rule matching module, configured to perform rule matching on the log sequence to be detected based on a predetermined rule library to generate an initial detection result; wherein, the predetermined rule library includes a plurality of predetermined rules for detecting server anomalies; the correlation degrees between the plurality of predetermined rules are less than a predetermined correlation threshold.
[0257] Figure 10The structural block diagram of the target model training device according to an embodiment of the present application is shown.
[0258] As Figure 10 shown, the target model training device 1000 may include: a second encoding module 1010, a second detection module 1020, and a training module 1030.
[0259] The second encoding module 1010 is configured to use the initial model to embed sample position encoding in the sample log sequence based on the timestamps of the sample logs in the sample log sequence, and generate features of the sample logs.
[0260] The second detection module 1020 is configured to use the initial model to fuse the features of any sample log in the sample log sequence with the features of at least two sample logs in the respective adjacent sample log sequences based on the attention mechanism, and generate a sample detection result; the log event types indicated by the at least two sample logs are different.
[0261] The training module 1030 is configured to train the initial model based on the target loss function, the sample detection result, and the sample label to generate a target model.
[0262] According to an embodiment of the present application, the second encoding module 1010 includes: a second obtaining sub-module, a second generating sub-module, and a second embedding sub-module.
[0263] The second obtaining sub-module is configured to obtain a first sample time interval between sample logs adjacent in position in the sample log sequence based on the timestamps of the sample logs in the sample log sequence.
[0264] The second generating sub-module is configured to generate sample position encoding based on the timestamps of the sample logs and the first sample time interval.
[0265] The second embedding sub-module is configured to embed the sample position encoding in the sample log sequence to generate features of the sample logs.
[0266] According to an embodiment of the present application, the second detection module 1020 includes: a second fusion sub-module and a second detection sub-module.
[0267] The second fusion sub-module is configured to fuse the features of the sample logs with the features of at least two respective adjacent sample logs based on the attention mechanism to generate sample fusion features.
[0268] The second detection sub-module is configured to detect the sample fusion features to generate a sample detection result.
[0269] According to an embodiment of the present application, the second fusion sub-module includes: a second obtaining unit and a second fusion unit.
[0270] A second obtaining unit, configured to obtain a sample penalty parameter according to timestamps and sample information of at least two sample logs; wherein, the sample information is used to indicate whether the types of log events indicated by the at least two sample logs are the same.
[0271] A second fusion unit, configured to fuse the features of a sample log with the features of at least two sample logs adjacent thereto based on an attention mechanism and the sample penalty parameter, and generate a sample fusion feature.
[0272] According to an embodiment of the present application, the second fusion unit includes: a second obtaining subunit and a second generating subunit.
[0273] The second obtaining subunit is configured to use an initial model to obtain a second time interval between at least two sample logs according to the timestamps of the at least two sample logs.
[0274] The second generating subunit is configured to use the initial model to generate a sample penalty parameter based on the second time interval and the sample information.
[0275] According to an embodiment of the present application, any of the first encoding module 910, the first detection module 920, and the analysis module 930, or any of the second encoding module 1010, the second detection module 1020, and the training module 1030 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present application, at least one of the first encoding module 910, the first detection module 920, and the analysis module 930, or at least one of the second encoding module 1010, the second detection module 1020, and the training module 1030 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Or, at least one of the first encoding module 910, the first detection module 920, and the analysis module 930, or at least one of the second encoding module 1010, the second detection module 1020, and the training module 1030 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0276] Figure 11 A block diagram of an electronic device suitable for implementing a server anomaly detection method or a training method of a target model according to an embodiment of the present application is shown.
[0277] As shown Figure 11 in FIG. 1, the electronic device 1100 according to an embodiment of the present application includes a processor 1101, which can perform various appropriate actions and processes according to a program stored in the ROM 1102 or a program loaded from the storage section 1108 into the RAM 1103. The processor 1101 may include, for example, a general-purpose microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1101 may also include on-board memory for caching purposes. The processor 1101 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application. The ROM is a read-only memory, and the RAM is a random access memory.
[0278] In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are stored. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. The processor 1101 performs various operations of the method flow according to an embodiment of the present application by executing the programs in the ROM 1102 and / or the RAM 1103. It should be noted that the program may also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 may also perform various operations of the method flow according to an embodiment of the present application by executing the programs stored in the one or more memories.
[0279] According to an embodiment of the present application, the electronic device 1100 may further include an input / output (I / O) interface 1105, and the I / O interface 1105 is also connected to the bus 1104. The electronic device 1100 may further include one or more of the following components connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the input / output (I / O) interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1110 as needed, so that a computer program read from it can be installed into the storage section 1108 as needed.
[0280] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present application are implemented.
[0281] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include one or more memories other than the above-described ROM 902 and / or RAM 1103 and / or ROM1102 and RAM1103.
[0282] An embodiment of the present application also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the server anomaly detection method or the training method of the target model provided by the embodiments of the present application.
[0283] When the computer program is executed by the processor 1101, the above functions defined in the system / apparatus of the embodiments of the present application are executed. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0284] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 909, and / or installed from the removable medium 1111. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0285] In such an embodiment, the computer program may be downloaded and installed from a network through the communication part 1109, and / or installed from the removable medium 1111. When the computer program is executed by the processor 1101, the above functions defined in the system of the embodiments of the present application are executed. According to the embodiments of the present application, the systems, devices, apparatuses, modules, units, etc. described above may be implemented by computer program modules.
[0286] According to the embodiments of the present application, the program code for executing the computer program provided by the embodiments of the present application may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (for example, by connecting through the Internet using an Internet service provider).
[0287] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0288] Those skilled in the art can understand that the features described in the various embodiments of the present application may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application may be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.
[0289] The embodiments of the present application have been described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present application. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all of these substitutions and modifications should fall within the scope of the present application.
Claims
1. A server anomaly detection method, characterized in that, The method includes: In response to determining that the initial detection result of the target server indicates an anomaly, using the target model, based on the timestamps of the logs to be detected in the log sequence to be detected, embedding position encoding in the log sequence to be detected, and generating features of the logs to be detected; Using the target model, based on the attention mechanism, fusing the features of any log to be detected with the features of at least two adjacent logs to be detected respectively, and generating a first detection result; wherein, the types of log events indicated by the at least two logs to be detected are different; and In response to determining that the first detection result includes at least two anomaly types with a correlation degree less than a predetermined threshold, using a pre-trained language model to analyze the log sequence to be detected, and generating a second detection result; wherein, the second detection result indicates the anomaly type of the server and the cause of the server anomaly.
2. The method according to claim 1, wherein The step of using the target model, based on the timestamps of the logs to be detected in the log sequence to be detected, embedding position encoding in the log sequence to be detected, and generating features of the logs to be detected includes: Based on the timestamps of the logs to be detected in the log sequence to be detected, obtaining a first time interval between the logs to be detected that are adjacent in position in the log sequence to be detected; Based on the timestamps of the logs to be detected and the first time interval, generating the position encoding; and Embedding the position encoding in the log sequence to be detected, and generating features of the logs to be detected.
3. The method according to claim 1, characterized in that, The step of using the target model, based on the attention mechanism, fusing the features of any log to be detected with the features of at least two adjacent logs to be detected respectively, and generating a first detection result includes: Based on the attention mechanism, fusing the features of the logs to be detected with the features of the at least two adjacent logs to be detected respectively, and generating a fused feature; and Detecting the fused feature to generate the first detection result.
4. The method according to claim 3, wherein The step of, based on the attention mechanism, fusing the features of the logs to be detected with the features of the at least two adjacent logs to be detected respectively, and generating a fused feature includes: According to the timestamps of at least two logs to be detected and target information, obtaining a penalty parameter; wherein, the target information is used to indicate whether the types of log events indicated by the at least two logs to be detected are the same; and Based on the attention mechanism and the penalty parameter, fusing the features of the logs to be detected with the features of the at least two adjacent logs to be detected respectively, and generating the fused feature.
5. The method according to claim 4, wherein The step of, according to the timestamps of at least two logs to be detected and target information, obtaining a penalty parameter includes: According to the timestamps of at least two logs to be detected, obtaining a second time interval between the at least two logs to be detected; and Based on the second time interval and the target information, generating the penalty parameter.
6. The method according to claim 1, wherein The pre-trained language model includes: a first pre-trained language model for performing conflict identification and a plurality of second pre-trained language models for performing logical analysis; In response to determining that the first detection result includes at least two abnormal types with a correlation degree less than a predetermined threshold, analyzing the to-be-detected log sequence by using a pre-trained language model to generate a second detection result, including: Based on the types of the to-be-detected logs in the to-be-detected log sequence, respectively performing logical analysis on the to-be-detected logs by using second pre-trained language models corresponding to the respective types to generate a plurality of initial analysis results; Using the first pre-trained language model to identify conflicts among the plurality of initial analysis results; and In the case where it is determined that there are no conflicts among the plurality of initial analysis results, generating a second detection result for fusing the plurality of initial analysis results.
7. The method according to claim 6, characterized in that, In response to determining that the first detection result includes at least two abnormal types with a correlation degree less than a predetermined threshold, analyzing the to-be-detected log sequence by using a pre-trained language model to generate a second detection result, further including: In the case where it is determined that there are conflicts among the plurality of initial analysis results, generating a to-be-corrected direction corresponding to the abnormal cause; Determining a target initial analysis result associated with the to-be-corrected direction from the plurality of initial analysis results; Using the second pre-trained language model for outputting the target initial analysis result to correct the target initial analysis result based on the to-be-corrected direction to generate an intermediate analysis result; Updating the target initial analysis result in the plurality of initial analysis results to the intermediate analysis result to obtain a plurality of intermediate analysis results; and Using the first pre-trained language model to identify conflicts among the plurality of intermediate analysis results, and in the case where it is determined that there are no conflicts among the plurality of intermediate analysis results, generating a second detection result for fusing the plurality of intermediate analysis results.
8. The method according to claim 6, wherein The identifying conflicts among the plurality of initial analysis results by using the first pre-trained language model includes: Based on a predetermined knowledge graph, obtaining association information associated with the plurality of initial analysis results; wherein the predetermined knowledge graph indicates the topological relationship and failure modes among the components configured in the server; and Using the first pre-trained language model to identify conflicts among the plurality of initial analysis results based on the association information.
9. The method according to claim 8, characterized in that The identifying conflicts among the plurality of initial analysis results by using the first pre-trained language model based on the association information includes: Using the first pre-trained language model to perform causal reasoning on the plurality of initial analysis results based on the association information to generate at least two causes for the server anomaly; In response to determining that there is an association between the at least two causes, determining that there are no conflicts among the plurality of initial analysis results; and In response to determining that there is no association between the at least two causes, determining that there are conflicts among the plurality of initial analysis results.
10. The method according to any one of claims 6-9, characterized in that, The method further includes: Receiving feedback information for the second detection result; and Adjust the processing strategy of the pre-trained language model based on the feedback information; wherein, the processing strategy includes at least one of the following: conflict recognition strategy, logical analysis strategy, cause analysis strategy, and correction strategy.
11. The method according to claim 1, wherein The method further includes: Obtain a compressed file of the log to be detected; Determine the type of the compressed file based on the content of the compressed file; Based on the type of the compressed file, use the pre-trained language model to call a first tool to decompress the compressed file, generating multiple log files; and Use the pre-trained language model to call a second tool to parse the multiple log files respectively, generating the log sequence to be detected.
12. The method according to claim 11, wherein The using the pre-trained language model to call a first tool to decompress the compressed file based on the type of the compressed file, generating multiple log files includes: Determine the parameter information of the first tool based on the type of the compressed file; Use the pre-trained language model to generate a first parameter for calling the first tool based on the parameter information of the first tool; and Call the first tool to decompress the compressed file based on the first parameter, generating multiple log files.
13. The method according to claim 11, wherein The using the pre-trained language model to call a second tool to parse the multiple log files respectively, generating the log sequence to be detected includes: Determine the parameter information of the second tool based on the parsing requirements of the log to be detected; Use the pre-trained language model to generate a second parameter for calling the second tool based on the parameter information of the second tool; and Call the second tool to parse the multiple log files based on the second parameter, generating the log sequence to be detected.
14. The method according to claim 1, characterized in that, The method further includes: Perform rule matching on the log sequence to be detected based on a predetermined rule library, generating the initial detection result; Wherein, the predetermined rule library includes multiple predetermined rules for detecting server anomalies; the correlation degrees between the multiple predetermined rules are less than a predetermined correlation threshold.
15. A training method for a target model, including: Using an initial model, based on the timestamps of the sample logs in the sample log sequence, embed sample position encodings in the sample log sequence, generating features of the sample logs; Using the initial model, based on an attention mechanism, fuse the features of any sample log in the sample log sequence with the features of at least two sample logs in the respective adjacent sample log sequences, generating a sample detection result; the log event types indicated by the at least two sample logs are different; Train the initial model based on a target loss function, the sample detection result, and sample labels, generating the target model as described in any one of claims 1 to 14; Wherein, the target loss function includes parameters for dynamically calling sample class weights and sample recognition difficulty weights; the sample labels indicate the anomaly categories of the sample log sequences.
16. The method according to claim 15, wherein The using the initial model, based on the timestamps of the sample logs in the sample log sequence, embed sample position encodings in the sample log sequence, generating features of the sample logs includes: Based on the timestamps of the sample logs in the sample log sequence, obtain a first sample time interval between adjacent sample logs in the sample log sequence; Generate the sample position encoding based on the timestamps of the sample logs and the first sample time interval; and Embed the sample position encoding in the sample log sequence to generate the features of the sample logs.
17. The method according to claim 15, wherein The step of using the initial model to generate a sample detection result by fusing the features of any sample log in the sample log sequence with the features of at least two sample logs in the respective adjacent sample log sequences based on the attention mechanism includes: Fuse the features of the sample log with the features of the at least two sample logs in the respective adjacent sample log sequences based on the attention mechanism to generate a sample fusion feature; and Detect the sample fusion feature to generate the sample detection result.
18. The method according to claim 17, characterized in that The step of fusing the features of the sample log with the features of the at least two sample logs in the respective adjacent sample log sequences based on the attention mechanism to generate a sample fusion feature includes: Obtain a sample penalty parameter according to the timestamps and sample information of at least two sample logs; wherein the sample information is used to indicate whether the types of the log events indicated by the at least two sample logs are the same; and Based on the attention mechanism and the sample penalty parameter, fuse the features of the sample log with the features of the at least two sample logs in the respective adjacent sample log sequences to generate the sample fusion feature.
19. The method according to claim 18, wherein The step of obtaining a sample penalty parameter according to the timestamps and sample information of at least two sample logs includes: Use the initial model to obtain a second time interval between at least two sample logs according to the timestamps of the at least two sample logs; and Use the initial model to generate the sample penalty parameter based on the second time interval and the sample information.
20. An electronic device, comprising: One or more processors; A memory for storing one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Fault prediction method and device, electronic equipment and storage medium
CN115328753A
Log detection method and device, electronic equipment and medium
CN115600607A
Service operation abnormity reason detection method and device, equipment and storage medium
CN117762678A