A model training, fault prediction method, device, medium and program product

Through multimodal fusion and knowledge distillation methods, a fault prediction model is built, which solves the problem of low fault detection efficiency and accuracy in microservice systems, and achieves efficient and accurate fault detection and positioning.

CN119204149BActive Publication Date: 2025-05-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411721322.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-13
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The low efficiency and accuracy of fault detection in microservice systems lead to frequent false alarms and missed alarms, and faults may propagate cascade along the call relationship, making it difficult to achieve accurate fault location.

Method used

By obtaining the operation data, log statistics and call delay information of the microservice system, the data processing module and teacher and student module are used to perform multimodal fusion and knowledge distillation to build a fault prediction model to improve the efficiency and accuracy of fault detection.

Benefits of technology

It realizes efficient detection and accurate prediction of microservice system failures, reduces false alarms and missed reports, simplifies the model structure and application process, and improves the comprehensiveness and accuracy of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204149B_ABST
    Figure CN119204149B_ABST
Patent Text Reader

Abstract

The present application discloses a model training, fault prediction method, device, medium and program product in the field of computer technology. The present application obtains the fusion information corresponding to each microservice by multimodal fusion of the operation data, log statistics data and call delay information of each microservice in the microservice system, and then trains the teacher module with the multimodal fusion information and the call graph structure formed by each microservice, which can capture the commonalities and differences between different modal data, and also allows complex call relationships to penetrate into the fault prediction process; trains the student module with multimodal fusion information, which provides a basis for realizing efficient fault detection; completes effective knowledge distillation through comparative learning in loss calculation, and finally obtains a fault prediction model including a data processing module, a student module and a fault prediction module, which can integrate the call relationship and multimodal data between each microservice, and improve the fault detection efficiency and accuracy of the microservice system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a model training, fault prediction method, device, medium and program product. Background Art

[0002] A microservice system includes many independent microservices (i.e., microservice instances), which form a complete application. Generally, a microservice occupies a container exclusively, and there are call relationships between different microservices. Due to the large number of microservices in a microservice system and the complex call relationships, the microservice system has various faults, and the causes and manifestations of faults are also different. There are often false positives and false negatives of faults; and faults may cascade along the call relationships, making it difficult to accurately locate faults.

[0003] Therefore, how to improve the fault detection efficiency and accuracy of microservice systems is a problem that technical personnel in this field need to solve. Summary of the invention

[0004] In view of this, the purpose of this application is to provide a model training, fault prediction method, device, medium and program product to improve the fault detection efficiency and accuracy of the microservice system. The specific scheme is as follows:

[0005] In a first aspect, the present application provides a model training method, comprising:

[0006] Obtain the operation data, log statistics and call delay information of each microservice included in the microservice system;

[0007] Use the data processing module to process the operation data, log statistics and call delay information of each microservice to obtain the fusion information corresponding to each microservice;

[0008] Input the call graph structure formed by each microservice and the fusion information corresponding to each microservice into the teacher module, so that the teacher module outputs the first state information containing graph structure knowledge for representing the operation status of each microservice for each microservice;

[0009] Inputting the fusion information corresponding to each microservice into the student module, so that the student module outputs the second state information without graph structure knowledge for representing the operation status of each microservice for each microservice;

[0010] Processing the first state information using a fault prediction module to predict fault prediction results corresponding to each microservice;

[0011] The comprehensive loss is calculated by using the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result;

[0012] If the comprehensive loss does not meet the preset model convergence conditions, the module parameters of the data processing module, the teacher module, the student module and the fault prediction module are iteratively updated using the comprehensive loss, and when the comprehensive loss meets the model convergence conditions, the current data processing module, the current student module and the current fault prediction module are constructed into a fault prediction model.

[0013] Optionally, the operation data, log statistics data, and call delay information of each microservice included in the microservice system are obtained, including:

[0014] Collect and record the computing resource usage information of each microservice at preset intervals to obtain the operation data of each microservice;

[0015] Count the occurrence frequency of different log templates in each microservice to obtain log statistics in each microservice;

[0016] Determine the delay information of each microservice calling other microservices within a period of time, and obtain the call delay information in each microservice.

[0017] Optionally, count the occurrence frequencies of different log templates in each microservice to obtain log statistics in each microservice, including:

[0018] Execute for each microservice: determine each fixed log template according to the original log data in the current microservice; count the frequency of occurrence of each log template in the original log data over a period of time to obtain log statistics data in the current microservice.

[0019] Optionally, before counting the occurrence frequency of each log template in the original log data within a period of time, the method further includes:

[0020] Discard the first N log templates that appear most frequently among all log templates.

[0021] Optionally, the data processing module includes: a running data processing submodule, a log processing submodule, a call delay processing submodule and a fusion submodule;

[0022] Accordingly, the data processing module is used to process the operation data, log statistics data and call delay information of each microservice to obtain the fusion information corresponding to each microservice, including:

[0023] Using the data processing submodule to convert the operation data of each microservice into a corresponding first embedding vector;

[0024] The log processing submodule is used to convert the log statistics of each microservice into a corresponding second embedding vector;

[0025] The call delay processing submodule is used to convert the call delay information of each microservice into a corresponding third embedding vector;

[0026] The fusion submodule is used to multiply the first embedding vector, the second embedding vector and the third embedding vector of each microservice to obtain fusion information corresponding to each microservice.

[0027] Optionally, the fusion information loss of the fusion information corresponding to each microservice is calculated according to the first formula;

[0028] Among them, the first formula is: ; represents the fusion information loss, Denotes the first embedding vector The transposed matrix of Denotes the second embedding vector The transposed matrix of represents the third embedding vector, Represents the preset hyperparameters.

[0029] Optionally, the state information loss between the first state information and the second state information is calculated according to a second formula;

[0030] Wherein, the second formula is: ; Indicates that the state information is lost, represents the second state information, represents the first state information, Represents the matrix norm.

[0031] Optionally, the prediction result loss of the fault prediction result is calculated according to a third formula;

[0032] Wherein, the third formula is: ; represents the prediction result loss, Indicates the fault prediction result With fault label The supervision loss between is the number of root cause instances output by the model, is the number of fault-free instances output by the model, which is optimized by the constraint .

[0033] Optionally, the comprehensive loss is calculated through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result, including:

[0034] The fusion information loss, the state information loss and the prediction result loss are weightedly fused to obtain the comprehensive loss.

[0035] Optionally, the comprehensive loss is calculated through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result, including:

[0036] The average value of the fusion information loss, the state information loss and the prediction result loss is calculated, and the average value is used as the comprehensive loss.

[0037] Optionally, it also includes:

[0038] The fault prediction model is deployed in each microservice.

[0039] In a second aspect, the present application provides a fault prediction method, which is applied to each microservice included in the microservice system, including:

[0040] Use the fault prediction model deployed in itself to predict faults and issue alarms in real time;

[0041] Wherein, the fault prediction model is obtained based on any of the model training methods described above.

[0042] In a third aspect, the present application provides an electronic device, including:

[0043] Memory for storing computer programs;

[0044] The processor is used to execute the computer program to implement the method disclosed above.

[0045] In a fourth aspect, the present application provides a non-volatile storage medium for storing a computer program, wherein the computer program implements the aforementioned disclosed method when executed by a processor.

[0046] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instructions, which implement the steps of the aforementioned disclosed method when executed by a processor.

[0047] It can be seen that the beneficial effects of the present application are as follows: based on the operation data, log statistics and call delay information of each microservice in the microservice system, the fusion information corresponding to each microservice is obtained through the multimodal fusion of these three types of information, and then the teacher module is trained with the call graph structure formed by the multimodal fusion information and each microservice, which can capture the commonalities and differences between different modal data, and also allow complex call relationships to penetrate into the fault prediction process; the student module is trained with multimodal fusion information, which provides a basis for achieving efficient fault detection; in the loss calculation, the output results of the teacher module and the output results of the student module are compared, so that effective knowledge distillation is completed through comparative learning, which helps to prevent the teacher module from overfitting. When the comprehensive loss meets the model convergence conditions, the current data processing module, the current student module and the current fault prediction module are constructed as a fault prediction model, which simplifies the model structure and the model application process, so that the fault prediction model can integrate the call relationship and multimodal data between each microservice, and improve the fault detection efficiency and accuracy of the microservice system.

[0048] Correspondingly, the model training, fault prediction device, equipment, medium and program product provided by the present application also have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0050] Figure 1 A flow chart of a model training method disclosed in this application;

[0051] Figure 2 A schematic diagram of a solution for efficiently detecting microservice system anomalies disclosed in this application;

[0052] Figure 3 A schematic diagram of a model training device disclosed in this application;

[0053] Figure 4 A schematic diagram of a fault prediction device disclosed in this application;

[0054] Figure 5 A schematic diagram of an electronic device disclosed in this application;

[0055] Figure 6 A server structure diagram provided for this application;

[0056] Figure 7 A terminal structure diagram provided for this application. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other examples obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0058] At present, there are many microservices in the microservice system and the calling relationships are complex, which leads to various faults in the microservice system, and the fault causes and fault manifestations are also different. There are often false alarms and missed alarms; and the faults may be cascaded along the calling relationship, making it difficult to accurately locate the fault. To this end, the present application provides a model training and fault prediction solution, which enables the fault prediction model to integrate the calling relationships and multimodal data between microservices, and improve the fault detection efficiency and accuracy of the microservice system.

[0059] See also Figure 1 As shown, the embodiment of the present application discloses a model training method, including:

[0060] S101. Obtain operation data, log statistics data, and call delay information of each microservice included in the microservice system.

[0061] For microservice systems, the observability of the system is generally reflected in three modal system observation data, including: metrics (corresponding to operating data), logs (corresponding to log statistics) and calls (trace, corresponding to call delay information). Metrics include data on changes in the system's CPU (Central Processing Unit) usage, memory occupancy, and other changes, expressed as multivariate time series. Logs are information automatically recorded by the system set by developers. A log consists of timestamps, log levels, log templates, and log variables, and is expressed as semi-structured text data. Call data records the mutual call relationships of all service instances contained in an application service, and consists of several call records, expressed as a call tree structure.

[0062] It should be noted that the observation data of the three modes can reflect the abnormality of the microservice system to a certain extent. Taking indicator data as an example, if the network traffic data suddenly drops, it indicates that the network port of the corresponding microservice instance may be faulty; taking log data as an example, when the log level is an error or the log template is an abnormal log template, it means that the corresponding microservice instance has a corresponding log fault; taking call data as an example, when the delay time of a call chain exceeds the expected value, it means that the relevant microservice instance may be abnormal.

[0063] In one embodiment, the operation data, log statistics and call delay information of each microservice included in the microservice system are obtained, including: collecting and recording the computing resource occupancy information (such as CPU usage, bandwidth occupancy and / or memory occupancy) of each microservice at preset intervals to obtain the operation data of each microservice; counting the frequency of occurrence of different log templates in each microservice to obtain log statistics in each microservice; determining the delay information of each microservice calling other microservices within a period of time to obtain the call delay information of each microservice. It should be noted that the operation data, log statistics and call delay information are all uniformly processed according to the corresponding standardized methods.

[0064] In one embodiment, the frequency of occurrence of different log templates in each microservice is counted to obtain log statistics in each microservice, including: performing the following steps for each microservice: determining each fixed log template according to the original log data in the current microservice; counting the frequency of occurrence of each log template in the original log data over a period of time, and obtaining log statistics in the current microservice. Wherein, before counting the frequency of occurrence of each log template in the original log data over a period of time, it also includes: discarding the first N log templates with the highest frequency of occurrence in each log template. For example: from the original log data in microservice 1, three fixed log templates are determined, which are arranged from high to low according to the frequency of occurrence: A, B, and C. If data redundancy is not considered, a frequency sequence can be obtained for log templates A, B, and C, respectively, and three frequency sequences can be obtained in microservice 1 as log statistics in microservice 1; but if it is considered that most of the frequently occurring logs are normal log data, then log template A can be discarded (assuming that N is 1), and only log templates B and C are retained. Then, for log templates B and C, two frequency sequences can be obtained as log statistics in microservice 1. That is to say: one log template corresponds to one occurrence frequency sequence, and the elements in the occurrence frequency sequence corresponding to log template A are: the frequency of occurrence of log template A at each statistical time point within the statistical time period.

[0065] It should be noted that the number of occurrence frequency sequences obtained in different microservices may be the same or different.

[0066] S102: Use the data processing module to process the operation data, log statistics data and call delay information of each microservice to obtain the fusion information corresponding to each microservice.

[0067] Step S102 is used to realize the mapping and fusion of each modal information to the embedding vector. In one embodiment, the data processing module includes: an operation data processing submodule, a log processing submodule, a call delay processing submodule and a fusion submodule; accordingly, the operation data, log statistics and call delay information of each microservice are processed by the data processing module to obtain the fusion information corresponding to each microservice, including: using the data processing submodule to convert the operation data of each microservice into a corresponding first embedding vector; using the log processing submodule to convert the log statistics of each microservice into a corresponding second embedding vector; using the call delay processing submodule to convert the call delay information of each microservice into a corresponding third embedding vector; using the fusion submodule to multiply the first embedding vector, the second embedding vector and the third embedding vector of each microservice to obtain the fusion information corresponding to each microservice.

[0068] Among them, the running data processing submodule, the log processing submodule, and the call delay processing submodule can all be implemented based on the Temporal Convolution Networks (TCN) and the self-attention mechanism, but the running data processing submodule, the log processing submodule, and the call delay processing submodule are independent of each other. The fusion submodule is used to multiply the first embedding vector, the second embedding vector, and the third embedding vector corresponding to the same microservice, so that each microservice finally corresponds to a fusion information.

[0069] S103. Input the call graph structure formed by each microservice and the fusion information corresponding to each microservice into the teacher module, so that the teacher module outputs the first state information containing graph structure knowledge for representing the operation status of each microservice for each microservice.

[0070] S104. Input the fusion information corresponding to each microservice into the student module, so that the student module outputs the second state information without graph structure knowledge for representing the operation status of each microservice.

[0071] S105: Use the fault prediction module to process the first state information to predict the fault prediction results corresponding to each microservice.

[0072] In this embodiment, the teacher module can be specifically a GAT (Graph Attention Network) model, the student module can be specifically an MLP (Multi-Layer Perceptron) model, and the fault prediction module can be specifically any neural network model.

[0073] S106. Calculate the comprehensive loss through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result.

[0074] In one implementation, the fusion information loss of the fusion information corresponding to each microservice is calculated according to a first formula; wherein the first formula is: ; represents the fusion information loss, Represents the first embedding vector The transposed matrix of represents the second embedding vector The transposed matrix of represents the third embedding vector, Represents the preset hyperparameters.

[0075] In one implementation, the state information loss between the first state information and the second state information is calculated according to a second formula; wherein the second formula is: ; Indicates loss of state information, Indicates the second state information, Indicates the first state information, Represents the matrix norm.

[0076] In one implementation, the prediction result loss of the fault prediction result is calculated according to a third formula; wherein the third formula is: ; Represents the prediction result loss, Indicates the fault prediction result With fault label The supervision loss between is the number of root cause instances output by the model, is the number of fault-free instances output by the model, which is optimized by the constraint The number of instances is the number of microservices.

[0077] In one embodiment, a comprehensive loss is calculated through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result, including: weighted fusion of the fusion information loss, the state information loss and the prediction result loss to obtain the comprehensive loss.

[0078] In one embodiment, a comprehensive loss is calculated through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result, including: calculating the average value of the fusion information loss, the state information loss, and the prediction result loss, and taking the average value as the comprehensive loss.

[0079] S107. If the comprehensive loss does not meet the preset model convergence conditions, the module parameters of the data processing module, the teacher module, the student module and the fault prediction module are iteratively updated using the comprehensive loss, and when the comprehensive loss meets the model convergence conditions, the current data processing module, the current student module and the current fault prediction module are constructed into a fault prediction model.

[0080] In this embodiment, after obtaining the fault prediction model, the method further includes: deploying the fault prediction model in each microservice.

[0081] It can be seen that this embodiment is based on the operation data, log statistics and call delay information of each microservice in the microservice system. Through the multimodal fusion of these three types of information, the fusion information corresponding to each microservice is obtained. Then, the teacher module is trained with the multimodal fusion information and the call graph structure formed by each microservice, which can capture the commonalities and differences between different modal data, and also allows complex call relationships to penetrate into the fault prediction process; the student module is trained with multimodal fusion information, which provides a basis for achieving efficient fault detection; in the loss calculation, the output results of the teacher module and the output results of the student module are compared, so that effective knowledge distillation is completed through comparative learning, which helps to prevent the teacher module from overfitting. When the comprehensive loss meets the model convergence conditions, the current data processing module, the current student module and the current fault prediction module are constructed as a fault prediction model, which simplifies the model structure and model application process, so that the fault prediction model can integrate the call relationship and multimodal data between each microservice, and improve the fault detection efficiency and accuracy of the microservice system.

[0082] A fault prediction method provided in an embodiment of the present application is introduced below. The fault prediction method described below can be referenced to other embodiments described in this document.

[0083] An embodiment of the present application discloses a fault prediction method, which is applied to each microservice included in a microservice system, including: using a fault prediction model deployed in itself to perform fault prediction and alarm in real time; wherein the fault prediction model is obtained based on the model training method described in other embodiments.

[0084] In this embodiment, the fault prediction model can integrate the call relationships between microservices and multimodal data to perform fault detection on a single microservice separately, thereby improving the fault detection efficiency and accuracy of the microservice system.

[0085] See also Figure 2 ,An efficient solution for detecting anomalies in microservice systems includes functional modules such as data collection, model building, model training and model deployment.

[0086] Multimodal observation data collection and preprocessing of microservice systems: Specifically, independent and identical data collection mechanisms can be deployed on different microservice instances in the microservice system to collect indicators, logs, and call data within a single service instance. The data collection tools, specific data storage, and transmission methods are not restricted here.

[0087] Indicator data collection and preprocessing: Monitor and collect various indicator data of the system at fixed time intervals, such as CPU usage, bandwidth usage, or memory usage. Furthermore, interpolation is used to reduce the impact of missing data, and then data standardization is used to unify the scale of different indicator data. , and its standardization method is: ,in, Respectively represent indicator data The mean and variance of Represents the observed value of the indicator data at time t.

[0088] In order to reduce the impact of seasonality or cyclicality, the data can be first-order differentiated, which is in the form of Assuming that in a microservice system, each microservice instance can collect m types of indicator data, the indicator data collected and preprocessed by a single microservice instance can be expressed as a multivariate time series ,in is the predefined data collection window length (e.g. one minute). For m types of indicator data collected at a time, collected S times within a period of time, the data can be collected eventually , including S .

[0089] Collection and preprocessing of log data: Use log collection tools to collect business logs generated by each microservice instance in the microservice system and store them in each microservice instance. Use log parsing tools to extract log templates from raw log data. Log templates refer to the fixed parts that remain unchanged when logs are generated. Then further count the frequency of occurrence of each log template and determine the log templates with the highest frequency of occurrence to filter and discard the corresponding raw log data and log templates, thereby reducing the redundancy of log data. The discard ratio is determined according to the specific situation. It is recommended to discard the top 50% of the logs and their templates with the highest frequency of occurrence. Assume that the number of remaining log templates is , the time series of the frequency of occurrence of each log template is expressed as: Specifically, Represents a log template At the moment Frequency of occurrence. After interpolation and normalization, a multivariate time series of log data can be obtained for a single microservice instance in a microservice system: ; The frequency of occurrence of l log templates collected at a time. If S log templates are collected over a period of time, the data can be collected eventually. , including S .

[0090] Collection and preprocessing of call chain data: Distributed tracing tools are used to collect call chain data in the microservice system. Furthermore, the call chain data is preprocessed in two steps: (1) Within a time window, the call chain data is used to construct a call relationship graph between microservice instances in the microservice system. , where V represents the set of microservice instances and E represents the set of edges between microservice instances; if and only if the microservice instance and microservice instances If there is a calling relationship between them, then (2)For a single microservice instance, assume that the number of microservice instances with which it has a scheduling relationship is , then the latency of calling the microservice instance is calculated as ,in Indicates at time Microservice Examples The call latency of calling the microservice instance. The call latency is also preprocessed according to the difference and standardization method described above, so the multivariate time series of the call chain data can be obtained. ; This includes the s delay data of a single microservice instance calling s microservice instances collected at one time, collected within a period of time w, and finally the data can be collected Assume that microservice instance I has a calling relationship with microservice instance E and microservice instance F, and at a time t, microservice instance I calls microservice instance E, and microservice instance F calls microservice instance I at the same time. Then at time t, for microservice instance I, the recorded delay is: the delay of microservice instance I calling microservice instance E; for microservice instance F, the recorded delay is: the delay of microservice instance F calling microservice instance I.

[0091] Multimodal data representation module: The pre-processed data are represented independently, that is, the indicator data are represented separately. Embedding as an indicator representation , log data Embedded as log representation , call chain data Embedding is an embedding representation The dimension of the representation vector is d. The value of d can be optimized according to the specific business data. In this example, .

[0092] Specifically, the following three embedding modules are set for the three modal data:

[0093] Embedding module for indicator data: The embedding of indicator data is realized by using Temporal Convolution Networks (TCN) and Self-Attention (SA). Specifically, for the input indicator data , the TCN model outputs convolution results of the same shape, that is, . SA is used to further aggregate the information of the time dimension, and the final output is Specifically, the SA module calculates: ,in is a learnable parameter of SA.

[0094] Log data embedding module: Thanks to the preprocessing stage, log data is processed into multivariate time series data. Log data is embedded in the same way as the indicator data embedding module. .

[0095] Embedding module for call chain data: Similarly, using the same model architecture to embed call chain data, we can get .

[0096] It should be noted that although the model architectures of the three embedding modules are the same, the model parameters are independent of each other, and are calculated and trained independently without sharing parameters.

[0097] Multimodal contrastive learning and fusion module: Since the information of the three modes is essentially a description of the same system from different angles, which is both shared and exclusive, semi-contrastive learning (SCL) is used to model the cross-nature of information. Specifically, for the multimodal representation within the same time window, , first perform The vector is normalized, and then the following semi-contrastive learning loss function is constructed: ,in is a hyperparameter. The loss function The purpose is to make the representations of the three modes calculated independently in the same time window as similar as possible (similarity is calculated using vector inner product), but retain a certain specificity (using hyperparameters control). Then for the multimodal representation in the same time window , you can get the microservice instance corresponding to the window (for example, microservice instance )’s fusion representation: ,in Indicates multiplication by vector elements. By performing the same processing and operation on different microservice instances in the microservice system, the fusion representation of multimodal data in each microservice instance can be further obtained. For example: For a system with four instances, we can get The four representations are used as fusion representation vectors (i.e., fusion information) of the four instances (microservice instances).

[0098] Teacher model (i.e. teacher module): The microservice system implements the complete application through the interconnected graph structure, which may cause the fault to propagate between microservice instances through the call relationship, making fault detection and location more difficult. This embodiment allows the model to perceive the global interconnected topology. Therefore, the GNN model is used to learn the interconnected structure of the system and the fusion representation information of the microservice instance. Based on the aforementioned microservice instance fusion representation vector and the call relationship graph , we can get the property graph representation of the entire microservice system in a single time window , Represents all the fused representation vectors obtained above. Specifically, the graph attention network GAT (GAT is a special case of GNN) can be used to realize the fusion of topological structure and microservice instance features. Specifically, the graph attention network GAT can be used to calculate ;in Represents a microservice instance In the model The representation vector of the layer, and ,and is a learnable weight matrix, Is a microservice instance The attention coefficient between can be obtained according to the general attention calculation method. is a nonlinear activation function. In each step of GAT, the instance representation of the previous step First, it is used to calculate all the connected instance pairs Attention coefficient ,Then and weight matrix And the coefficient Multiply and sum the results according to the instance adjacency relationship. , the summation result is passed through a nonlinear activation function After that, we get the next instance representation ,go through After several iterations, we can finally get the representation of each microservice instance. , that is, the first state information containing graph structure knowledge ,in Indicates the number of layers of the model used, which can be determined according to the specific problem and data. . That is to say: Input to the GNN model or GAT model, we can get .

[0099] Student model (i.e., student module): The student model abandons the global interconnection structure and only uses the internal information of a single microservice instance for representation learning. The benefits of doing so are: 1) reducing computing requirements, 2) enabling parallel computing; 3) reducing data transmission between microservice instances; 4) crucial for lightweight deployment of the model. However, due to the lack of global interconnection information, fault detection and diagnosis will inevitably be inferior to the teacher model. For this reason, this example gives each microservice instance a unique position encoding so that the student model also has instance perception capabilities; and uses contrastive learning to learn the global from the teacher model. Furthermore, the multilayer perceptron MLP, which serves as a student model, uses a learnable encoding vector as the position encoding of the microservice instance. Let , indicating an instance The learnable position embedding of , then the student model calculates ,in , Represents the vector Add according to the corresponding dimensions, note is a learnable parameter vector, and MLP means that the input vector is sequentially subjected to linear transformation and nonlinear activation function. In other words: Input to the student model, we can get .

[0100] This embodiment transfers the knowledge distillation of the teacher model to the student model, so that the lightweight, local and parallelized student model has global knowledge and fault handling capabilities. To achieve knowledge transfer, Represents the matrix norm.

[0101] This embodiment enables each microservice instance to independently detect and locate faults. The characterization vector is , if fault detection also uses the MLP model structure, then the output of fault detection is: , where MLP is as described above; for the input , the output of SoftMax is .when , it means that the system has a fault, and the fault cause is ; Otherwise, it is determined that the system has no fault. The loss function used for fault detection is: Among them, the first term represents the supervision loss, and the latter two terms constrain the output results There is only one root cause instance. Indicates the fault labels marked in the historical data.

[0102] Specifically, the offline model training process includes: collecting historical labeled data and dividing it into different batches; for each batch, calculating the total loss function: , here we are not limited to averaging the values ​​of the three loss functions, we can set more flexible hyperparameters to add the three loss functions, for example, using learnable weight coefficients. The gradient calculation is performed on the learnable parameters of each model involved above, and the Adam optimizer is used to complete the update of the learnable parameters. All hyperparameters of the model can be tuned according to specific data. Repeat the above steps on all batches until the loss function value L no longer decreases, and then output the trained model architecture and parameters.

[0103] On the basis of the trained model, retain the three embedding modules, fusion module, student model and fault prediction module, that is, remove the two comparative learning modules and the teacher model, and you can get a completely localized fault prediction model. Deploy the fault prediction model to the container where each microservice instance is located as a real-time detection tool for the system. Then, in each microservice instance, regularly collect multimodal data within a fixed time window, and use the fault prediction model to localize the collected multimodal data, including data preprocessing, data fusion, data representation, detection and positioning, and finally get a two-dimensional vector ; Based on the above results, if there is a fault in the microservice instance, an abnormal alarm is generated and the alarm information is provided to the server operation and maintenance administrator.

[0104] This embodiment can capture the commonalities and differences between data of different modalities, make full use of information of different modalities, and realize lightweight model deployment and efficient reasoning; it aligns data representations of different modalities through comparative learning methods, and passes global information to the lightweight student model through knowledge distillation technology, thereby realizing efficient fault detection and location of the microservice system.

[0105] In the data modeling stage, this embodiment characterizes the data of different modes through their respective embedding modules for the three modal information of indicators, logs and call data, uses contrastive learning to improve the information interactivity of different modal representations, and then obtains the overall information representation of the microservice instance through the feature fusion module. In addition, by extracting the call data, the graph structure between the microservice instances is constructed. In the fault detection and location stage, the graph attention network GAT is used as the teacher model to learn the information of the instance itself and the scheduling graph structure between the instances at the same time, and the instance encoding and MLP are used as the student model to learn the information of the instance itself. The knowledge transfer from the teacher model to the student model is realized through contrastive learning, so that the student model also has global information. In the model training stage, the end-to-end training of the model is realized through supervised learning and two contrastive learning losses. In the model deployment and reasoning stage, only efficient and lightweight student models are deployed, which reduces the computational overhead and communication requirements during reasoning, realizes distributed fault detection and location, and improves the real-time and efficiency of the system. Through the integration of contrastive learning and multimodal data, the abnormal performance in the system can be better captured, false positives and false negatives can be reduced, and the comprehensiveness and accuracy of anomaly detection can be improved, thereby improving the accuracy of fault detection and root cause location. This lightweight architecture is suitable for large-scale microservice systems and can significantly reduce computing and storage costs without sacrificing accuracy.

[0106] A model training device provided in an embodiment of the present application is introduced below. The model training device described below can be referenced to other embodiments described in this document.

[0107] See also Figure 3 As shown, the embodiment of the present application discloses a model training device, including:

[0108] The acquisition module 301 is used to obtain the operation data, log statistics data and call delay information of each microservice included in the microservice system;

[0109] The fusion module 302 is used to process the operation data, log statistics data and call delay information of each microservice by using the data processing module to obtain the fusion information corresponding to each microservice;

[0110] The first processing module 303 is used to input the call graph structure formed by each microservice and the fusion information corresponding to each microservice into the teacher module, so that the teacher module outputs the first state information containing graph structure knowledge for representing the operation status of each microservice for each microservice;

[0111] The second processing module 304 is used to input the fusion information corresponding to each microservice into the student module, so that the student module outputs the second state information without graph structure knowledge for representing the operation status of each microservice for each microservice;

[0112] The prediction module 305 is used to process the first state information using the fault prediction module to predict the fault prediction results corresponding to each microservice;

[0113] The calculation module 306 is used to calculate the comprehensive loss through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result;

[0114] The updating module 307 is used to iteratively update the module parameters of the data processing module, the teacher module, the student module and the fault prediction module using the comprehensive loss if the comprehensive loss does not meet the preset model convergence conditions, and to construct the current data processing module, the current student module and the current fault prediction module into a fault prediction model when the comprehensive loss meets the model convergence conditions.

[0115] In one implementation, the acquisition module is specifically used to:

[0116] Collect and record the computing resource usage information of each microservice at preset intervals to obtain the operation data of each microservice;

[0117] Count the occurrence frequency of different log templates in each microservice to obtain log statistics in each microservice;

[0118] Determine the delay information of each microservice calling other microservices within a period of time, and obtain the call delay information in each microservice.

[0119] In one implementation, the acquisition module is specifically used to:

[0120] Execute for each microservice: determine fixed log templates based on the original log data in the current microservice; count the frequency of occurrence of each log template in the original log data over a period of time to obtain log statistics data in the current microservice.

[0121] In one implementation, the acquisition module is specifically used to:

[0122] Discard the first N log templates that appear most frequently among all log templates.

[0123] In one embodiment, the data processing module includes: a running data processing submodule, a log processing submodule, a call delay processing submodule and a fusion submodule;

[0124] Accordingly, the fusion module is specifically used for:

[0125] Using the data processing submodule to convert the operation data of each microservice into a corresponding first embedding vector;

[0126] The log processing submodule is used to convert the log statistics of each microservice into a corresponding second embedding vector;

[0127] The call delay processing submodule is used to convert the call delay information of each microservice into a corresponding third embedding vector;

[0128] The fusion submodule is used to multiply the first embedding vector, the second embedding vector and the third embedding vector of each microservice to obtain fusion information corresponding to each microservice.

[0129] In one embodiment, the computing module is specifically configured to:

[0130] The fusion information loss, state information loss and prediction result loss are weighted and fused to obtain the comprehensive loss.

[0131] In one embodiment, the computing module is specifically configured to:

[0132] Calculate the average of the fusion information loss, state information loss, and prediction result loss, and take the average as the comprehensive loss.

[0133] In one embodiment, it further includes:

[0134] The deployment module is used to deploy the fault prediction model in each microservice.

[0135] Among them, for more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0136] It can be seen that this embodiment provides a fault prediction device, which is based on the operation data, log statistics and call delay information of each microservice in the microservice system, and obtains the fusion information corresponding to each microservice through the multimodal fusion of these three types of information. Then, the teacher module is trained with the multimodal fusion information and the call graph structure formed by each microservice, which can capture the commonalities and differences between different modal data, and also allows complex call relationships to penetrate into the fault prediction process; the student module is trained with multimodal fusion information, which provides a basis for achieving efficient fault detection; the output results of the teacher module and the output results of the student module are compared in the loss calculation, so that effective knowledge distillation is completed through comparative learning, which helps to prevent the teacher module from overfitting. When the comprehensive loss meets the model convergence condition, the current data processing module, the current student module and the current fault prediction module are constructed as a fault prediction model, which simplifies the model structure and model application process, so that the fault prediction model can integrate the call relationship and multimodal data between each microservice, and improve the fault detection efficiency and accuracy of the microservice system.

[0137] A fault prediction device provided in an embodiment of the present application is introduced below. The fault prediction device described below can be referenced to other embodiments described in this document.

[0138] See also Figure 4 As shown, an embodiment of the present application discloses a fault prediction device, which is applied to each microservice included in a microservice system, including: a prediction and alarm module, which is used to perform fault prediction and alarm in real time using a fault prediction model deployed in itself; wherein the fault prediction model is obtained based on the model training method described in other embodiments.

[0139] In this embodiment, the fault prediction model can integrate the call relationships between microservices and multimodal data to perform fault detection on a single microservice separately, thereby improving the fault detection efficiency and accuracy of the microservice system.

[0140] An electronic device provided in an embodiment of the present application is introduced below. The electronic device described below can be referenced to other embodiments described in this document.

[0141] See also Figure 5 As shown, the embodiment of the present application discloses an electronic device, including:

[0142] Memory 501, used for storing computer programs;

[0143] The processor 502 is used to execute the computer program to implement the method disclosed in any of the above embodiments.

[0144] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: obtaining the operating data, log statistics data and call delay information of each microservice included in the microservice system; using the data processing module to process the operating data, log statistics data and call delay information of each microservice to obtain the fusion information corresponding to each microservice; inputting the call graph structure formed by each microservice and the fusion information corresponding to each microservice into the teacher module, so that the teacher module outputs the first state information containing graph structure knowledge for representing the operating status of each microservice for each microservice; inputting the fusion information corresponding to each microservice into the student module, so that the student module outputs the first state information containing graph structure knowledge for each microservice The second state information without graph structure knowledge representing the operation status of each microservice is used; the fault prediction module is used to process the first state information to predict the fault prediction results corresponding to each microservice; the comprehensive loss is calculated through the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result; if the comprehensive loss does not meet the preset model convergence conditions, the module parameters of the data processing module, the teacher module, the student module and the fault prediction module are iteratively updated using the comprehensive loss, and when the comprehensive loss meets the model convergence conditions, the current data processing module, the current student module and the current fault prediction module are constructed into a fault prediction model.

[0145] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: collecting and recording the computing resource occupancy information in each microservice at preset intervals to obtain the operation data in each microservice; counting the frequency of occurrence of different log templates in each microservice to obtain log statistical data in each microservice; determining the delay information of each microservice calling other microservices within a period of time to obtain the call delay information in each microservice.

[0146] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: for each microservice, the following steps are executed: determining each fixed log template based on the original log data in the current microservice; counting the frequency of occurrence of each log template in the original log data over a period of time to obtain log statistical data in the current microservice.

[0147] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: discarding the first N log templates with the highest frequency of occurrence among the log templates.

[0148] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: using a data processing submodule to convert the operating data of each microservice into a corresponding first embedding vector; using a log processing submodule to convert log statistical data of each microservice into a corresponding second embedding vector; using a call delay processing submodule to convert the call delay information of each microservice into a corresponding third embedding vector; using a fusion submodule to multiply the first embedding vector, the second embedding vector and the third embedding vector of each microservice to obtain fusion information corresponding to each microservice.

[0149] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: weighted fusion of fusion information loss, state information loss and prediction result loss to obtain a comprehensive loss.

[0150] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: calculating the average value of the fusion information loss, the state information loss, and the prediction result loss, and taking the average value as the comprehensive loss.

[0151] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: deploying the fault prediction model in each microservice.

[0152] Furthermore, the present application also provides an electronic device. The electronic device can be Figure 6 The server shown can also be Figure 7 Terminal shown. Figure 6 and Figure 7 All of them are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the diagrams cannot be regarded as any limitation on the scope of use of the present application.

[0153] Figure 6 A schematic diagram of the structure of a server provided in an embodiment of the present application. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the relevant steps in the fault prediction disclosed in any of the aforementioned embodiments.

[0154] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0155] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.

[0156] The operating system is used to manage and control the hardware devices and computer programs on the server to realize the operation and processing of the data in the memory by the processor, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to computer programs that can be used to complete the corresponding methods disclosed in any of the aforementioned embodiments, computer programs can also further include computer programs that can be used to complete other specific tasks. In addition to data such as application update information, data can also include data such as application developer information.

[0157] Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of the present application, the terminal may specifically include but is not limited to a smart phone, a tablet computer, a laptop computer or a desktop computer, etc.

[0158] Generally, the terminal in this embodiment includes: a processor and a memory.

[0159] Among them, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0160] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is at least used to store the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the corresponding method performed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, update information of the application.

[0161] In some embodiments, the terminal may also include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0162] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the terminal, and may include more or fewer components than those shown in the figure.

[0163] A non-volatile storage medium provided in an embodiment of the present application is introduced below. The non-volatile storage medium described below can be cross-referenced with other embodiments described herein.

[0164] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the corresponding method disclosed in the above-mentioned embodiment. The non-volatile storage medium is a computer-readable non-volatile storage medium, which, as a carrier for storing resources, may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system, a computer program and data, etc., and the storage method may be temporary storage or permanent storage.

[0165] A computer program product provided in an embodiment of the present application is introduced below. The computer program product described below can be cross-referenced with other embodiments described in this document.

[0166] A computer program product comprises a computer program / instruction, which implements the steps of the corresponding method disclosed above when executed by a processor.

[0167] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0168] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.

[0169] Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A model training method, characterized in that: include: Obtain the operation data, log statistics and call delay information of each microservice included in the microservice system; Use the data processing module to process the operation data, log statistics and call delay information of each microservice to obtain the fusion information corresponding to each microservice; Input the call graph structure formed by each microservice and the fusion information corresponding to each microservice into the teacher module, so that the teacher module outputs the first state information containing graph structure knowledge for representing the operation status of each microservice for each microservice; Inputting the fusion information corresponding to each microservice into the student module, so that the student module outputs the second state information without graph structure knowledge for representing the operation status of each microservice for each microservice; Processing the first state information using a fault prediction module to predict fault prediction results corresponding to each microservice; The comprehensive loss is calculated by using the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result; If the comprehensive loss does not meet the preset model convergence conditions, the module parameters of the data processing module, the teacher module, the student module and the fault prediction module are iteratively updated using the comprehensive loss, and when the comprehensive loss meets the model convergence conditions, the current data processing module, the current student module and the current fault prediction module are constructed into a fault prediction model.

2. The method according to claim 1, characterized in that Obtain the operation data, log statistics, and call delay information of each microservice in the microservice system, including: Collect and record the computing resource usage information of each microservice at preset intervals to obtain the operation data of each microservice; Count the occurrence frequency of different log templates in each microservice to obtain log statistics in each microservice; Determine the delay information of each microservice calling other microservices within a period of time, and obtain the call delay information in each microservice.

3. The method according to claim 2, characterized in that Count the occurrence frequency of different log templates in each microservice and obtain log statistics in each microservice, including: Execute for each microservice: determine each fixed log template according to the original log data in the current microservice; count the frequency of occurrence of each log template in the original log data over a period of time to obtain log statistics data in the current microservice.

4. The method according to claim 3, characterized in that Before counting the occurrence frequency of each log template in the original log data within a period of time, the following is also included: Discard the first N log templates that appear most frequently among all log templates.

5. The method according to claim 1, characterized in that The data processing module includes: a running data processing submodule, a log processing submodule, a call delay processing submodule and a fusion submodule; Accordingly, the data processing module is used to process the operation data, log statistics data and call delay information of each microservice to obtain the fusion information corresponding to each microservice, including: Using the data processing submodule to convert the operation data of each microservice into a corresponding first embedding vector; The log processing submodule is used to convert the log statistics of each microservice into a corresponding second embedding vector; The call delay processing submodule is used to convert the call delay information of each microservice into a corresponding third embedding vector; The fusion submodule is used to multiply the first embedding vector, the second embedding vector and the third embedding vector of each microservice to obtain fusion information corresponding to each microservice.

6. The method according to claim 5, characterized in that Calculate the fusion information loss of the fusion information corresponding to each microservice according to the first formula; Among them, the first formula is: ; represents the fusion information loss, Denotes the first embedding vector The transposed matrix of Denotes the second embedding vector The transposed matrix of represents the third embedding vector, Represents the preset hyperparameters.

7. The method according to claim 1, characterized in that Calculate the state information loss between the first state information and the second state information according to a second formula; Wherein, the second formula is: ; Indicates that the state information is lost, represents the second state information, represents the first state information, Represents the matrix norm.

8. The method according to claim 1, characterized in that Calculate the prediction result loss of the fault prediction result according to the third formula; Wherein, the third formula is: ; represents the prediction result loss, Indicates the fault prediction result With fault label The supervision loss between is the number of root cause instances output by the model, is the number of fault-free instances output by the model, which is optimized by the constraint .

9. The method according to claim 1, characterized in that: The comprehensive loss is calculated by using the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result, including: The fusion information loss, the state information loss and the prediction result loss are weightedly fused to obtain the comprehensive loss.

10. The method according to claim 1, characterized in that The comprehensive loss is calculated by using the fusion information loss of the fusion information corresponding to each microservice, the state information loss between the first state information and the second state information, and the prediction result loss of the fault prediction result, including: The average value of the fusion information loss, the state information loss and the prediction result loss is calculated, and the average value is used as the comprehensive loss.

11. The method according to any one of claims 1 to 10, characterized in that: Also includes: The fault prediction model is deployed in each microservice.

12. A fault prediction method, characterized in that: Applied to each microservice included in the microservice system, including: Use the fault prediction model deployed in itself to predict faults and issue alarms in real time; Wherein, the fault prediction model is obtained based on the model training method described in any one of claims 1 to 11.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 12.

14. A non-volatile storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Internet of Things anomaly detection method and device based on graph structure learning

    CN117688504A

  • Internal threat anomaly detection method based on heterogeneous knowledge distillation

    CN118070282A