Micro-service anomaly detection method based on dynamic calling
By building a microservice system state graph sequence model and using a spatiotemporal feature extraction model based on attention mechanism and a deep support vector data description model, the problem of dynamically changing call relationships and method-level anomaly detection between microservice instances is solved, more accurate anomaly detection is achieved, and the usability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510051330.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to handle the dynamically changing call relationship between microservice instances, and it is impossible to detect exceptions from the method level, resulting in low accuracy of exception detection.
The microservice exception detection method based on dynamic calls is adopted, and the call dependency between microservice instance features and services is extracted through the indicator data and call chain data in the microservice system, and the microservice system state chart sequence model is constructed, and the spatial-temporal feature extraction model and deep support vector data description model based on attention mechanism are used for abnormal detection.
A more accurate detection of abnormal microservice instances is achieved, improving the usability and reliability of the system, and reducing the negative impact on the user experience.
Smart Images

Figure CN119988074A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data center microservice anomaly detection, and in particular relates to a microservice anomaly detection method in a data center cluster. Background Art
[0002] With the booming development of cloud computing, the software system architecture of data centers is gradually evolving from a monolithic model to a service-oriented model. In a monolithic architecture, the complete business logic is implemented as an executable binary file. However, with the continuous enrichment of software system functions and the continuous increase in the demand for agile development, the monolithic architecture faces huge challenges and is difficult to flexibly adapt to the fluctuation of business capacity and the elastic expansion and contraction requirements. Service-oriented architecture has become a popular trend today. Microservice architecture is a typical representative of this trend. By decomposing complex applications into fine-grained, lightweight, elastic and independently maintainable microservices, each microservice is responsible for a single simple function, thus achieving the advantage of easy horizontal expansion. The adoption of this architecture has significantly improved system performance, scalability and flexibility, and has been widely used in fields such as the Internet of Things (IoT) and cloud native. Anomaly detection in microservice systems refers to timely identifying faults or abnormal conditions that may affect their normal functions by monitoring the operating status of microservices. The operating status of microservices is mainly reflected by the following three data sources: key performance indicators (KPIs), logs, and call chain data. In the complex ecology of microservice architecture, anomaly detection plays a vital role. It is the cornerstone for ensuring stable operation of the system and timely troubleshooting of hidden dangers. Anomaly detection can keenly capture abnormal signals in services and prevent potential failures by deeply analyzing key performance indicators (KPIs) and complex call chain data. This process not only relies on efficient data processing capabilities, but also requires the use of intelligent algorithms to accurately analyze massive amounts of data, so as to achieve rapid response to problems and minimize the negative impact on user experience. An effective anomaly detection mechanism is an indispensable part of the microservice architecture. It ensures the reliability and availability of the system and provides solid support for business continuity.
[0003] The existing data center microservice anomaly detection work has the following defects: 1) It is unable to handle the dynamically changing call relationships between microservice instances. In a microservice system, the call relationship between microservice instances is not static, but dynamically adjusted as requests change. This feature brings significant challenges to the identification of abnormal microservices. Specifically, since the system processes multiple requests at the same time and each request may flow through different microservice instances, the request types and their combinations on the microservice instances will change over time. 2) It is mostly targeted at the microservice level and cannot detect anomalies at the method level. The response time feature is not only significantly different in its normal and abnormal range due to different call types, but also affected by the request intensity and the type of calling method. Judging the abnormal state of the microservice system only from the response time interval in a single request scenario leads to a high risk of misjudgment. 3) Microservice indicators are mostly detected from a single request perspective, but microservices are also affected by other requests in the microservice system, which affects the accuracy of anomalies. Summary of the invention
[0004] In order to solve the problems of the above-mentioned prior art, the present invention proposes a microservice anomaly detection method based on dynamic calling. Aiming at the multi-request concurrent scenario, the present invention utilizes the indicator data and call chain data in the microservice system, extracts the microservice instance features and the call dependency between services, and constructs a microservice system state diagram sequence model, uses the spatiotemporal feature extraction model based on the attention mechanism to learn the instance embedding vector, and then detects abnormal microservice instances through the deep support vector data description model. The model constructed by this method can more accurately detect abnormal microservice instances, thereby improving the system availability and reliability.
[0005] The proposed method for data center microservice anomaly detection consists of four steps: initialization, microservice system state graph construction, spatiotemporal feature extraction model construction based on attention mechanism, and anomaly detection based on deep support vector data description model. In this method, the following important parameters are used: number of iterations β, batch size b, spatiotemporal feature extraction neural network learning rate step, sliding serial port size wd, node feature dimension nd, graph attention head number muti, instance feature vector dimension lf. β is 128-512, b is 64, step0.001, wd is 8, nd is 32-64, muti is 8, and lf is 32.
[0006] Before executing this method, the indicator data and call chain data need to be converted into a processable form. The indicator data can be read with the help of the data processing library, and after data cleaning (processing missing values, converting data types, etc.), it is saved as a pandas DataFrame object. The call chain data is read using the JSON parsing library, and key information such as timestamp, source service, target service, and operation name are extracted and converted into the data format and stored as a dictionary.
[0007] (1) Initialization
[0008] The present invention uses indicator data and call chain data when performing data initialization processing. First, define the entire attribute set of indicator data A = {a1, a1...a F}, divided into CPU, memory, IO and network categories. Among them, the CPU category indicators include: the total number of time cycles of the container under the completely fair scheduler; the number of time cycles that the container is restricted under the completely fair scheduler; the total time that the container is restricted; the average CPU load of the container in the past 10 seconds; the CPU time used by the container in system mode; the total CPU usage time of the container; the CPU time used by the container in user mode, etc. The memory category indicators include: the memory cache size of the container; the number of memory allocation failures of the container; related to memory allocation failures; the memory size occupied by the container mapping file; the maximum memory usage of the container; the resident set size of the container; the swap space size used by the container; the current memory usage of the container; the working set size of the container, etc. File system I / O category indicators: the current number or rate of file system I / O operations of the container; the total time of the file system I / O operation of the container; the total time after weighting the file system I / O time; the file system capacity limit of the container; the time spent by the container to read the file; the number of file read operations of the container; the total size of the files read by the container; the number of file read operations merged by the container, etc. Network indicators: the number of network packets received by the container; the number of errors when the container sends network packets; the number of network packets received by the container that are discarded; the total amount of network data sent by the container; the number of network packets sent by the container; the number of network packets sent by the container that are discarded, etc. All indicator data sequences are subjected to maximum and minimum standardization operations to eliminate dimensional differences, and then the constant attributes and attributes not directly related to performance analysis are screened out to obtain a refined attribute subset K = {k1, k2…k S}, This subset is a subset of the entire attribute set. The recording time span of the indicator data starts from the earliest time point T min To the latest time point T max , using the sliding serial port size wd, the step size is 1, and the sliding window technology is used to cut all indicator data to form a series of time series, where the time series can be expressed as mt = (mt1,...,mt i ,...,mt wd ). i It contains the performance indicators of all microservice instances. The time series set corresponding to any microservice instance m is recorded as Contains all performance indicators, among which the time series corresponding to the indicator kr can be expressed as For the call chain data, the present invention adopts the same cutting strategy as the indicator data, dividing it into time series with the same time span of wd minutes. The call chain consists of call events, including information such as the caller, the callee, the response time, the call method, and the call type. In particular, the response time is an important feature reflecting the internal performance of the microservice, and its granularity is as fine as 1 millisecond. In order to smooth the response time data and reduce noise, the present invention calculates the mean of all response times of the same microservice instance using the same call method within each minute to obtain the relative response time. Based on this, the relative response time series of the microservice instances in the window can be constructed, and the specific form of the series can be expressed as The combined indicator attributes and relative response time form an attribute subset.
[0009] (2) Construction of microservice system state diagram
[0010] 2.1) For the microservice system w in any time period i ,1≤i≤TD, construct the system state diagram of the microservice system based on the call chain data and indicator data in the time period. The system state diagram of the microservice system consists of the indicator characteristics in the indicator data and the response time and dependency relationship in the call chain. First, define its dependency relationship through the adjacency matrix. The adjacency matrix is as follows:
[0011]
[0012] The number of rows and columns of the matrix is H. The number of rows and columns of the matrix is equal to the number of microservice instances in the system. Each row represents a microservice instance, and each column represents a microservice instance. r,p ,1≤r≤H,1≤p≤H,When microservice instance p depends on microservice instance r, a r,p =1; when microservice instance p does not depend on microservice instance r, a r,p =0.
[0013] 2.2) For a microservice system within a period of time, the fixed time interval is 1 minute, and the entire time period contains TD time intervals. Each time interval corresponds to the state w of a microservice system i , a matrix is used to define the characteristics of the microservice instances in the microservice system within each interval, which is in the following form:
[0014]
[0015] The number of rows of the matrix is H, the number of columns is S, and any element f of the matrix gq ,1≤g≤H,1≤q≤S, represents the microservice system w i In the example wt ig In the indicator kr qThe number of rows of the matrix is the number of microservice instances H in the microservice system, each row corresponds to a microservice instance, the number of columns of the matrix is the number of all attributes of the initialized attribute subset S, and each column corresponds to an attribute in the attribute subset TR
[0016] 2.3) For all microservice systems w in each time period i , according to w i Initialize the adjacency matrix A with the number of microservice instances H in w i The microservice instances in the traversal are constructed, and the adjacency matrix A is constructed according to the dependencies between the instances. For the microservice system w in each time period i , initialize the feature matrix X and fill it with data according to the number of microservice instances H it contains and the number of elements S in the attribute subset TR.
[0017] (3) Construction of spatiotemporal feature extraction model based on attention mechanism
[0018] 3.1) Using the spatiotemporal feature extraction model as the design structure of the instance feature extraction model, the construction of the spatiotemporal feature extraction model based on the attention mechanism includes four parts: graph attention layer, sorting layer, temporal attention layer, and output layer. The graph attention layer consists of three graph attention layers. The sorting layer mainly rearranges and combines the features obtained through the graph attention layer to form a sequence of microservice instances in the time dimension, which is convenient for the subsequent time encoder to extract time features. The temporal attention layer consists of three self-attention layers, and the output layer consists of two fully connected layers. The input of the model is wd H×H adjacency matrices A and wd H×S feature matrices X, where H is the number of microservice instances in the microservice system, S is the feature dimension of the microservice instance, and the output is the graph attention layer output feature vector of all microservice instances in the microservice system. The instance feature aggregates the time features of the microservice instance in the topological dependency relationship of the dependency structure and the performance index. The learning rate of the graph neural network is set to step, and the batch size of the training is b.
[0019] The graph attention layer uses a spatial domain-based graph attention calculation method, as shown in formula (1), x u and x v is the node feature vector of instance u and instance v obtained by 2.2), A uv is the element in the adjacency matrix obtained from 2.1). ss ∈R F*F′ The graph attention weight matrix is used to calculate the weight of the microservice instance to each neighbor instance. This weight matrix is the relationship between the input F features and the output F′ features, which plays a mapping role. uvrepresents the attention coefficient of instance u to instance v, σ(·) is the LeakyRELU nonlinear activation function, and a T is a learnable weight vector, and T is the transpose of the vector. Its function is to multiply the intermediate feature vector after transformation and concatenation to obtain the intermediate calculation result of the attention coefficient. When calculating the attention value, x u and x v , using the attention weight matrix W ss Do the mapping and concatenate the resulting vectors. Then, use the feedforward neural network a T Map the concatenated vector to real numbers.
[0020] e uv =σ(A uv ·a T [W ss x u ||W s x v ])#(1)
[0021] Next, the value is normalized and the calculation formula is shown in formula (2). uυ is the normalized attention coefficient of instance u to instance v. exp(·) is an exponential function, exp(e uυ ) represents the original attention coefficient e uυ Perform exponential operation, ∑ represents the sum operation, N v represents the set of neighbor nodes of instance v. Formula (2) represents the sum of the attention coefficients of all neighbor instances of instance v after exponential operation, and finally obtains the attention coefficient α uυ .
[0022]
[0023] Get the attention coefficient α uυ After that, the weighted sum of neighbors can be performed to obtain the graph output feature vector of microservice instance v as shown in formula (3). υ is the output feature vector of the graph attention layer of instance v, σ(·) is the LeakyRELU nonlinear activation function, ∑ represents the summation operation, N v Represents the set of neighbor nodes of instance v. u is the feature vector of instance u, W ss is the attention weight matrix, α uυ It is the coefficient obtained by applying the LeakyRELU function to the neighbors of each node, which indicates the contribution of microservice instance u to microservice instance v in the current microservice system state diagram.
[0024]
[0025] The output of the graph attention layer is in units of multiple microservice systems. Each microservice system corresponds to an instance representation matrix at different times, forming a matrix sequence. In order to facilitate the processing of the subsequent sequential attention layer, the output of the graph attention layer needs to be reordered. The instance representation matrix output by the graph attention layer is Z wd,xm,xd , where wd represents the time index (wd=1,2,...,WD), xm represents the microservice instance index (xm=1,2,...,XM), and xd represents the feature dimension index (xd=1,2,...,XD). After reordering, the graph feature extraction representation matrix Z of different instances is formed I ∈R WD×XD .
[0026] The representation matrices of these different instances will enter the temporal attention layer. In the temporal attention layer, the weight matrix W is first used q ,W k ,W v Perform a linear transformation on the representation vector to obtain the query vector Q, key vector K, and value vector V, which can be specifically expressed as formula (4)(5)(6), where W q Query weight matrix, W k is the key weight matrix, W v is the value weight matrix.
[0027] Q=Z i W q #(4)
[0028] K=Z i W k #(5)
[0029] V=Z i W v #(6)
[0030] Then, the query vector Q and the key vector K are used to calculate the representation of different nodes in the time series according to formula (7), where Z i is the graph feature extraction representation vector of instance i at time point, is the attention value of time j to time i. ((Z i W q )(Z i W q ) T ) ij It is a linear transformation of the query weight matrix and the representation vector. T is the transposition operation, and then the inner product of the two transformed representation matrices is calculated to measure the correlation or similarity between nodes. F′ is the normalization factor, M ijis the position encoding matrix, which is used to solve the problem of temporal feature disappearance caused by parallel input. It is defined as formula (8). Then all attention values are softmax normalized. The specific expression is shown in formula (9). exp(·) is an exponential function. Represents the original attention coefficient Perform exponential operations, It is the sum of the exponential transformation values of the attention weight of the node at all times. The normalized attention β is obtained ij According to the normalized attention value β ij The sum value vector V gives the node representation Z i ′, as shown in formula (10).
[0031]
[0032] Z′ i =β ij (Z i W v )#(10)
[0033] In Z′ i After inputting into the multi-layer perceptron (MLP) layer for feature integration learning, the output representation of each microservice instance is a vector h. First, it is a matrix processed by the temporal attention layer, which contains the feature representation of the microservice instance at different times. When it is input into the MLP layer, the MLP can further integrate and learn these features through a series of linear transformations and nonlinear activation functions. The MLP consists of two linear layers and a ReLU activation function. For the input matrix Z′ i , first pass through the first linear layer, as shown in formula (11). Where W1 is the weight matrix from the input layer to the hidden layer, b1 is the bias vector from the input layer to the hidden layer, Y1 is the intermediate variable of the hidden layer, and Y′1 is the intermediate variable of the activated hidden layer. Then it is transformed nonlinearly through the ReLU activation function. Then pass through the second linear layer, the linear transformation formula is (13), where W2 is the weight matrix from the hidden layer to the output layer, b2 is the bias vector from the hidden layer to the output layer, and the final output vector h of the microservice instance v is obtained.
[0034] Y1=Z′ i W1+b1#(11)
[0035] Y′1=ReLU(Y1)#(12)
[0036] h=Y′1W2+b2#(13)
[0037] 3.2) Use the microservice system state set W to train the spatiotemporal feature extraction model AF.
[0038] 3.2.1) Train the constructed spatiotemporal feature extraction network AF and use all microservice system states in the microservice system state set W as sample data. These sample data are used to generate the adjacency matrix and feature matrix in the system state set. Specifically, the adjacency matrix A is constructed based on the dependency relationship between the microservice instances in the sample data. i , construct the feature matrix X based on the performance indicators of the microservice instances in the sample data i Each training input is 64 samples, and the adjacency matrix and feature matrix are used as the input values of the model. Each sample contains wd adjacency matrices and feature matrices. The adjacency matrix and feature matrix are combined to represent the state of the microservice system within 1 minute. i is the adjacency matrix of the ith minute, X i is the feature matrix of the ith minute. The model parameters are updated through the forward propagation algorithm and the Adam optimizer for training, and the input is repeated until all microservice systems are trained.
[0039] 3.2.2) Repeat the process of 3.2.1) β times. The hyperparameter β represents the number of rounds of parameter updates during model training, and its value range is 128-512. In practical applications, different β values are tried through multiple experiments. According to the performance of the model on the validation set (such as accuracy, recall rate, F1 value and other evaluation indicators), the β value that makes the model performance optimal is selected as the final parameter setting. The model is updated with multiple rounds of parameters. The parameters here include the graph attention weight matrix, query weight matrix, key weight matrix, value weight matrix, position encoding matrix, input layer to hidden layer weight matrix, and hidden layer to output layer weight matrix. After the parameter update is completed, the microservice system training is completed, and the corresponding spatiotemporal feature extraction network is constructed.
[0040] (4) Anomaly detection based on deep support vector data description model
[0041] 4.1) Take the microservice system instance output vector h as the attribute, use the deep support vector data description model to perform anomaly classification, project the microservice instance feature vector h into the hypersphere, calculate the distance between the feature vector and the center c, and initialize the randomly selected hypersphere with the center point c and radius R.
[0042] 4.2) Use the microservice system dataset to collaboratively train the spatiotemporal feature extraction model and anomaly detection model based on the attention mechanism. The specific process is as follows: First, the sample data in the microservice system dataset W is input into the spatiotemporal feature extraction model, and the feature vector representation of the microservice instance is obtained through the forward propagation algorithm. At the same time, these feature vectors are input into the anomaly detection model as attributes. During the training process, the reconstruction loss and clustering loss are calculated jointly, as shown in formula (14). Among them, R is the radius of the hypersphere, and the loss function attempts to minimize this value so that the hypersphere can cover as many data points as possible. The hyperparameter μ is used to balance the volume and boundary violations of the hypersphere, allowing some data to be outside the hypersphere. max{0,∥hc∥ 2 -R 2} is a penalty term used to penalize data points outside the hypersphere. c represents the center of the hypersphere. ∥hc∥ 2 Calculate data point h o The square of the distance to the center c of the hypersphere. represents a regularization term, and λ is a regularization parameter used to control the strength of regularization. It represents the square of the Frobenius norm of the weight matrix W of the lth layer, which is used to measure the size of the weight matrix of this layer, that is, the complexity of the model. Indicates the sum of the weight matrices of all layers to ensure that the regularization term can take into account the complexity of each layer in the model. For the values of the hyperparameters μ and λ, users can choose different μ and λ in the range of (0-1) to experiment and obtain a set of μ and λ with the fastest decrease in the loss function of the user's current model. The data set is divided into a training set and a validation set, and each value is co-trained. The accuracy, recall rate, F1-score and other indicators are used to evaluate the validation set, and the value that optimizes the overall performance of the model is selected. It may be necessary to adjust the value based on multiple experiments according to the data distribution and model characteristics. The loss function is minimized by adjusting the parameters of the spatiotemporal feature extraction model and the anomaly detection model.
[0043]
[0044] For the spatiotemporal feature extraction model, first initialize its parameters and the initial radius of the hypersphere, and determine the center of the hypersphere through forward propagation. Then, fix the radius and the center of the sphere and focus on training the spatiotemporal feature extraction network. Use the Adam optimizer to minimize the difference between the data points mapped to the surface and the center of the hypersphere. Based on the parameters of the current spatiotemporal feature extraction network, recalculate the representation of the data points on the hypersphere. Users can select different percentile values in the range of [50%, 95%] to update the radius according to the effect and needs of the model. After every 64 samples of feature extraction, the distance Distance (h, c) from all instance output vectors in the current sample to the center of the hypersphere is processed. The formula is shown in (15), where h is the instance output vector and c is the center of the sphere. First, convert these distances into tensors and flatten them, then sort these distances from small to large. After finding the distances sorted from small to large, select the distance corresponding to the percentile such as 90% as the radius, which can ensure that 90% of the data points are within the hypersphere.
[0045] For the anomaly detection model, the feature vector output by the spatiotemporal feature extraction model is used to determine whether a data point is abnormal based on its distance from the center of the hypersphere. For the output vector of an instance, the distance Distance(h,c) from the center of the hypersphere is calculated. If this distance is greater than the radius of the hypersphere, the data point is considered an abnormal point, and points outside the hypersphere are considered abnormal. By continuously adjusting the parameters of the anomaly detection model, mainly the radius of the hypersphere, the radius of the hypersphere is adjusted by minimizing the loss function. The loss function is shown in formula (14). According to the distance from the data point to the center of the hypersphere and the current radius, the radius and all weight matrices in the spatiotemporal feature extraction model are adjusted through the Adam optimization algorithm to minimize the loss function and improve the accuracy of anomaly detection.
[0046] During the entire collaborative training process, the two models influence and optimize each other, jointly improving the feature extraction and anomaly detection performance of the microservice system.
[0047] Distance(h,c)=||hc|| 2 #(15)
[0048] 4.3) Repeat step 4.2) until the radius of the hypersphere no longer changes, thus completing the spatiotemporal feature extraction model and anomaly detection model based on the attention mechanism.
[0049] (5) Microservice instance feature extraction
[0050] 5.1) For any microservice system w i , input its adjacency matrix A i With the feature matrix X i, use the spatiotemporal feature extraction model based on the attention mechanism for feature extraction, and output the output vector h of all microservice instances in the microservice system.
[0051] 5.2) Repeat step 5.1) until feature extraction of all microservice systems is completed.
[0052] (6) Microservice instance exception classification
[0053] 6.1) Traverse the instance feature set and calculate the distance between any microservice output vector h and the center point according to formula (15) to determine whether the data point is within the hypersphere. If it is within the hypersphere, it is normal, otherwise it is abnormal. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 A cluster platform for anomaly detection methods in microservice systems.
[0055] Figure 2 This is a schematic diagram of the present invention.
[0056] Figure 3 It is a flow chart of the present invention.
[0057] Figure 4 Flowchart built for microservice system statechart.
[0058] Figure 5 Flowchart built for the spatiotemporal feature extraction model.
[0059] Figure 6 A flowchart built for a microservice instance anomaly detector. DETAILED DESCRIPTION
[0060] The present invention is described below in conjunction with the accompanying drawings and specific embodiments.
[0061] The data center microservice system anomaly detection method proposed in the present invention is built on multiple connected servers and is implemented by writing corresponding functions. Figure 1It is a deployment diagram of the platform built by this method. The platform is composed of multiple computer servers (platform nodes), which are connected through a network to store data and execute tasks in a distributed manner. Platform nodes are divided into two categories: one management node and multiple computing nodes. The platform built by the method of the present invention includes three types of core software modules: a resource management module, a data receiving module, and a data processing module. Among them, the resource management module is responsible for allocating the required indicators and call chain data to the data receiving module, and collecting and managing the data results, and is only deployed on the management node; the data receiving module is responsible for pulling the required indicators and call chain data, and needs to be deployed on each computing node; the data processing module is responsible for running the corresponding algorithm and returning the results to the resource management module, which is deployed on the computing node. The above three types of software modules are all deployed and run when the platform starts.
[0062] Figure 2 This is the architecture diagram of the method of the present invention. The present invention uses data center indicators and call chain data as input, and first generates a corresponding adjacency matrix and feature matrix for each microservice system. Based on the generated adjacency matrix and feature matrix set, a spatiotemporal feature extraction model is constructed. By training the model, a feature extraction model of the microservice instance is obtained. The adjacency matrix and feature matrix of the samples in any microservice system are input into the spatiotemporal feature extraction model to obtain the feature vectors of all instances. An anomaly detection model is constructed for the instance vector set, and the center point and radius of the hypersphere are obtained after training. The feature vector of any microservice instance is input into the anomaly detection model to obtain the anomaly detection result of the instance.
[0063] Combine the following Figure 3 Invention content The overall process describes the specific implementation method of this method. In this implementation method, the basic parameters are set as follows: number of iterations β = 256, batch size b = 64, spatiotemporal feature extraction neural network learning rate step = 0.001, sliding serial port size wd = 8, node feature dimension nd = 32, graph attention head number multi = 8, instance feature vector dimension lf = 32.
[0064] The specific implementation method can be divided into the following steps:
[0065] 1. Initialization
[0066] The indicators used in this invention have a total of 65 attributes, including the full set of attributes A = {a1, a1…a 65}, perform the maximum and minimum standardization operations on all indicator data sequences to eliminate dimensional differences. Then, by screening out constant attributes and attributes that are not directly related to performance analysis, 31 refined attributes are obtained, including 10 CPU indicators, 11 memory indicators, 5 IO indicators, and 5 network indicators. The indicator data contains sample data from 86,340 microservice systems. Take the microservice system w1 corresponding to one of the sample data as an example. The system contains 32 attributes, 31 of which are extracted from indicators and 1 from the call chain. The start time of each microservice system is time o , for example, the start time of the metric and response time attributes of w1 is 1647705600.
[0067] 2. Construction of microservice system state diagram
[0068] 2.1) For the microservice system w in any time period i , complete the construction of adjacency matrix and feature matrix. Taking microservice system w5 as an example, there are 40 microservice instances belonging to w5, w5 = {wt 1,1 ,wt 1,2 ,…,wt 1,40}, attribute subset tr={tr1,tr2...,tr 32}, so according to the method of steps 2.1) and 2.2) in the invention content, its adjacency matrix A5 and feature matrix X5 are defined as follows:
[0069]
[0070] 2.2) For each microservice system sample w i , based on the model constructed in 2.1) i and X i Initialization and data filling.
[0071] 2.2.1) Select a microservice system w that does not generate an adjacency matrix and a feature matrix i , taking microservice system w5 as an example, the adjacency matrix and feature matrix of microservice system w5 are established by the method of step 2.3) in the invention content. Instance 40 in microservice system w5 depends on 1, so the corresponding element value is 1, and the adjacency matrix A5 is constructed as follows:
[0072]
[0073] Continue to build the feature matrix and fill the matrix according to the feature values of the microservice instances in each microservice system. The feature matrix is as follows:
[0074]
[0075] 2.2.2) Repeat step 2.2.1) until the adjacency matrix and feature matrix are established for all microservice systems.
[0076] 3. Construction of spatiotemporal feature extraction model based on attention mechanism
[0077] 3.1) Construct the basic structure of the spatiotemporal feature extraction network, set the neural network learning rate to 0.001, the number of iterations to 256, and the training batch size to 64. The first two layers of the graph attention layer are multi-head attention layers, each layer consists of 8 graph attention heads, each head has an input feature dimension of 32 and an output feature dimension of 32. The third layer is a graph attention layer that integrates the output of multi-head attention, with an input feature dimension of 256 and an output feature dimension of 32.
[0078] 3.2) Train the spatiotemporal feature extraction network.
[0079] 3.2.1) Taking the microservice system sample set W as an example, 64 microservice system samples {w5,w 12 ,…,w 131}As a sample for training once, the input of the sample is 64×8 adjacency matrices and feature matrices, and the size of the adjacency matrix of different samples is different. Training is performed by the method of step 3.2) in the content of the invention. First, all the states of the microservice systems in the microservice system sample set W are used as sample data, and b samples are input for each training, where b is the training batch size, which is set to 64 here. The input of the sample is the adjacency matrix and the feature matrix, which represent the dependencies and features between the microservice instances in the microservice system. Then, the model parameters are updated by the forward propagation algorithm and the Adam optimizer for training. During the training process, the sample is input repeatedly until all microservice system snapshots are trained.
[0080] 3.3) Repeat the process of 3.2.1) 256 times, perform multiple rounds of parameter updates on the model, and complete the training after the parameter updates, completing the construction of the spatiotemporal feature extraction model AF.
[0081] 4. Construction of anomaly detector based on deep support vector data description model
[0082] 4.1) Construct a deep support vector data description model and initialize the model. Randomly select a sample in the instance of the microservice system as the center point of the hypersphere and randomly set the radius.
[0083] 4.2) Train the deep support vector data description model.
[0084] 4.2.1) Update the radius of the hypersphere in the model according to the method in step 4.2) in the invention content. Use the microservice system dataset W to collaboratively train the spatiotemporal feature extraction model and the anomaly detection model based on the attention mechanism. Formula (15) is used to jointly calculate the reconstruction loss and the clustering loss. Where R is the radius of the hypersphere, and the loss function attempts to minimize this value so that the hypersphere can cover as many data points as possible. The hyperparameter μ is used to balance the volume and boundary violations of the hypersphere, allowing some data to be outside the hypersphere. max{0,∥hc∥ 2 -R 2} is a penalty term used to penalize data points outside the hypersphere. c represents the center of the hypersphere. ∥hc∥ 2 Calculate data point h o The square of the distance to the center c of the hypersphere. represents a regularization term, and λ is a regularization parameter used to control the strength of regularization. It represents the square of the Frobenius norm of the weight matrix W of the lth layer, which is used to measure the size of the weight matrix of this layer, that is, the complexity of the model. represents the sum of the weight matrices of all layers to ensure that the regularization term takes into account the complexity of each layer in the model.
[0085] The optimization objective is to make the network learn the parameters W, R, c to ensure that the data points are tightly mapped to the center of the hypersphere. During the training process, the parameters of the spatiotemporal feature extraction network and the initial radius of the hypersphere are first initialized, and the center of the hypersphere is determined by forward propagation. Next, the radius is fixed and the focus is on training the spatiotemporal feature extraction network, using the Adam optimizer to minimize the difference between the mapping of data points to the surface and the center of the hypersphere. Based on the parameters of the current spatiotemporal feature extraction network, the representation of the data points on the hypersphere is recalculated, and the radius is dynamically adjusted according to the set percentile parameter to ensure that it can tightly surround most of the data points.
[0086] 4.2.2) Repeat step 4.2.1) until the radius in the model no longer changes, completing the training of the anomaly detection model.
[0087] 5. Microservice instance feature extraction
[0088] 5.1) Taking the microservice system w5 in the microservice system W as an example, its adjacency matrix A5 and feature matrix X5 are input into the spatiotemporal feature extraction model. The spatiotemporal feature extraction network outputs the feature vector V5 of all microservice instances of the microservice system w5, which is in the following form:
[0089] V5={0.1,1.2,-3.1,…,2.1}
[0090] 5.2) Repeat step 5.1) until the feature extraction of microservice instances in all microservice systems is completed.
[0091] 6. Microservice instance exception classification
[0092] 6.1) For any microservice system w i The microservice instances in the invention are classified according to the method of step 6.1) in the invention content, and the instance output feature vector set is traversed. The distance between any microservice output feature vector and the center point is calculated according to formula (15) to determine whether the data point is located in the hypersphere. If it is in the hypersphere, it is judged as normal, otherwise it is judged as abnormal. The type of the instance is obtained, and the microservice instance classification is completed.
[0093] According to the anomaly detection method for data center microservices proposed by the present invention, the inventors conducted relevant performance tests. The test results show that the method of the present invention is applicable to AIOps2022 with a huge data volume. The method can accurately classify anomalies of data center microservices.
[0094] The performance test compares this method with the existing anomaly detection method to show the advantages of the proposed method in terms of interference degree prediction accuracy. The comparison method is as follows:
[0095] (1) Anomaly detection method for single request environment
[0096] This method only considers the key performance indicators and call dependencies of microservices from the perspective of a single request, and does not consider that the microservice is affected by other requests in the microservice system in addition to the current request.
[0097] (2) Anomaly detection methods that do not consider the characteristics of the calling method in the call chain
[0098] This method does not consider the response time in the call chain as a clue for anomaly detection, but only detects abnormal microservices through indicator data or simply considers the call relationship.
[0099] The performance tests were run on a computer with the following configurations: processor: Intel Core i7-9700; memory capacity: 32GB; hard disk storage capacity: 1TB; operating system: Windows.
[0100] Precision, recall, and F1-score are often used to measure the effectiveness of anomaly detection.
[0101] The performance test selects batch processing jobs in Alibaba logs. The evaluation results of the dependent graph structure features (including the number of nodes, the longest path, the maximum number of parallel operations, and the edge density) are shown in Table 1.
[0102] Table 1. Comparison of anomaly detection results
[0103]
[0104] It can be concluded from the data in Table 1 that, compared with the comparison method, the method of the present invention has a higher improvement in anomaly detection accuracy, with an improvement of 17.1% and 8.0% in precision, 10.0% and 12.6% in recall, and 13.7% and 10.4% in F1_score respectively; in summary, the method of the present invention has a significant improvement in the accuracy of microservice anomaly detection.
[0105] Finally, it should be noted that the above examples are only used to illustrate the present invention and are not intended to limit the technology described in the present invention. All technical solutions and improvements that do not deviate from the spirit and scope of the invention should be included in the scope of the claims of the present invention.
Claims
1. A microservice anomaly detection method based on dynamic invocation, characterized in that: The parameters are as follows: number of iterations β, batch size b, learning rate of the spatiotemporal feature extraction neural network step, sliding serial port size wd, node feature dimension nd, number of graph attention heads muti, instance feature vector dimension lf; β is 128-512, b is 64, step 0.001, wd is 8, nd is 32-64, muti is 8, and lf is 32; Before execution, the indicator data and call chain data need to be converted into a processable form; the indicator data is read with the help of the data processing library, and after data cleaning, it is stored as a pandas DataFrame object; the call chain data is read using the JSON parsing library, and the timestamp, source service, target service, and operation name are extracted and converted into the data format and stored as a dictionary; (1) Initialization When performing data initialization processing, the indicator data and call chain data are used; first, the entire attribute set of the indicator data is defined as A = {a1, a1…a F }, divided into CPU, memory, IO and network categories; among them, the CPU category indicators include: the total number of time periods of the container under the completely fair scheduler; the number of time periods that the container is restricted under the completely fair scheduler; the total length of time the container is restricted; the average CPU load of the container in the past 10 seconds; the CPU time used by the container in system mode; the total CPU usage time of the container; The CPU time used by the container in user mode; memory indicators include: the size of the container's memory cache; the number of times the container's memory allocation failed; related to memory allocation failure; the memory size occupied by the container's mapped files; the maximum memory usage of the container; the size of the container's resident set; the size of the swap space used by the container; the current memory usage of the container; the working set size of the container; file system I / O indicators: the number or rate of the container's current file system I / O operations; the total time of the container's file system I / O operations; the total time after weighting the file system I / O time; the file system capacity limit of the container; the time it takes the container to read files; the number of file read operations of the container; the total size of files read by the container; the number of file read operations merged by the container; network indicators: the number of network packets received by the container; the number of errors when the container sends network packets; the number of network packets received by the container that are discarded; the total amount of network data sent by the container; the number of network packets sent by the container; the number of network packets sent by the container that are discarded; perform maximum and minimum normalization operations on all indicator data sequences to eliminate dimensional differences, and then filter out constant attributes and attributes that are not directly related to performance analysis to obtain a refined attribute subset K = {k1, k2…k S }, This subset is a subset of the entire attribute set; the recording time span of the indicator data starts from the earliest time point T min To the latest time point T max , using the sliding serial port size wd, the step size is 1, and the sliding window technology is used to cut all indicator data to form a series of time series, where the time series is represented by mt = (mt1,...,mt i ,...,mt wd );mt i It contains the performance indicators of all microservice instances. The time series set corresponding to any microservice instance m is recorded as Contains all performance indicators, where the time series corresponding to the indicator kr is expressed as For the call chain data, the same cutting strategy as the indicator data is adopted to split it into time series with the same time span of wd minutes; the call chain consists of call events, including the caller, the callee, the response time, the call method and the call type information; in particular, the response time is an important feature reflecting the internal performance of the microservice, and its granularity is as fine as 1 millisecond; in order to smooth the response time data and reduce noise, the average of all response times of the same microservice instance using the same call method within each minute is calculated to obtain the relative response time; based on this, the relative response time series of the microservice instances in the window can be constructed, and the specific form of the series is expressed as The combined indicator attributes and relative response time form an attribute subset; (2) Construction of microservice system state diagram 2.1) For the microservice system w in any time period i ,1≤i≤TD, construct the system state diagram of the microservice system according to the call chain data and indicator data in the time period. The system state diagram of the microservice system is composed of the indicator characteristics in the indicator data and the response time and dependency relationship in the call chain; first, define its dependency relationship through the adjacency matrix, and the adjacency matrix form is as follows: The number of rows and columns of the matrix is H. The number of rows and columns of the matrix is equal to the number of microservice instances in the system. Each row represents a microservice instance, and each column represents a microservice instance. Any element a in the matrix r,p ,1≤r≤H,1≤p≤H,When microservice instance p depends on microservice instance r, a r,p =1; when microservice instance p does not depend on microservice instance r, a r,p =0; 2.2) For a microservice system within a period of time, the fixed time interval is 1 minute, and the entire time period contains TD time intervals; each time interval corresponds to the state w of a microservice system i , a matrix is used to define the characteristics of the microservice instances in the microservice system within each interval, which is in the following form: The number of rows of the matrix is H, the number of columns is S, and any element f of the matrix gq ,1≤g≤H,1≤q≤S, represents the microservice system w i In the example wt ig In the indicator kr q The number of rows of the matrix is the number of microservice instances H in the microservice system, each row corresponds to a microservice instance, the number of columns of the matrix is the number of all attributes of the initialized attribute subset S, and each column corresponds to an attribute in the attribute subset TR 2.3) For all microservice systems w in each time period i , according to w i Initialize the adjacency matrix A with the number of microservice instances H in w i The microservice instances in the traversal are constructed, and the adjacency matrix A is constructed according to the dependencies between the instances. For the microservice system w in each time period i , initialize the feature matrix X and fill it with data according to the number of microservice instances H and the number of elements S in the attribute subset TR; (3) Construction of spatiotemporal feature extraction model based on attention mechanism 3.1) Using the spatiotemporal feature extraction model as the design structure of the instance feature extraction model, the construction of the spatiotemporal feature extraction model based on the attention mechanism includes four parts: graph attention layer, sorting layer, temporal attention layer, and output layer; the graph attention layer is composed of three layers of graph attention layers. The sorting layer mainly rearranges and combines the features obtained through the graph attention layer to form a sequence of microservice instances in the time dimension, which is convenient for the subsequent time encoder to extract time features; the temporal attention layer is composed of three layers of self-attention layers, and the output layer is composed of two layers of fully connected layers; the input of the model is wd H×H adjacency matrices A and wd H×S feature matrices X, where H is the number of microservice instances in the microservice system, S is the feature dimension of the microservice instance, and the output is the graph attention layer output feature vector of all microservice instances in the microservice system. The instance feature aggregates the time features of the microservice instance in the topological dependency relationship of the dependency structure and the performance index. The learning rate of the graph neural network is set to step, and the training batch size is b; The graph attention layer uses a spatial domain-based graph attention calculation method, as shown in formula (1), x u and x v is the node feature vector of instance u and instance v obtained by 2.2), A uv is the element in the adjacency matrix obtained by 2.1); W ss ∈R F*F′ The graph attention weight matrix is used to calculate the weight of the microservice instance to each neighbor instance; This weight matrix is the relationship between the input F features and the output F′ features, which plays a mapping role; uv represents the attention coefficient of instance u to instance v, σ(·) is the LeakyRELU nonlinear activation function, and a T is a learnable weight vector, and T is the transpose of the vector; its function is to multiply the intermediate feature vector after transformation and concatenation to obtain the intermediate calculation result of the attention coefficient; when calculating the attention value, x u and x v , using the attention weight matrix W ss Do the mapping and concatenate the resulting vectors; then, use the feedforward neural network a T Map the concatenated vector to real numbers; euv=σ(Auv·aT[Wssxu||Wsxv])#(1) Next, the value is normalized, and the calculation formula is shown in formula (2); α uυ is the normalized attention coefficient of instance u to instance v; exp(·) is an exponential function, exp(e uυ ) represents the original attention coefficient e uυ Perform exponential operation, ∑ represents the sum operation, N v represents the set of neighbor nodes of instance v; Formula (2) represents the sum of the attention coefficients of all neighbor instances of instance v after exponential operation, and finally obtains the attention coefficient α uυ ; Get the attention coefficient α uυ After that, the weighted sum of neighbors can be performed to obtain the graph output feature vector of microservice instance v as shown in formula (3); υ is the output feature vector of the graph attention layer of instance v, σ(·) is the LeakyRELU nonlinear activation function, ∑ represents the summation operation, N v represents the set of neighbor nodes of instance v; x u is the feature vector of instance u, W ss is the attention weight matrix, α uυ is the coefficient obtained by applying the LeakyRELU function to the neighbors of each node, which represents the contribution of microservice instance u to microservice instance v in the current microservice system state diagram; The output of the graph attention layer is based on multiple microservice systems. Each microservice system corresponds to an instance representation matrix at different times, forming a matrix sequence. In order to facilitate the processing of the subsequent temporal attention layer, the output of the graph attention layer needs to be reordered. The instance representation matrix output by the graph attention layer is Z wd,xm,xd , where wd represents the time index (wd=1,2,...,WD), xm represents the microservice instance index (xm=1,2,...,XM), and xd represents the feature dimension index (xd=1,2,...,XD); after reordering, the graph feature extraction representation matrix Z of different instances is formed I ∈R WD×XD ; The representation matrices of these different instances will enter the temporal attention layer; in the temporal attention layer, the weight matrix W is first used q ,W k ,W v Perform a linear transformation on the representation vector to obtain the query vector Q, key vector K, and value vector V, which can be specifically expressed as formula (4)(5)(6), where W q Query weight matrix, W k is the key weight matrix, W v is the value weight matrix; Q=Z i W q #(4) K=Z i IN k #(5) V=Z i W v #(6) Then, the query vector Q and the key vector K are used to calculate the representation of different nodes in the time series according to formula (7), where Z i is the graph feature extraction representation vector of instance i at time point, is the attention value of time j to time i; ((Z i W q )(Z i W q ) T ) ij It is a linear transformation of the query weight matrix and the representation vector. T is the transposition operation and then the inner product of the two transformed representation matrices is calculated to measure the correlation or similarity between nodes. F′ is the normalization factor, M ij is the position encoding matrix, which is used to solve the problem of temporal feature disappearance caused by parallel input. It is defined as formula (8). Then all attention values are softmax normalized. The specific expression is shown in formula (9). exp(·) is an exponential function. Represents the original attention coefficient Perform exponential operations, is the sum of the exponential transformation values of the attention weights of the nodes at all times; the normalized attention β is obtained ij ; According to the normalized attention value β ij The sum value vector V is used to get the node representation Z′ i , as shown in formula (10); WITH' i =β ij (WITH i IN v )#(10) In Z′ i After inputting into the multi-layer perceptron (MLP) layer for feature integration learning, the output representation of each microservice instance is a vector h. First, it is a matrix processed by the temporal attention layer, which contains the feature representation of the microservice instance at different times; when it is input into the MLP layer, the MLP can further integrate and learn these features through a series of linear transformations and nonlinear activation functions; the MLP consists of two linear layers and the ReLU activation function. For the input matrix Z′ i , first pass through the first linear layer, as shown in formula (11); where W1 is the weight matrix from the input layer to the hidden layer, b1 is the bias vector from the input layer to the hidden layer, Y1 is the intermediate variable of the hidden layer, and Y′1 is the intermediate variable of the activated hidden layer, and then pass through the ReLU activation function for nonlinear transformation; then pass through the second linear layer, the linear transformation formula is (13), where W2 is the weight matrix from the hidden layer to the output layer, b2 is the bias vector from the hidden layer to the output layer, and the final output vector h of the microservice instance v is obtained; Y1=Z′ i W1+b1#(11) Y′1=ReLU(Y1)#(12) h=Y′1W2+b2#(13) 3.2) Use the microservice system state set W to train the spatiotemporal feature extraction model AF; 3.2.1) Train the constructed spatiotemporal feature extraction network AF and use all microservice system states in the microservice system state set W as sample data; These sample data are used to generate the adjacency matrix and feature matrix in the system state set; Specifically, the adjacency matrix A is constructed based on the dependencies between microservice instances in the sample data. i , construct the feature matrix X based on the performance indicators of the microservice instances in the sample data i ; Each training input is 64 samples, and the adjacency matrix and feature matrix are used as the input values of the model. Each sample contains wd adjacency matrices and feature matrices. The adjacency matrix and feature matrix are combined to represent the state of the microservice system within 1 minute. A i is the adjacency matrix of the ith minute, X i is the feature matrix of the i-th minute; Update model parameters through forward propagation algorithm and Adam optimizer for training, and repeat input until all microservice systems are trained; 3.2.2) Repeat the process of 3.2.1) for β times. The hyperparameter β represents the number of rounds of parameter updates during model training, and its value range is 128-512. Perform multiple rounds of parameter updates on the model, where the parameters include graph attention weight matrix, query weight matrix, key weight matrix, value weight matrix, position encoding matrix, input layer to hidden layer weight matrix, hidden layer to output layer weight matrix; After the parameter update is completed, the microservice system training is completed, and the corresponding spatiotemporal feature extraction network is constructed; (4) Anomaly detection based on deep support vector data description model 4.1) Take the microservice system instance output vector h as an attribute, use the deep support vector data description model to perform anomaly classification, project the microservice instance feature vector h into the hypersphere, calculate the distance between the feature vector and the center c, and initialize the randomly selected hypersphere center point c with radius R; 4.2) Use the microservice system dataset to co-train the attention-based spatiotemporal feature extraction model and the anomaly detection model; The specific process is as follows: First, the sample data in the microservice system dataset W is input into the spatiotemporal feature extraction model, and the feature vector representation of the microservice instance is obtained through the forward propagation algorithm; At the same time, these feature vectors are input into the anomaly detection model as attributes; during the training process, the reconstruction loss and clustering loss are calculated together, as shown in formula (14); where R is the radius of the hypersphere, and the hyperparameter μ is used to balance the volume and boundary violations of the hypersphere, allowing some data to be outside the hypersphere; max{0,∥hc∥ 2 -R 2 } is a penalty term used to penalize data points outside the hypersphere; c represents the center of the hypersphere; ∥hc∥ 2 Calculate data point h o The square of the distance to the center c of the hypersphere; represents a regularization term, and λ is the regularization parameter; Represents the square of the Frobenius norm of the weight matrix W of the lth layer, which is used to measure the size of the weight matrix of this layer, that is, the complexity of the model; represents the sum of the weight matrices of all layers to ensure that the regularization term can take into account the complexity of each layer in the model; For the values of hyperparameters μ and λ, users can choose different μ and λ in the range of (0-1) to conduct experiments, and obtain a set of μ and λ that reduces the loss function of the user's current model the fastest. The data set is divided into a training set and a validation set, and the model is co-trained for each value. By adjusting the parameters of the spatiotemporal feature extraction model and the anomaly detection model, the loss function is minimized; For the spatiotemporal feature extraction model, first initialize its parameters and the initial radius of the hypersphere, and determine the center of the hypersphere through forward propagation; then, fix the radius and the center of the sphere and focus on training the spatiotemporal feature extraction network, and use the Adam optimizer to minimize the difference between the data points mapped to the surface and the center of the hypersphere; based on the parameters of the current spatiotemporal feature extraction network, recalculate the representation of the data points on the hypersphere. Users can select different percentile values in the range of [50%, 95%] to update the radius according to the effect and needs of the model. After the feature extraction of every 64 samples is completed, the distance Distance(h, c) from the output vector of all instances in the current sample to the center of the hypersphere is processed; the formula is shown in (15), where h is the instance output vector and c is the center of the sphere; first convert these distances into tensors and flatten them, then sort these distances from small to large, and after finding the distances sorted from small to large, select the distance corresponding to the percentile such as 90% as the radius to ensure that 90% of the data points are within the hypersphere; For the anomaly detection model, the feature vector output by the spatiotemporal feature extraction model is used to determine whether a data point is abnormal based on its distance from the center of the hypersphere. For the output vector of an instance, the distance Distance(h,c) from the center of the hypersphere is calculated. If this distance is greater than the radius of the hypersphere, the data point is considered an abnormal point, and points outside the hypersphere are considered abnormal. The parameters of the anomaly detection model are continuously adjusted, mainly the radius of the hypersphere. The radius of the hypersphere is adjusted by minimizing the loss function. The loss function is shown in formula (14). According to the distance from the data point to the center of the hypersphere and the current radius, the radius and all weight matrices in the spatiotemporal feature extraction model are adjusted through the Adam optimization algorithm to minimize the loss function. During the entire collaborative training process, the two models influence and optimize each other, jointly improving the feature extraction and anomaly detection performance of the microservice system; Distance(h,c)=||h-c|| 2 #(15) 4.3) Repeat step 4.2) until the radius of the hypersphere no longer changes, thus completing the spatiotemporal feature extraction model and anomaly detection model based on the attention mechanism; (5) Microservice instance feature extraction 5.1) For any microservice system w i , input its adjacency matrix A i With the feature matrix X i , use the spatiotemporal feature extraction model based on the attention mechanism to extract features and output the output vector h of all microservice instances in the microservice system; 5.2) Repeat step 5.1) until feature extraction of all microservice systems is completed; (6) Microservice instance exception classification 6.1) Traverse the instance feature set and calculate the distance between any microservice output vector h and the center point according to formula (15) to determine whether the data point is within the hypersphere. If it is within the hypersphere, it is normal, otherwise it is abnormal.
Citation Information
Cited By
Microservice abnormal node detection method and system based on hierarchical graph neural network
CN120956634A