Drift detection method for micro-service cluster and electronic device

By constructing incremental subgraphs and combining multi-layer graph and hypergraph algorithms, drifting nodes in microservice clusters are automatically detected, solving the problem of lack of automated awareness during the runtime of microservice clusters and realizing real-time architecture state awareness and improved adaptive capabilities.

CN120935070BActive Publication Date: 2025-12-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511429897.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-12-16
Estimated Expiration
2045-09-30

AI Technical Summary

Technical Problem

In existing technologies, microservice clusters lack automated drift awareness during runtime, requiring repeated manual intervention and adjustments, resulting in insufficient system flexibility and stability.

Method used

By collecting link tracing data from microservice clusters, an incremental subgraph is constructed. Feature changes are extracted using a joint embedding algorithm combining multi-layer graphs and hypergraphs. Drifting nodes in the microservice cluster are automatically detected, and the full node features of adjacent periods are compared to achieve automated drift detection.

Benefits of technology

It achieves automated detection of runtime drift in microservice clusters, breaking through the bottleneck of traditional static offline partitioning, and enabling real-time perception of architectural state changes, thereby improving the system's adaptability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935070B_ABST
    Figure CN120935070B_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods for detecting the drift of microservice cluster and electronic equipment, it is related to server technical field, including: by gathering link tracking data, identifying drift node and constructing incremental subgraph, combining multi-layer graph and hypergraph joint embedding algorithm extracts feature variation, further generates current detection period full node feature;By comparing adjacent period full node feature, the feature difference of each entity node is accurately captured, so as to determine whether the running period drift of microservice cluster occurs.The scheme realizes the automatic detection of microservice running period drift, breaks the technical bottleneck that traditional static offline division needs artificial repeated intervention, can real-time perceive architecture state change, effectively improves the adaptive ability of microservice cluster to respond to business dynamic change, guarantees system flexibility and stability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a drift detection method for a micro-service cluster and an electronic device. BACKGROUND

[0002] Cloud computing and mobile internet drive the explosive growth of enterprise business, and micro-service architecture becomes the mainstream solution to meet the flexibility, scalability and high concurrency requirements of the system due to "divide and rule". However, in related technologies, micro-service division is mostly based on static offline algorithms, which only output single division suggestions, and manual repeated intervention is required for adjustment during the running period of micro-service, lacking automatic running period drift perception. SUMMARY

[0003] The present application provides a drift detection method for a micro-service cluster and an electronic device to at least solve the problem of lacking automatic running period drift perception in related technologies during the running period of micro-service.

[0004] The present application provides a drift detection method for a micro-service cluster, comprising:

[0005] In the current detection period, link tracking data in the running process of the micro-service cluster is collected, and the link tracking data indicates the performance of the micro-service cluster in the running process;

[0006] Based on the link tracking data, an incremental subgraph is obtained, the incremental subgraph indicating a drift node in the micro-service cluster and an entity node associated with the drift node, the drift node being an entity node with a node change degree greater than a first threshold in the running period of the micro-service cluster, the node change degree indicating the degree of change of the association relationship of the entity node in the running period of the micro-service cluster;

[0007] A joint embedding algorithm of multi-layer graph and hypergraph is adopted to extract features from the incremental subgraph to obtain a feature change amount, the feature change amount indicating the node feature change of the nodes in the incremental subgraph compared with the last detection period;

[0008] The feature change amount and a first feature are fused to obtain a second feature, the first feature being a node feature of all nodes in the last detection period, and the second feature being a node feature of all nodes in the current detection period;

[0009] Based on the first feature and the second feature, drift detection is performed to obtain a detection result, the detection result indicating whether a running period drift of the micro-service cluster occurs.

[0010] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above drift detection methods for a micro-service cluster.

[0011] The drift detection method for the microservice cluster and the electronic device provided by the embodiments of the present application can identify the drift node by collecting link tracking data and construct an incremental subgraph, extract feature changes by combining a multi-layer graph and a hypergraph joint embedding algorithm, and further generate full node features in the current detection period. By comparing the full node features in adjacent periods, the feature differences of each entity node are accurately captured, so as to determine whether the microservice cluster has run-time drift. The scheme realizes the automatic detection of the microservice run-time drift, breaks the technical bottleneck of the traditional static offline division requiring repeated human intervention, can realize real-time perception of the architecture state change, effectively improves the adaptive ability of the microservice cluster to respond to the business dynamic change, and guarantees the flexibility and stability of the system. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] Figure 1 A drift detection system for a microservice cluster is provided for the embodiments of the present application.

[0014] Figure 2 A flowchart of a drift detection method for a microservice cluster is provided for the embodiments of the present application. Figure 1 ;

[0015] Figure 3 A flowchart of a drift detection method for a microservice cluster is provided for the embodiments of the present application. Figure 2 ;

[0016] Figure 4 A flowchart of a drift detection method for a microservice cluster is provided for the embodiments of the present application.

[0017] Figure 5 A flowchart of a microservice routing decision is provided for the embodiments of the present application.

[0018] Figure 6 A flowchart of a microservice merging decision is provided for the embodiments of the present application.

[0019] Figure 7 A flowchart of a microservice drift detection and merging is provided for the embodiments of the present application.

[0020] Figure 8 A flowchart of a microservice architecture processing is provided for the embodiments of the present application.

[0021] Figure 9A structural schematic diagram of a drift detection device for a micro-service cluster provided by an embodiment of the present application is shown in the figure.

[0022] Figure 10 A structural schematic diagram of an electronic device provided by the present application is shown in the figure. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be clearly and completely described in the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0024] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0025] Cloud computing and mobile internet drive the explosive growth of enterprise business, and micro-service architecture becomes the mainstream solution to meet the flexibility, scalability and high concurrency requirements of the system due to “division and autonomy”. However, in related technologies, micro-service division is mostly static offline algorithm, which only outputs single division suggestion, and manual repeated intervention is required during the running period of micro-service, lacking automatic running period drift perception. The drift detection method for micro-service cluster provided by the present application can detect and rollback during the running process of the micro-service cluster, which is a key link in the whole life cycle management of micro-service architecture. After the micro-service is split and deployed, the method continuously collects data, incrementally calculates model embedding vectors, quantitatively analyzes vector distribution differences, identifies the “drift” phenomenon caused by service coupling relationship deviating from the initial splitting benchmark due to business changes, traffic fluctuations, etc., and triggers the closed-loop management mechanism of automatic merging, data synchronization and routing adjustment. The core goal is to break through the limitations of traditional “one-time splitting”, realize the observability and controllability of architecture evolution, and ensure the continuous effectiveness of “stable splitting” after micro-service splitting. In order to enable a person skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0026] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the drift detection method for the micro-service cluster depends, the specific application environment architecture or specific hardware architecture is described herein. For reference Figure 1 , Figure 1This is a drift detection system for microservice clusters. The system includes server 101 and server 102, and a communication connection is established between server 101 and server 102. Server 101 deploys the microservice cluster; server 102 is used to obtain link tracing data during the operation of the microservice cluster from server 101, thereby enabling drift detection of the microservice cluster.

[0027] Figure 2 A flowchart illustrating the drift detection method for microservice clusters provided in this application embodiment. Figure 1 ,like Figure 2 As shown, embodiments of this application provide a drift detection method for microservice clusters. The method is described in detail below:

[0028] S201. During the current detection period, collect link tracing data during the operation of the microservice cluster. The link tracing data indicates the performance of the microservice cluster during operation.

[0029] In this embodiment, the microservice cluster includes multiple microservices, which are used to run an application or platform to achieve its functionality. During the operation of the microservice cluster, drift may occur, meaning that the actual state of a microservice (such as node coupling relationships, call characteristics, and business associations) gradually deviates from its initial or previous cycle-established baseline state. This is caused by factors such as business iteration and traffic changes, which may lead to performance degradation and coupling issues. Therefore, detection and repair are needed to stabilize the microservice cluster. By performing drift detection during the operation of the microservice cluster, changes can be detected in a timely manner, allowing for timely adjustments to ensure its performance.

[0030] The trace data consists of automatically generated and reported trace data during the operation of the microservice cluster using Open Telemetry (an open-source observability framework). The trace data includes service names, operation names, call relationships, start / end timestamps, and execution time within the microservice cluster. The detection period can be arbitrary, for example, one day or one week.

[0031] S202. Based on the link tracing data, obtain the incremental subgraph. The incremental subgraph indicates the drifting nodes in the microservice cluster and the entity nodes associated with the drifting nodes. The drifting nodes are entity nodes whose node change degree is greater than the first threshold during the operation of the microservice cluster. The node change degree indicates the degree of change in the association relationship of entity nodes during the operation of the microservice cluster.

[0032] In the embodiments of the present application, each microservice in the microservice cluster corresponds to at least one entity node, which is a node corresponding to a function, a class or an interface, and the function, the class or the interface corresponding to the microservice is used to implement the running of the microservice. During the running of the microservice cluster, not all microservices or nodes in the microservices may drift, therefore, based on the link tracking data, only the drift nodes and the entity nodes associated with the drift nodes that have drifted are obtained, so as to reduce the amount of data required to be obtained, and thus ensure the detection efficiency. The drift node refers to an entity node whose node association relationship (such as static call, dynamic interaction, coupling change, semantic similarity, etc.) changes significantly during the running of the microservice cluster, resulting in a change degree exceeding a first threshold value; the first threshold value is an arbitrary numerical value.

[0033] S203, a joint embedding algorithm of a multi-layer graph and a hypergraph is adopted to perform feature extraction on the incremental subgraph, to obtain a feature change amount, which indicates the node feature change of the nodes in the incremental subgraph compared with the last detection period.

[0034] In the embodiments of the present application, the incremental subgraph includes the drift nodes and the associated entity nodes. Since the association relationship of the drift nodes changes, and the node feature of any node is related to itself and the associated nodes, the node features of the drift nodes and the associated entity nodes in the incremental subgraph change, and the feature extraction method can be used to determine the node feature change amount of each node in the incremental subgraph. The joint embedding algorithm of the multi-layer graph and the hypergraph refers to constructing a multi-layer graph network and a hypergraph, and then combining the multi-layer graph network and the hypergraph to obtain the feature change amount of the incremental subgraph. The feature change amount can be represented in any form, for example, the feature change amount is represented in the form of an embedding vector.

[0035] S204, the feature change amount and the first feature are fused to obtain a second feature, the first feature being the node feature of all nodes in the last detection period, and the second feature being the node feature of all nodes in the current detection period.

[0036] In the embodiments of the present application, all nodes refer to all nodes in the microservice cluster, and the first feature is the node feature of each node in the microservice cluster in the last detection period. Considering that only the node features of the nodes in the incremental subgraph change in the current detection period, and the node features of the remaining nodes do not change, the node feature of all nodes in the current detection period, i.e., the second feature, can be obtained based on only the feature change amount and the node feature of all nodes in the last detection period. The first feature and the second feature can be represented in any form, for example, the first feature and the second feature are represented in the form of an embedding vector.

[0037] S205, drift detection is performed based on the first feature and the second feature, and a detection result is obtained, which indicates whether the runtime drift of the microservice cluster occurs.

[0038] In the embodiment of the application, when the first feature in the last detection period and the second feature in the current detection period are obtained, the feature change characteristics of each entity node in the two detection periods can be compared, and it is determined whether the runtime drift of the microservice cluster occurs compared with the last detection period, that is, the detection result is obtained, which realizes the scheme of automatically performing drift detection in the running process of the microservice cluster, and can timely detect the case that the runtime drift of the microservice cluster occurs.

[0039] In the scheme provided by the embodiment of the application, the drift node is identified by collecting link tracking data and constructing an incremental subgraph, and the feature change is extracted by combining a multi-layer graph and a hypergraph joint embedding algorithm to further generate the full node feature in the current detection period; by comparing the full node features in adjacent periods, the feature difference of each entity node is accurately captured, so as to determine whether the runtime drift of the microservice cluster occurs. The scheme realizes the automatic detection of the runtime drift of the microservice, breaks through the technical bottleneck of traditional static offline division requiring repeated human intervention, can realize real-time perception of architecture state change, effectively improves the adaptive ability of the microservice cluster to respond to business dynamic change, and guarantees the flexibility and stability of the system.

[0040] On the basis of the above Figure 2 The embodiment of the application provides a drift detection method for a microservice cluster, and the drift detection method is described in detail as follows. Figure 3 The drift detection method for the microservice cluster provided by the embodiment of the application is shown in the flowchart Figure 1 As shown in Figure 3 The embodiment of the application provides a drift detection method for a microservice cluster, and the drift detection method is described in detail as follows.

[0041] S301, in a current detection period, link tracking data in the running process of the microservice cluster is collected, and the link tracking data indicates the performance of the microservice cluster in the running process.

[0042] In the embodiment of the application, the data collection module can be used to collect the link tracking data. For example, an application integrates Open Telemetry SDK (a software development kit of a distributed tracking data collection and stream processing architecture) or Open Telemetry Agent (an agent of a distributed tracking data collection and stream processing architecture), generates and reports Trace (tracking) data automatically when the microservice cluster runs, and the Trace data is the link tracking data.

[0043] S302, acquire a node change degree based on the link tracking data, the node change degree indicating a change degree of an association relationship of an entity node in a running period of the micro service cluster.

[0044] In the embodiment of the present application, the link tracking data can reflect the calling conditions between the entity nodes in the micro service cluster during the running process, and then based on the link tracking data, the change conditions of the association relationship of each entity node can be acquired, that is, the node change degree is obtained. The association relationship of the entity node indicates the calling relationship or the dependency relationship between the entity node and other nodes.

[0045] In a possible implementation manner, the step S302 comprises: acquiring a first calling relationship of the entity node in the micro service cluster in a previous detection period; acquiring a calling relationship of the entity node in the micro service cluster in a current detection period based on the link tracking data; and determining the node change degree based on the difference between the first calling relationship and the second calling relationship.

[0046] In the embodiment of the present application, in each detection period, the calling relationship of each entity node corresponding to the micro service cluster in each detection period can be acquired based on the link tracking data in each detection period. In the current detection period, the calling relationship of the entity node in the micro service cluster in the previous detection period can be directly acquired, and the calling relationship of the entity node in the micro service cluster in the current detection period can be acquired based on the link tracking data in the current detection period, so as to compare the change conditions of the calling relationship of each entity node in the two detection periods, and then determine the node change degree, so that the node change degree can accurately indicate the change conditions of the calling relationship of each entity node, to ensure the accuracy of the node change degree. In the embodiment of the present application, in the process of detecting the drift of the micro service cluster every interval of a detection period, after the calling relationship of each entity node in the micro service cluster is acquired in each detection period, the calling relationship is cached for calling in the next detection period. After the first calling relationship and the second calling relationship are obtained, the difference between the first calling relationship and the second calling relationship is quantitatively analyzed, such as the calling link of the new message, the change of the calling frequency, the fluctuation of the association weight, and the change value of the node association state is obtained. The calling relationship not only refers to the calling relationship between the micro services, but also includes the calling relationship in multiple dimensions such as static, dynamic, change and semantics.

[0047] S303, determine a drift node and an entity node associated with the drift node based on the node change degree and the link tracking data, the drift node being an entity node with a node change degree greater than a first threshold in a running period of the micro service cluster.

[0048] In the embodiments of the present application, the condition that the node change degree is greater than the first threshold value is used as the condition for screening the drift node. After the node change degree of each entity node is obtained, the drift node can be determined from the entity nodes according to the condition.

[0049] In the embodiments of the present application, the entity nodes associated with the drift node include downstream nodes called by the drift node, upstream nodes calling the drift node, and nodes associated through shared resources (such as database tables and configuration files); the link tracking data ensures that the identification of the associated nodes covers static dependencies (such as code calls) and dynamic interactions (such as runtime request links), avoiding missing implicit associations. Both the "core drift node with significant changes" can be accurately locked, and the influence range of the drift can be identified and grasped through the associated nodes, providing accurate analysis objects for subsequent construction of the incremental subgraph (focusing on drift-related nodes rather than all nodes), making the feature extraction more focused on the key areas of architecture changes, and improving the efficiency and accuracy of drift detection.

[0050] In a possible implementation manner, the entity nodes associated with the drift node are also referred to as the transaction influence domain corresponding to the drift node, and the associated nodes of the drift node are screened according to a preset screening range. For example, if the preset range is 3-Hop (hop) neighbor representation, the entity nodes associated with the drift node include 1-hop neighbors (upstream and downstream nodes of the domain drift node), 2-hop neighbors (upstream nodes of the 1-hop neighbors), and 3-hop neighbors (upstream nodes of the 2-hop neighbors).

[0051] S304, based on the drift node and the entity nodes associated with the drift node, an incremental subgraph is obtained, and the incremental subgraph indicates the drift node and the entity nodes associated with the drift node in the microservice cluster.

[0052] In the embodiments of the present application, the drift node is taken as the core, and the entity nodes associated with the drift node (such as upstream calling nodes, downstream called nodes, and shared resource associated nodes) are combined to form a node set of the incremental subgraph. At the same time, the calling relationship (such as static dependencies, dynamic interactions, and multi-dimensional associations) between these nodes is reserved as the edge of the subgraph to form a complete local association structure, that is, the incremental subgraph. In this way, only the drift node and the associated nodes are concerned, rather than all nodes, which can filter irrelevant node interference, accurately capture the local area with significant changes in the architecture, and make the subsequent feature extraction and drift detection more focused on the real change core.

[0053] It should be noted that the embodiments of the present application are described by taking the combination of node change degree to determine the drift node and then obtain the incremental subgraph as an example, and in another embodiment, the steps S302-S304 described above are not required to be performed, but other ways are adopted to obtain the incremental subgraph based on the link tracking data.

[0054] S305, obtain multi-dimensional data corresponding to the incremental subgraph, the multi-dimensional data indicating source code, performance in a running process, historical change situation and business meaning of nodes in the incremental subgraph.

[0055] In the embodiments of the present application, the multi-dimensional data of each node in the incremental subgraph is obtained, so as to provide a basis for subsequent extraction of node features of each node in the incremental subgraph. The source code can indicate the underlying code information corresponding to each entity node in the incremental subgraph, reflecting the static structure of the node. The performance in the running process refers to the real-time performance index of the node in the incremental subgraph when running in the current detection period, reflecting the dynamic running state of the node. The historical change situation indicates the historical modification record of the code, configuration and other assets associated with the node in the incremental subgraph, reflecting the change trajectory of the node. The business meaning indicates the business function and business association information corresponding to the node in the incremental subgraph, reflecting the business attribute of the node, avoiding the "business disconnection" caused by the analysis from the technical dimension.

[0056] In a possible implementation manner, the multi-dimensional data includes code data, link tracking data, historical change data and semantic data, the code data indicating the source code, the link tracking data indicating the performance, the historical change data indicating the historical change situation, and the semantic information indicating the business meaning.

[0057] S306, based on the multi-dimensional data, construct a multi-layer graph network and a hypergraph, the multi-layer graph network including a plurality of graph layers, the graph layer including a plurality of entity nodes corresponding to the incremental subgraph and the association relationship between the plurality of entity nodes, and the hypergraph including a plurality of hyperedges, the hyperedge indicating the entity nodes belonging to the same business process in the incremental subgraph.

[0058] In the embodiment of the present application, since the multi-dimensional data can describe the incremental subgraph from multiple dimensions, a multi-layer graph network is constructed based on the multi-dimensional data, and the multiple graph layers contained in the multi-layer graph network correspond one-to-one to the multiple dimensions corresponding to the multi-dimensional data, so that each graph layer in the multi-layer graph network can reflect the incremental subgraph described by the data in one dimension of the multi-dimensional data, so as to fully describe the relationship of each entity node corresponding to the incremental subgraph in different dimensions, so as to enrich the basis for drift detection of the micro-service cluster. Considering that the implementation of the same business process during the running of the incremental subgraph may involve multiple different entity nodes, which can also reflect the association relationship between different entity nodes, a hypergraph is constructed based on the multi-dimensional data, so as to fully describe the entity nodes in the incremental subgraph involved in the implementation of each business process. In the multi-layer graph network (Multiplex Graph Network), the entity nodes corresponding to the multiple graph layers are the same. In the hypergraph (Hyper Graph), each hyperedge can connect any number of entity nodes, and the number of entity nodes contained in different hyperedges may be the same or different. The entity nodes corresponding to the incremental subgraph are the nodes corresponding to the business entities in the incremental subgraph, and the business entities are methods, classes, APIs or sub-services, etc. in the incremental subgraph. The business process defines a series of actions of a business execution.

[0059] In a possible implementation manner, the multi-dimensional data includes code data, link tracking data, historical change data and semantic data; the step S306 includes: constructing a multi-layer graph network based on the code data, the link tracking data, the historical change data and the semantic data; and constructing a hypergraph based on the link tracking data.

[0060] In the embodiment of the present application, the code data, the link tracking data, the historical change data and the semantic data can describe the incremental subgraph from different dimensions, so the multi-layer graph network is constructed based on the four kinds of data, so that each graph layer in the multi-layer graph network can describe the incremental subgraph from different dimensions, and through the targeted use of different dimensional data, the precise modeling of technical association and business association is realized, the data source isolation of technical association (multi-layer graph) and business association (hypergraph) is realized, the cross interference of data is avoided, the double value of the link tracking data supports the technical interaction modeling of the dynamic layer, and provides the actual running track of the business process for the hypergraph, ensures that the business association is consistent with the real running state, and provides precise and layered feature input for the subsequent joint embedding algorithm.

[0061] Optionally, the process of constructing the multi-layer graph network comprises: respectively parsing code data, link tracking data, historical change data and semantic data to obtain a graph layer corresponding to the code data, a graph layer corresponding to the link tracking data, a graph layer corresponding to the historical change data and a graph layer corresponding to the semantic data; and combining the graph layer corresponding to the code data, the graph layer corresponding to the link tracking data, the graph layer corresponding to the historical change data and the graph layer corresponding to the semantic data to form the multi-layer graph network.

[0062] In the embodiments of the present application, the code data, the link tracking data, the historical change data and the semantic data can describe the original service from different dimensions, and therefore, the four kinds of data are parsed respectively to fully understand the original service according to the data of each dimension, so as to obtain the graph layer corresponding to each dimension, and then the multiple graph layers are superimposed to obtain the multi-layer graph network, so as to ensure the accuracy of the obtained multi-layer graph network; each graph layer strictly corresponds to a specific data type, avoiding cross interference of the association relationship of different dimensions, so that the subsequent feature extraction can be accurately positioned to a specific association type (such as analyzing whether the drift is caused by dynamic call change); the design of independent graph layers facilitates the addition of new dimensions (such as the addition of a security dependency graph layer later), without the need to reconstruct the overall structure; the multi-layer aggregation ensures that the node features cover the four dimensions of static structure, dynamic operation, change history and semantic function, providing comprehensive association relationship input for the subsequent joint embedding algorithm, and improving the accuracy of drift detection.

[0063] Optionally, the process of constructing the hypergraph comprises: extracting operation information corresponding to at least one business process and the number of occurrences of the at least one business process from the link tracking data, the operation information indicating operations performed when implementing the business process; traversing the operations in the link tracking data to determine the services to which the operations belong; and constructing a hyperedge by using the services to which the operation information corresponding to the same business process belongs, and determining the number of occurrences of the business process as the weight of the hyperedge corresponding to the business process, to obtain the hypergraph. In the embodiments of the present application, based on the link tracking data, the operations involved in the business process are determined by means of operation recognition, and then the services involved in each business process are determined, and the number of occurrences of the business process is combined to construct the hypergraph, so as to enrich the content contained in the hypergraph and improve the accuracy of the hypergraph. The number of occurrences refers to the number of occurrences of the business process. The hyperedge represents a set of nodes, rather than a call sequence.

[0064] In the embodiments of the present application, the link tracking data can reflect the entity nodes involved in the implementation of each business process, therefore, a hypergraph is constructed based on the link tracking data to illustrate the entity nodes involved in each business process, and then the association relationship between each entity node is reflected. The hypergraph takes a complete business use case (i.e. a business process) as an indivisible edge, and through distributed link tracing (Tracing) and hypergraph modeling (Hypergraph Modeling), the business use case (rather than a single call pair) is taken as a basic construction unit (Hyperedge) of the hypergraph, which is closer to the business semantics and integrity requirements.

[0065] S307, feature extraction is performed on the multi-layer graph network and the hypergraph to obtain node features of nodes in the incremental subgraph.

[0066] In the embodiments of the present application, the feature extraction is performed on the multi-layer graph network and the hypergraph to fully consider the relationship between each entity node in the incremental subgraph, and then the node features of each entity node in the incremental subgraph are obtained. The node features can be represented in any form, for example, the node features are represented in the form of embedding vectors.

[0067] In a possible implementation manner, the step S307 includes: performing feature extraction on a plurality of graph layers in the multi-layer graph network to obtain graph layer features corresponding to the plurality of graph layers; performing weighted fusion on the graph layer features corresponding to the plurality of graph layers to obtain fused features; performing feature extraction on the hypergraph to obtain hypergraph features; splicing the fused features and the hypergraph features to obtain spliced features; and performing linear transformation on the spliced features to obtain the node features of the nodes in the incremental subgraph.

[0068] In the embodiments of the present application, through the hierarchical extraction and weighted fusion of the multi-layer graph network, the association details of a single graph layer are not lost, the importance of different dimensions is embodied through weight adjustment, the introduction of the hypergraph features avoids the problem of focusing only on technical association and deviating from the business scenario, the node features reflect both “technical attributes” and “business roles”, and finally linear transformation is adopted to ensure that the final node features are in a unified vector space. Through the four-step process of “hierarchical extraction - weighted fusion - cross-structure splicing - linear transformation”, the graph structure information is converted into quantifiable node features (such as embedding vectors), which not only retains multi-dimensional association details but also integrates business process attributes.

[0069] S308, the difference between the node features of the nodes in the incremental subgraph and the node features of the nodes in the incremental subgraph in the last detection period is determined as a feature change amount, and the feature change amount indicates the change of the node features of the nodes in the incremental subgraph compared with the last detection period.

[0070] In the embodiment of the present application, the node features of each node in the incremental subgraph are incremented by the node features in the last detection period, so as to determine the change of the node features of each node in the incremental subgraph, and further determine the feature change amount. Only the nodes in the incremental subgraph are calculated for the feature change amount (not all nodes), which reduces the calculation overhead and improves the efficiency and pertinence of drift detection.

[0071] It should be noted that the embodiment of the present application takes the multi-dimensional data based on the incremental subgraph to obtain the feature change amount as an example for description. In another embodiment, the steps S305-S308 are not performed, but other ways are adopted. The joint embedding algorithm of multi-layer graph and hypergraph is adopted to extract features from the incremental subgraph to obtain the feature change amount.

[0072] S309, fusing the feature change amount and the first feature to obtain a second feature, the first feature being the node feature of all nodes in the last detection period, and the second feature being the node feature of all nodes in the current detection period.

[0073] In the embodiment of the present application, since the first feature is the node feature of all nodes in the micro-service cluster in the last detection period, and the feature change amount is the node feature increment of the nodes changed in the current period of the micro-service cluster, the feature change amount and the first feature are fused, and the node feature of all nodes in the micro-service cluster in the current detection period can be obtained.

[0074] In one possible implementation manner, the second feature satisfies the following relationship:

[0075]

[0076] wherein, is used to represent the second feature; is used to represent the first feature, is used to represent the feature change amount; is used to represent a smoothing coefficient, and the value range is [0, 1].

[0077] In the embodiment of the present application, the incremental MHGE (multi-layer graph and hypergraph joint embedding algorithm) mechanism is adopted, and it is not necessary to calculate the node feature of all nodes each time. Only the local nodes that have changed significantly are updated for the node feature in each detection period, which greatly reduces the calculation overhead and uses the dynamic change characteristics of micro services.

[0078] S310, based on the first feature and the second feature, obtaining a first probability distribution corresponding to the first feature and a second probability distribution corresponding to the second feature.

[0079] The first feature and the second feature are node features of all nodes in the microservice cluster in a previous detection period and a current detection period, and the first feature and the second feature are discrete node feature sets. Therefore, a statistical learning method is adopted to convert the discrete node features into continuous probability distributions to reflect the overall state of the all-node features, which facilitates subsequent comparison of the first feature and the second feature. For example, the first probability distribution and the second probability distribution are Gaussian distributions, or probability distributions obtained through kernel density estimation.

[0080] In the embodiments of the present application, the first probability distribution or the second probability distribution can be obtained by kernel density estimation, and the bandwidth is automatically selected when performing kernel density estimation. In addition, a 128-dimensional space is used for frequency division bucketing to pre-process high-dimensional features, and the feature values of each dimension are divided into several buckets according to the frequency (for example, each bucket contains the same number of node features).

[0081] S311, determining a distribution difference value of the first probability distribution and the second probability distribution.

[0082] In the embodiments of the present application, when one probability distribution and the second probability distribution are obtained, the distance between the two probability distributions is calculated by a mathematical method, and then the distribution difference value of the two probability distributions is obtained to reflect the feature change characteristics of the all nodes in the microservice cluster in the adjacent two detection periods.

[0083] In one possible implementation, the KL (Kullback-Leibler Divergence, relative entropy) divergence is used to determine the distribution difference value between the first probability distribution and the second probability distribution. The smaller the KL divergence value, the more similar the two probability distributions are; the larger the KL divergence value, the less similar the two probability distributions are, and the KL divergence satisfies the following relationship:

[0084]

[0085] wherein, is used to represent the distribution difference value of the first probability distribution and the second probability distribution, is used to represent the first probability distribution; is used to represent the second probability distribution, is used to represent a sample point in the first probability distribution and the second probability distribution; is used to represent a sample point is used to represent the probability density or probability value of the sample point in the first probability distribution; is used to represent a sample point is used to represent the probability density or probability value of the sample point in the second probability distribution.

[0086] S312, based on the distribution difference value, obtaining a detection result, the detection result indicating whether the micro-service cluster has a running period drift.

[0087] In the embodiment of the present application, since the distribution difference value can reflect the feature change characteristics of all nodes in the micro-service cluster within two adjacent detection periods, based on the distribution difference value, whether the micro-service cluster has a running period drift can be determined, so as to ensure the accuracy of the obtained detection result. In the embodiment of the present application, the probability distribution has a smoothing effect on the abnormal characteristics of a single node, reduces the interference of single-point abnormality on the overall determination, and improves the reliability of the detection result; by converting the node feature change into an architecture drift conclusion, the automatic and accurate determination of the micro-service running period drift is realized.

[0088] In a possible implementation manner, the step S312 comprises: obtaining a per-second query relative growth rate, the per-second query relative growth rate indicating a change amplitude of the per-second query in the current detection period compared with the last detection period in the running process of the micro-service cluster; and based on the distribution difference value and the per-second query relative growth rate, obtaining a detection result.

[0089] In the embodiment of the present application, considering that the micro-service cluster will appear normal business traffic fluctuation during running, these fluctuations will also cause the features of each entity node in the micro-service cluster to change, and the relative growth rate can distinguish the normal business traffic fluctuation from the traffic anomaly caused by the architecture drift, therefore, the detection result is obtained in combination with the distribution difference value and the per-second query relative growth rate, which can improve the accuracy of drift detection, avoid single-index misjudgment, and provide more reliable determination basis for subsequent starting of the drift repair process.

[0090] Optionally, the process of obtaining the detection result comprises: obtaining a first detection result in a case that the distribution difference value is greater than a second threshold value and the relative growth rate of the query per second is greater than a third threshold value, the first detection result indicating that the microservice cluster has runtime drift; obtaining a second detection result in a case that the distribution difference value is less than or equal to the second threshold value, or in a case that the relative growth rate of the query per second is less than or equal to the third threshold value, the second detection result indicating that the microservice cluster does not have runtime drift. The second threshold value is a drift judgment threshold value of the distribution difference value, used to measure whether the deviation of the feature distribution exceeds the normal range, for example, the second threshold value is 0.3. The third threshold value is a drift judgment threshold value of the relative growth rate of the query per second, used to measure whether the fluctuation amplitude of the business traffic exceeds the normal range, for example, the third threshold value is 0.4. In the embodiment of the present application, only in the case that the distribution difference value is greater than the second threshold value and the relative growth rate of the query per second is greater than the third threshold value, it can be determined that the microservice cluster has runtime drift, the judgment logic is simple and clear, without complex fuzzy judgment, which is convenient for engineering implementation, and the threshold value can be dynamically adjusted according to the business scenario, which has high flexibility and can also ensure the accuracy of drift detection.

[0091] For example, the process of drift detection for the microservice cluster is as shown in Figure 4 For the first probability distribution and the second probability distribution, the probability density estimation is performed on the first probability distribution and the second probability distribution to convert the discrete features into continuous probability distributions, the KL divergence between the two probability distributions is calculated to quantify the difference between them, and the calculated KL divergence is compared with the second threshold value to determine whether there is a significant distribution difference.

[0092] It should be noted that the embodiment of the present application is described by taking the way of calculating the probability distribution to obtain the detection result as an example, and in another embodiment, without performing the steps S310-S312, other ways are adopted to perform drift detection based on the first feature and the second feature to obtain the detection result.

[0093] S313, in a case that the first detection result is obtained, starting a microservice merging process to merge the microservice cluster, the first detection result indicating that the microservice cluster has runtime drift.

[0094] In the embodiment of the present application, in a case that the microservice cluster is detected to have runtime drift, the microservice merging process is started to eliminate the abnormal association of the drift node, reconstruct the stable service boundary, integrate the scattered drift node and the associated entity, make the microservice cluster return to the stable architecture from the deviation state, and finally restore the stability and business carrying capacity of the microservice cluster.

[0095] In a possible implementation, the process of merging the microservice cluster includes: locating the code repository of the drifting service; automatically merging the code branches according to the class / method call relationship of the static layer, and resolving the conflicts in interface definition, dependency injection and the like by using a "baseline version priority + incremental change adaptation" strategy; verifying the call relationship integrity of the merged code, ensuring that there is no isolated node and avoiding introducing new structural defects due to merging; determining the database tables associated with the drifting service, identifying shared tables and exclusive tables from the database tables; directly migrating the exclusive tables to the database instance of the merged service; using synchronous writing to the new and old tables for the shared tables to ensure data consistency during migration; performing foreign key conversion for the associated tables to reconstruct the association relationship between the tables; generating an incremental migration script, automatically adding a verification trigger to ensure that there is no data loss or tampering before and after migration; analyzing the dependency relationship of the microservice, merging the third-party dependency relationship, removing redundant dependency items, and ensuring dependency version compatibility; adjusting the performance configurations such as thread pool size and timeout time, updating the configuration of the registration center, and enabling the merged service to be normally registered and discovered; adopting a double-write switching mechanism to synchronize the historical data in the old database of the drifting service to the new database of the merged service, and using offline batch execution in the synchronization process to reduce the impact on the running business; starting a real-time synchronization program to detect the write operation of the old database, and synchronizing the new data generated during synchronization to the new database in real time to ensure that the data of the new and old databases is consistent in real time; adjusting the traffic routing rules to gradually direct the read requests of the client from the old service cluster to the new service cluster after merging; after the read traffic is stable, switching the write request to the new service cluster to complete the full traffic migration; modifying the routing configuration file, deleting the upstream nodes of the old service cluster, and uniformly pointing the routing target of the client request to the instance address of the merged service cluster; verifying whether the traffic completely flows to the new cluster and whether there is no request access to the old service cluster; after confirming that the new cluster is running stably, stopping the old service instance, releasing the server resources, and completing the whole merging and migration process. In the embodiment of the application, the traffic scheduling strategy is performed at 5% of the initial traffic, 20% during verification, and 100% during switching, the automatic rollback mechanism is to monitor the key indicators (error rate > 1% or RT > 500ms) and automatically trigger the rollback script, the service start order is controlled, and the database migration is checked in advance to realize dependency coordination.

[0096] For example, the microservice routing decision process in the routing adjustment process is as shown in Figure 5 The client sends a request, the routing decision module determines whether the request should be routed to the new merged service cluster or the old service cluster, the request is processed by the new merged service cluster and a response is returned if the request is routed to the new merged service cluster, and the request is processed by the old service cluster and a response is returned if the request is routed to the old service cluster.

[0097] For example, the microservice merging decision process is as shown in Figure 6As shown, the drift services A and B are identified; the merging decision maker receives the related information of the drift services A and B, analyzes their code, data, configuration, and other associations; the merging decision maker performs code merging operation to integrate the code of the two drift services; the merging decision maker performs data migration to migrate the database tables, cache, and other data related to the drift services to the storage corresponding to the merged service; and the merging decision maker performs configuration integration to unify the system configuration of the merged service.

[0098] In the scheme provided by the embodiments of the present application, the drift nodes are identified by collecting link tracking data and constructing an incremental subgraph, and the feature changes are extracted by combining a multi-layer graph and a hypergraph joint embedding algorithm to further generate full-node features of the current detection period. By comparing the full-node features of adjacent periods, the feature differences of each entity node are accurately captured to determine whether the microservice cluster has run-time drift. The scheme realizes the automatic detection of microservice run-time drift, breaks through the technical bottleneck of traditional static offline division requiring repeated human intervention, can realize real-time perception of architecture state changes, effectively improves the adaptive ability of the microservice cluster to respond to business dynamic changes, and guarantees the flexibility and stability of the system.

[0099] On the basis of the above-mentioned embodiments, the embodiments of the present application further provide a microservice drift detection and merging process, as shown in Figure 7 The method includes: continuously collecting Trace data of microservice call links; running an incremental multi-layer graph and hypergraph joint embedding (MHGE) algorithm every week; calculating the KL divergence between the baseline probability distribution and the current probability distribution based on the updated embedding features; performing drift detection to determine whether the conditions of KL divergence greater than 0.15 and QPS (Queries Per Second) growth exceeding 50% are met; if the conditions are met, triggering an alarm; if the conditions are not met, waiting for the next detection period; after the alarm is triggered, performing a one-key merging service operation to merge the drift microservices; using a double-write switching mechanism to write data to the old service and the merged new service at the same time to ensure data consistency; and adjusting the routing configuration to gradually migrate traffic to the merged service cluster.

[0100] It should be noted that the embodiments of the present application only take the drift detection and merging of the microservice cluster as an example for description, and in another embodiment, before this, a microservice division method can be taken to divide the original services to obtain a microservice cluster, and then the drift detection method provided in the above-mentioned embodiments is used for detection during the operation of the microservice cluster. The embodiments of the present application provide a microservice architecture processing whole process, as shown in Figure 8 The method includes:

[0101] Step 1, extract information from source code repository, build static matrix; get change information from Git repository, build change matrix; build dynamic matrix by collecting Trace data; build semantic matrix by annotating documents.

[0102] Step 2, build multi-layer graph network based on static matrix, change matrix, dynamic matrix, and semantic matrix, and build hypergraph based on Trace data.

[0103] Step 3, perform MHGE joint embedding on multi-layer graph network and hypergraph to generate 128-dimensional node vectors.

[0104] Step 4, make clustering decisions based on node vectors to generate node cluster information and determine the clustering results of microservices.

[0105] Step 5, generate code based on clustering results, and prepare repositories and scripts, etc.

[0106] Step 6, perform gray-scale deployment on the generated code, and let part of the traffic access the newly deployed service first.

[0107] Step 7, perform drift detection and calculate KL divergence to determine whether drift has occurred.

[0108] Step 8, if drift is detected, trigger automatic rollback or merge operation to ensure the stability of the microservice architecture.

[0109] In the embodiments of the present application, the hypergraph encapsulates the "one service use case" as a hyperedge, solving the problem of "group coupling" division in the traditional graph model; the multi-layer graph unifies the embedding of static, dynamic, change, and semantic heterogeneous information in layers, avoiding the loss of information caused by simple weighting, and the two are coordinated rather than superimposed; the embedding stage simultaneously preserves the integrity of the multi-layer graph network and the hypergraph hyperedge, generates a single low-dimensional vector, and solves the mutual exclusion problem of hypergraph and multi-layer graph embedding; the business constraint injection clustering mechanism is adopted: the "Must-link / Cannot-link" constraints (derived from cross-library transactions, compliance domains, etc.) are injected into clustering in the form of configurable rules, and the rules and algorithms are deeply integrated through Constrained Louvain + Pareto Frontier to replace post-hoc manual verification; the running period closed-loop management mechanism is adopted, and the trace is continuously collected online, and the MHGE is re-run incrementally every week; the KL divergence + business KPI is used to jointly determine the drift threshold, and the rollback script is automatically generated and grayly takes effect, forming a "drift detection-automatic rollback" closed loop; the non-runtime dimension quantification modeling mechanism is adopted, the "change coupling" (Git co-change / commit frequency) and "semantic similarity" (SBERT / ERNIE-Sim cosine value) are quantified and included in the graph weight, and the calling graph is jointly modeled to reduce the maintenance cost after disassembly; the double-index Pareto optimization is adopted: taking "cross-service transaction rate" and "future drift risk" as objective functions, a set of compromise solutions is provided through Pareto optimization to solve the contradiction between "accurate disassembly but difficult to maintain" and "stable disassembly but excessive coupling"; the industrial-grade tool chain adaptation is adopted, and the hypergraph partition tool chain (KaHyPar / hMETIS) is introduced into the microservice field, and is connected with Zipkin, Git, and NLP, supporting million-node scale and realizing minute-level production-level solutions. Based on the scheme provided in the embodiments of the present application, the hypergraph ensures the integrity of the service use case, and the number of distributed transactions is reduced by more than 80% on average; the multi-layer graph fuses four-dimensional coupling information, and the POD granularity disassembly error is reduced by 45%; online drift detection + rollback enables the architecture evolution to have CI / CD-level automation and observability; the system supports million-level nodes and minute-level calculation, meeting the needs of large enterprises.

[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.

[0111] Figure 9 The structure schematic diagram of the drift detection device for the microservice cluster provided in the embodiments of the present application is shown in FIG. 1. Figure 9 As shown in FIG. 1, the embodiments of the present application also provide a drift detection device for a microservice cluster, and the device comprises:

[0112] The collection module 901 is configured to collect link tracking data in a running process of the micro-service cluster in a current detection period, the link tracking data indicating performance of the micro-service cluster in the running process.

[0113] The acquisition module 902 is configured to acquire an incremental subgraph based on the link tracking data, the incremental subgraph indicating a drift node and an entity node associated with the drift node in the micro-service cluster, the drift node being an entity node with a node change degree greater than a first threshold in a running period of the micro-service cluster, the node change degree indicating a change degree of an association relationship of the entity node in the running period of the micro-service cluster.

[0114] The extraction module 903 is configured to perform feature extraction on the incremental subgraph by using a joint embedding algorithm of a multi-layer graph and a hypergraph, to obtain a feature change amount, the feature change amount indicating a node feature change of a node in the incremental subgraph compared with a previous detection period.

[0115] The fusion module 904 is configured to fuse the feature change amount and a first feature to obtain a second feature, the first feature being a node feature of an all-amount node in the previous detection period, and the second feature being a node feature of an all-amount node in the current detection period.

[0116] The detection module 905 is configured to perform drift detection based on the first feature and the second feature, to obtain a detection result, the detection result indicating whether a running period drift occurs in the micro-service cluster.

[0117] In a possible implementation, the acquisition module 902 is configured to acquire the node change degree based on the link tracking data; determine the drift node and the entity node associated with the drift node based on the node change degree and the link tracking data; and acquire the incremental subgraph based on the drift node and the entity node associated with the drift node.

[0118] In another possible implementation, the acquisition module 902 is configured to acquire a first call relationship of an entity node in the micro-service cluster in a previous detection period; acquire a second call relationship of the entity node in the micro-service cluster in a current detection period based on the link tracking data; and determine the node change degree based on a difference between the first call relationship and the second call relationship.

[0119] In another possible implementation, the detection module 905 is configured to acquire a first probability distribution corresponding to the first feature and a second probability distribution corresponding to the second feature based on the first feature and the second feature; determine a distribution difference value of the first probability distribution and the second probability distribution; and acquire the detection result based on the distribution difference value.

[0120] In another possible implementation, the acquisition module 902 is further configured to acquire a query-per-second relative growth rate, the query-per-second relative growth rate indicating a change range of the query per second in the current detection period compared with the last detection period during running of the microservice cluster.

[0121] The detection module 905 is configured to acquire a detection result based on the distribution difference value and the query-per-second relative growth rate.

[0122] In another possible implementation, the detection module 905 is configured to obtain a first detection result in a case where the distribution difference value is greater than the second threshold value and the query-per-second relative growth rate is greater than a third threshold value, the first detection result indicating that the microservice has a running-time drift; and obtain a second detection result in a case where the distribution difference value is less than or equal to the second threshold value, or in a case where the query-per-second relative growth rate is less than or equal to the third threshold value, the second detection result indicating that the microservice cluster does not have a running-time drift.

[0123] In another possible implementation, the apparatus further includes:

[0124] The merging module is configured to start a microservice merging process to merge the microservice cluster in a case where the first detection result is obtained, the first detection result indicating that the microservice cluster has a running-time drift.

[0125] In another possible implementation, the extraction module 903 is configured to acquire multi-dimensional data corresponding to the incremental subgraph, the multi-dimensional data indicating source code, performance in a running process, historical change conditions and business meanings of nodes in the incremental subgraph; construct a multi-layer graph network and a hypergraph based on the multi-dimensional data, the multi-layer graph network including a plurality of graph layers, the graph layers including a plurality of entity nodes corresponding to the incremental subgraph and association relationships between the plurality of entity nodes, the hypergraph including a plurality of hyperedges, the hyperedges indicating entity nodes belonging to a same business process in the incremental subgraph; perform feature extraction on the multi-layer graph network and the hypergraph to obtain node features of the nodes in the incremental subgraph; and determine a difference value between the node features of the nodes in the incremental subgraph and node features of the nodes in the last detection period as a feature change amount.

[0126] In another possible implementation, the extraction module 903 is configured to perform feature extraction on a plurality of graph layers in the multi-layer graph network to obtain graph layer features corresponding to the plurality of graph layers; perform weighted fusion on the graph layer features corresponding to the plurality of graph layers to obtain fused features; perform feature extraction on the hypergraph to obtain hypergraph features; splice the fused features and the hypergraph features to obtain spliced features; and perform linear transformation on the spliced features to obtain the node features of the nodes in the incremental subgraph.

[0127] The features of the embodiments of the drift detection apparatus for the micro-service cluster can be referred to the related descriptions of the embodiments of the drift detection method for the micro-service cluster, which will not be repeated here.

[0128] Figure 10 The structural schematic diagram of an electronic device is provided in the present application. As shown in the figure, the electronic device 1000 provided in the present embodiment comprises at least one processor 1001 and a memory 1002. Optionally, the electronic device 1000 further comprises a communication component 1003. Wherein, the processor 1001, the memory 1002 and the communication component 1003 are connected through a bus. Figure 10

[0129] In the specific implementation process, the at least one processor 1001 executes the computer execution instructions stored in the memory 1002, so that the at least one processor 1001 executes the drift detection method for the micro-service cluster described above.

[0130] The specific implementation process of the processor 1001 can be referred to the method embodiments described above, which has similar implementation principles and technical effects, and will not be repeated here.

[0131] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC) and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor and the like. The steps of the method disclosed in the application can be directly embodied as hardware processor execution, or executed by hardware and software module combination in the processor.

[0132] The memory can contain a random access memory (RAM), and can also include a non-volatile memory (NVM), for example, at least one disk memory.

[0133] ​The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.

[0134] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above-mentioned embodiments of the drift detection method for a micro-service cluster when running.

[0135] In an example embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0136] Embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned embodiments of the drift detection method for a micro-service cluster.

[0137] Embodiments of the present application also provide another computer program product, which includes a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned embodiments of the drift detection method for a micro-service cluster.

[0138] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0139] The above describes in detail the method for detecting drift of the micro-service cluster and the electronic device provided by the present application. The principles and implementation modes of the present application are described by using specific examples, and the above description of the examples is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A drift detection method for microservice clusters, characterized in that, The method includes: During the current detection period, link tracing data is collected during the operation of the microservice cluster, and the link tracing data indicates the performance of the microservice cluster during operation; Based on the link tracing data, an incremental subgraph is obtained. The incremental subgraph indicates the drifting nodes in the microservice cluster and the entity nodes associated with the drifting nodes. The drifting nodes are entity nodes whose node change degree is greater than a first threshold during the operation of the microservice cluster. The node change degree indicates the degree of change in the association relationship of entity nodes during the operation of the microservice cluster. A joint embedding algorithm of multi-layer graph and hypergraph is adopted to extract features from the incremental subgraph and obtain the feature change amount. The feature change amount indicates the change of node features of the nodes in the incremental subgraph compared with the previous detection period. The feature change and the first feature are fused to obtain the second feature. The first feature is the node feature of all nodes in the previous detection period, and the second feature is the node feature of all nodes in the current detection period. Drift detection is performed based on the first feature and the second feature to obtain a detection result, which indicates whether the microservice cluster has experienced runtime drift. The drift detection based on the first feature and the second feature, to obtain the detection result, includes: Based on the first feature and the second feature, obtain the first probability distribution corresponding to the first feature and the second probability distribution corresponding to the second feature; Determine the distribution difference between the first probability distribution and the second probability distribution; Obtain the relative growth rate of queries per second, which indicates the magnitude of change in the number of queries per second in the current detection period compared to the previous detection period during the operation of the microservice cluster; The detection results are obtained based on the distribution difference value and the relative growth rate of queries per second; The method employs a joint embedding algorithm combining multi-layer graphs and hypergraphs to extract features from the incremental subgraph, obtaining the feature change, including: Obtain the multidimensional data corresponding to the incremental subgraph, wherein the multidimensional data indicates the source code of the nodes in the incremental subgraph, their performance during operation, historical changes, and business implications; Based on the multidimensional data, a multi-layer graph network and a hypergraph are constructed. The multi-layer graph network includes multiple layers, each layer including multiple entity nodes corresponding to the incremental subgraph and the relationships between the multiple entity nodes. The hypergraph includes multiple hyperedges, each hyperedge indicating entity nodes in the incremental subgraph that belong to the same business process. Feature extraction is performed on the multilayer graph network and the hypergraph to obtain the node features of the nodes in the incremental subgraph; The difference between the node features of a node in the incremental subgraph and the node features of a node in the incremental subgraph during the previous detection period is determined as the feature change.

2. The method according to claim 1, characterized in that, The step of obtaining an incremental subgraph based on the link tracing data includes: Based on the link tracing data, the node change degree is obtained; Based on the node change degree and the link tracing data, the drifting node and the entity node associated with the drifting node are determined; The incremental subgraph is obtained based on the drift node and the entity node associated with the drift node.

3. The method according to claim 2, characterized in that, The step of obtaining node change degree based on the link tracing data includes: Obtain the first call relationship of the entity nodes in the microservice cluster during the previous detection period; Based on the link tracing data, the second call relationship of the entity nodes in the microservice cluster within the current detection period is obtained; The node change degree is determined based on the difference between the first call relationship and the second call relationship.

4. The method according to claim 1, characterized in that, The process of obtaining the detection result based on the distribution difference value and the relative growth rate of queries per second includes: If the distribution difference value is greater than the second threshold and the relative growth rate of queries per second is greater than the third threshold, a first detection result is obtained, and the first detection result indicates that the microservice has experienced runtime drift. If the distribution difference value is less than or equal to the second threshold, or if the relative growth rate of queries per second is less than or equal to the third threshold, a second detection result is obtained, which indicates that the microservice cluster has not experienced runtime drift.

5. The method according to claim 1, characterized in that, After obtaining the detection result by performing drift detection based on the first feature and the second feature, the method further includes: Upon receiving the first detection result, the microservice merging process is initiated to merge the microservice cluster, whereby the first detection result indicates that the microservice cluster has experienced runtime drift.

6. The method according to claim 1, characterized in that, The step of extracting features from the multilayer graph network and the hypergraph to obtain node features of nodes in the incremental subgraph includes: Feature extraction is performed on multiple layers in the multi-layer graph network to obtain the layer features corresponding to the multiple layers; The layer features corresponding to the multiple layers are weighted and fused to obtain the fused features; Feature extraction is performed on the hypergraph to obtain hypergraph features; The fused feature is concatenated with the hypergraph feature to obtain the concatenated feature; A linear transformation is performed on the splicing features to obtain the node features of the nodes in the incremental subgraph.

7. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the drift detection method for a microservice cluster as described in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Micro-service testing method, device and equipment and computer readable storage medium

    CN115904892A

  • Dependent task unloading method based on reliability perception of topology reconstruction in industrial internet edge computing

    CN120353590A