Full-link monitoring system and method based on micro-service and cloud platform architecture
Through distributed tracking and recording systems and deep learning algorithms, full-link monitoring is carried out in the microservice architecture, which solves the problem that traditional monitoring methods are difficult to track and parse cross-service calls, and realizes efficient monitoring and abnormal detection of the microservice architecture.
Patent Information
- Application Number
- CN202510281257.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In microservice architecture, traditional monitoring methods are difficult to effectively track and parse the complete path of cross-service calls, resulting in inefficiency in positioning inter-service dependencies, analyzing performance bottlenecks and failure sources.
A distributed tracking and recording system is used to record the detailed information of each service call, and a deep learning-based data analysis algorithm is introduced in the full-link monitoring center to perform semantic analysis and timing semantic correlation analysis to extract the deep semantic features and business logical relationships of microservice call information.
It realizes full-link monitoring and abnormal detection of the microservice architecture, improves the accuracy and efficiency of monitoring, can effectively identify and warn of business abnormalities, and ensures the stable operation of the microservice architecture.
Smart Images

Figure CN120200932A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of business monitoring, and more specifically, to an end-to-end monitoring system and method based on a microservices and cloud platform architecture. Background Art
[0002] With the rapid development of Internet technology and cloud computing, the microservices architecture, as a new software design pattern, has gradually become the mainstream. By splitting a large monolithic application into a series of small, autonomous services, each service runs and deploys independently and communicates through well-defined APIs, greatly improving the scalability, maintainability, and flexibility of the system. However, this distributed system architecture also brings new challenges, especially in system monitoring and troubleshooting.
[0003] In traditional monolithic applications, since all components run in the same process, the monitoring of application performance and health status is relatively simple, mainly focusing on the performance metrics monitoring of individual services or components, such as CPU usage, memory occupancy, response time, etc. However, in a microservices-based distributed system, the service call relationships are intricate, and a single user request may involve multiple calls between multiple microservices. Traditional monitoring methods are difficult to effectively track and analyze the complete paths of these cross-service calls, resulting in low efficiency in locating service dependencies, analyzing performance bottlenecks, and identifying the root causes of failures.
[0004] In addition, most existing monitoring tools perform anomaly detection based on rules, relying on preset thresholds or pattern matching. For the complex and ever-changing microservices environment, such static rules are difficult to adapt to dynamic business scenarios, easily causing false alarms or missed alarms. Especially when facing the implicit business logic relationships and timing dependencies in microservices calls, traditional methods are even more difficult to accurately identify and warn of potential business anomalies.
[0005] Therefore, there is a need for an end-to-end monitoring system and method based on a microservices and cloud platform architecture. Summary of the Invention
[0006] To solve the above technical problems, the present application is proposed. Embodiments of the present application provide a full-link monitoring system and method based on a microservices and cloud platform architecture. It uses a distributed tracing recording system to record the detailed information of each service call involved in a business request. At the same time, a data analysis algorithm based on deep learning is introduced in the full-link monitoring center to perform semantic parsing on the detailed information of microservices calls, so as to extract the deep semantic features of each microservices call information. Then, the dependency relationship between services is further constructed to perform time-series semantic association analysis based on the full link for each microservices call information, and mine the business logic relationship and time-series dependency characteristics in microservices calls, thereby realizing full-link monitoring and anomaly detection of the business state. In this way, business anomalies in the microservices architecture can be effectively identified and warned, and the accuracy and efficiency of monitoring can be improved, thus providing strong technical support for the stable operation of the microservices architecture.
[0007] Correspondingly, according to one aspect of the present application, there is provided a full-link monitoring method based on a microservices and cloud platform architecture, which includes:
[0008] Using a distributed tracing recording system to record the detailed information of each service call to obtain a time queue of microservices call detailed information;
[0009] Uploading the time queue of the microservices call detailed information to a full-link monitoring center based on a cloud platform;
[0010] In the full-link monitoring center based on the cloud platform, performing semantic encoding on each microservices call detailed information in the time queue of the microservices call detailed information to obtain a time queue of microservices call detailed information semantic encoding vectors;
[0011] In the full-link monitoring center based on the cloud platform, performing distinguishable strengthening based on essential features on the time queue of the microservices call detailed information semantic encoding vectors to obtain a time queue of microservices call detailed information enhanced semantic encoding vectors;
[0012] In the full-link monitoring center based on the cloud platform, based on the dependency relationship between services, performing service full-link semantic association encoding on the time queue of the microservices call detailed information enhanced semantic encoding vectors to obtain a service full-link time-series semantic encoding matrix;
[0013] In the full-link monitoring center based on the cloud platform, based on the service full-link semantic association encoding to obtain the service full-link time-series semantic encoding matrix, determining the recognition result of the business state.
[0014] According to another aspect of the present application, there is provided a full-link monitoring system based on a microservices and cloud platform architecture, which includes:
[0015] A service call information recording module, which is used to record the detailed information of each service call by using a distributed tracing recording system to obtain a time queue of the detailed information of microservice calls;
[0016] A data transmission module, which is used to upload the time queue of the detailed information of the microservice calls to a full-link monitoring center based on a cloud platform;
[0017] A service call information semantic encoding module, which is used to perform semantic encoding on each piece of detailed information of the microservice calls in the time queue of the detailed information of the microservice calls respectively in the full-link monitoring center based on a cloud platform to obtain a time queue of semantic encoding vectors of the detailed information of the microservice calls;
[0018] A semantic discrimination enhancement module, which is used to perform distinguishable enhancement based on essential features on the time queue of the semantic encoding vectors of the detailed information of the microservice calls in the full-link monitoring center based on a cloud platform to obtain a time queue of enhanced semantic encoding vectors of the detailed information of the microservice calls;
[0019] A full-link semantic association encoding module, which is used to perform service full-link semantic association encoding on the time queue of the enhanced semantic encoding vectors of the detailed information of the microservice calls based on the dependency relationship between services in the full-link monitoring center based on a cloud platform to obtain a service full-link time-series semantic encoding matrix;
[0020] A business status recognition module, which is used to determine the recognition result of the business status based on the service full-link semantic association encoding to obtain the service full-link time-series semantic encoding matrix in the full-link monitoring center based on a cloud platform.
[0021] Compared with the prior art, the full-link monitoring system and method based on the microservice and cloud platform architecture provided by the present application first obtain a user data set and a commodity data set, and use a distributed tracing recording system to record the detailed information of each service call involved in a business request. At the same time, a data analysis algorithm based on deep learning is introduced into the full-link monitoring center to perform semantic parsing on the detailed information of the microservice calls to extract the deep semantic features of each piece of microservice call information. Then, the dependency relationship between services is further constructed to perform time-series semantic association analysis based on the full link on each piece of microservice call information, and the business logic relationship and time-series dependency characteristics in the microservice calls are mined, so as to realize the full-link monitoring and anomaly detection of the business status. In this way, business anomalies in the microservice architecture can be effectively identified and warned, the accuracy and efficiency of monitoring can be improved, and strong technical support can be provided for the stable operation of the microservice architecture. Description of the Drawings
[0022] The above and other objects, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail with reference to the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 It is a flowchart of a full-link monitoring method based on a microservices and cloud platform architecture according to an embodiment of the present application.
[0024] Figure 2 It is a schematic diagram of data flow of a full-link monitoring method based on a microservices and cloud platform architecture according to an embodiment of the present application.
[0025] Figure 3 It is a flowchart of step S140 in a full-link monitoring method based on a microservices and cloud platform architecture according to an embodiment of the present application.
[0026] Figure 4 It is a flowchart of step S150 in a full-link monitoring method based on a microservices and cloud platform architecture according to an embodiment of the present application.
[0027] Figure 5 It is a block diagram of a full-link monitoring system based on a microservices and cloud platform architecture according to an embodiment of the present application. Detailed Embodiments
[0028] Next, exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0029] Figure 1 It is a flowchart of a full-link monitoring method based on a microservices and cloud platform architecture according to an embodiment of the present application. Figure 2 It is a schematic diagram of data flow of a full-link monitoring method based on a microservices and cloud platform architecture according to an embodiment of the present application. As Figure 1 and Figure 2As shown, the full-link monitoring method based on the microservice and cloud platform architecture according to the embodiments of the present application includes the steps of: S110, using a distributed tracing recording system to record the detailed information of each service call to obtain a time queue of microservice call detailed information; S120, uploading the time queue of the microservice call detailed information to the full-link monitoring center based on the cloud platform; S130, in the full-link monitoring center based on the cloud platform, semantically encoding each microservice call detailed information in the time queue of the microservice call detailed information to obtain a time queue of microservice call detailed information semantic encoding vectors; S140, in the full-link monitoring center based on the cloud platform, performing distinguishable enhancement based on the essential features on the time queue of the microservice call detailed information semantic encoding vectors to obtain a time queue of microservice call detailed information enhanced semantic encoding vectors; S150, in the full-link monitoring center based on the cloud platform, based on the dependency relationship between services, performing service full-link semantic association encoding on the time queue of the microservice call detailed information enhanced semantic encoding vectors to obtain a service full-link time-series semantic encoding matrix; S160, in the full-link monitoring center based on the cloud platform, determining the recognition result of the business status based on the service full-link semantic association encoding to obtain the service full-link time-series semantic encoding matrix.
[0030] In the above full-link monitoring method based on the microservice and cloud platform architecture, in step S110, a distributed tracing recording system is used to record the detailed information of each service call to obtain a time queue of microservice call detailed information. It should be understood that the Distributed Tracing System is a tool for monitoring and diagnosing distributed systems, which can track the flow path of a transaction or request among multiple services. In a distributed environment, a single operation of a user may trigger multiple internal calls across multiple different services. Through the distributed tracing system, it is possible to accurately track all the service (i.e., microservice) nodes and their order that each request experiences from start to end, forming a time queue of microservice call detailed information, thereby providing a detailed data basis for subsequent full-link data monitoring and analysis.
[0031] Specifically, the distributed tracing system infrastructure usually consists of multiple key components. First is the tracing data collector, which is deployed in each microservice instance and is responsible for collecting call information within the service. These tracing data collectors need to be deeply integrated with the microservice code so that data can be captured at various key nodes of the service call. For example, in a microservice system built with Spring Cloud, by introducing a dedicated distributed tracing SDK such as Spring Cloud Sleuth, data collection logic can be implanted at the entry and exit of microservice methods, as well as at the initiation and reception of remote calls.
[0032] The distributed tracing record system relies on a globally unique tracing identifier (Trace ID) to uniquely identify the entire call chain of a transaction or request, which runs through the entire request chain, from the client initiating the request to the final response being returned to the client. Whenever a new business request enters the system, a unique Trace ID is generated. This ID is passed between all relevant microservices, enabling each participating service to identify that the request belongs to the same transaction or process. At the same time, for each internal service call, a locally unique operation identifier (SpanID) is also assigned to distinguish different operation steps under the same trace.
[0033] When a microservice receives a request with a specific Trace ID and Span ID, it creates a new Span to represent this call and passes its own Span ID as the parent Span ID to the next called service. The purpose of this is to establish a complete call tree structure, where each node represents a specific microservice call, and the edges reflect the call order and hierarchical relationship between them. In this way, even in a highly distributed environment, the specific path of each request can be clearly traced.
[0034] Next, to record the detailed information of each service call, a series of data points related to this call need to be collected. These data include but are not limited to: service name, call timestamp, execution duration, input parameters, output results, status code, network transmission information, relevant details of database access, and any other content that can reflect the service behavior characteristics. All this information will be packaged into an event log entry and attached to the corresponding Span.
[0035] Specifically, the data points include: the start time and end time of the call: the start time can determine the initiation moment of the service call, while the end time marks the completion of the call. The difference between the two is the duration of this service call. For example, in a microservice call for processing user order queries, recording the time interval from receiving the query request to returning the query result is crucial for analyzing the response performance of this service.
[0036] The source microservice identifier and target microservice identifier of the call: Through unique identification information (such as microservice name, instance ID, or custom encoding), each microservice can be clearly distinguished, and the direction of the call can be determined. In an e-commerce system, for example, the order service may call the inventory service to check the product inventory. Then the order service is the source microservice, and the inventory service is the target microservice. Accurately recording these identifiers helps to build the service call topology map of the entire system.
[0037] The unique identifier of the call: Generate a globally unique identifier for each service call. This identifier runs through the entire service call link and can associate multiple microservice calls involved in a complete user request. For example, when a user initiates a product purchase request at the front end, a series of microservice calls starting from the user authentication service, to the order creation service, payment service, etc. can be connected through this unique identifier to form a complete tracking link.
[0038] The hierarchical information of the call: In a complex distributed system, there may be multiple levels of nested service calls. For example, microservice A calls microservice B, and microservice B further calls microservice C. Recording the hierarchical information of the call helps to intuitively understand the depth and complexity of the service call. A digital or specific encoding method can be used to represent the levels. For example, the level of the root node is 0, and the level of the microservice directly called by it is 1, and so on.
[0039] The parameter information of the request: Record in detail the content of the request parameters passed during the service call. For different types of microservices, the forms and contents of the request parameters vary. For example, in a user registration microservice, the request parameters may include user information such as username, password, and email; while in a file upload microservice, the request parameters may include file metadata (file name, file size, file type, etc.) and the binary data of the file (which can be summarized or partially sampled for recording). Recording the request parameters helps to restore the call scene during fault troubleshooting and analyze whether the service call fails due to abnormal request data.
[0040] Status Code and Data Digest of the Response: After the service call is completed, the status code of the response can intuitively reflect whether the call is successful. For example, the common HTTP status code 200 indicates success, and 500 indicates an internal server error, etc. At the same time, to avoid recording a large amount of response data (which may contain sensitive information or be too large in size), the response data can be subjected to digest calculation, such as using a hash algorithm (such as MD5, SHA-1, etc.) to generate a fixed-length digest information. In this way, during subsequent analysis, if detailed response data needs to be viewed, the original response data can be obtained from the storage system according to the digest information for further analysis.
[0041] Over time, as more and more services participate in the same tracing process, these log entries scattered in various places gradually form a continuous timeline - namely, the "time queue of microservice call details". This timeline is not just a simple collection of logs, but rather a data sequence with internal logical connections, which can reveal how each link in the business process occurs and develops sequentially. For the convenience of subsequent processing, these original logs are usually preprocessed, such as removing redundant information, formatting fields, adding metadata tags, etc., so as to construct a dataset with a more standardized structure and easier to parse.
[0042] In the above full-link monitoring method based on the microservice and cloud platform architecture, in step S120, the time queue of the microservice call details is uploaded to the full-link monitoring center based on the cloud platform. That is, in order to achieve the global view management and efficient processing of microservice call information, the time queue of the microservice call details is further uploaded to the full-link monitoring center based on the cloud platform. The full-link monitoring center based on the cloud platform is equipped with powerful computing resources and storage capabilities, supports the real-time processing and storage of large-scale data, and deploys data processing and analysis algorithms based on deep learning, which can perform real-time analysis and processing on the uploaded microservice call details, helping to more clearly understand the interaction situation between each microservice and the overall business process.
[0043] Specifically, in this process, choosing a suitable cloud platform is crucial for the overall data processing performance. Modern cloud service platforms such as Alibaba Cloud, AWS (Amazon Web Services), Google Cloud Platform, etc. all offer rich function and service options. For example, Alibaba Cloud has powerful big data processing capabilities and is suitable for complex data analysis tasks; while AWS is known for its extensive geographical coverage and highly customizable solutions; Google Cloud Platform emphasizes its advantages in machine learning and artificial intelligence. When choosing a specific platform, factors such as business characteristics, technical requirements, budget constraints, etc. need to be comprehensively considered. At the same time, the security, compliance, and compatibility with other existing IT infrastructures of each platform also need to be evaluated.
[0044] After determining the target cloud platform, a reasonable data transmission plan needs to be designed. Considering factors such as network bandwidth, latency, packet loss rate, etc. that may affect the transmission efficiency, optimized transmission protocols such as HTTP / 2 or gRPC are usually adopted to ensure that data arrives at the destination quickly and stably. Especially for application scenarios with high real-time requirements, two-way communication protocols such as WebSocket also need to be introduced for instant feedback and interaction. In addition, in order to reduce unnecessary load, the original log can be compressed and encoded before uploading, such as using the Gzip algorithm, which not only saves bandwidth but also speeds up the transmission.
[0045] At the same time, security is the most important aspect that cannot be ignored in any operation involving sensitive data. Therefore, in the process of uploading the time queue of the microservice call details to the cloud platform, a series of strict protection measures must be taken. First of all, all transmitted data should be encrypted to ensure that the content cannot be interpreted even if intercepted during network transmission. This can be achieved through an end-to-end encrypted connection using the SSL / TLS protocol, or using a higher-level encryption standard such as AES-256. Secondly, access control policies are also very important, and only authorized users and services can read or write specific data sets. Finally, regularly reviewing and updating security settings and promptly patching known vulnerabilities are also an indispensable part of maintaining long-term stable operation.
[0046] After successfully uploading to the cloud platform, it is necessary to further manage and process the large amount of incoming data properly and efficiently. On the one hand, a unique identifier needs to be assigned to each batch of uploaded data for subsequent query and management; on the other hand, the newly arrived data should be classified, filtered, and preliminarily cleaned according to preset rules to eliminate invalid or duplicate information and retain valuable records. In addition, considering the possible large-scale data analysis in the future, the data warehouse structure can be planned in advance, a suitable database type can be selected (such as the relational database MySQL, the NoSQL database MongoDB), and indexes can be set reasonably to improve the retrieval efficiency.
[0047] As more and more historical data accumulates, the scale and complexity of the data warehouse will also increase. To address this challenge, a message queue system such as Apache Kafka can be used as a buffer to smooth the input and output rate differences; an Elasticsearch-based full-text search engine can be built to facilitate the quick location of specific events; or even a streaming processing framework such as SparkStreaming or Flink can be combined to achieve real-time data analysis and anomaly detection.
[0048] In the above full - link monitoring method based on the microservice and cloud - platform architecture, in step S130, in the full - link monitoring center based on the cloud platform, semantic encoding is respectively performed on each microservice call detail in the time queue of microservice call details to obtain a time queue of microservice call detail semantic encoding vectors. In a specific example of the present application, a language model based on the Transformer architecture is used to respectively perform semantic encoding on each microservice call detail in the time queue of microservice call details to obtain the time queue of microservice call detail semantic encoding vectors. It should be understood that considering that the original microservice call details usually contain a large amount of unstructured or semi - structured data (such as API requests and responses, log records, etc.), in order to effectively use deep - learning algorithms for data analysis on them, it is necessary to further convert the microservice call details into a structured data form suitable for the input of deep - learning models. Based on this, the present application uses a language model based on the Transformer architecture to respectively perform semantic encoding on each microservice call detail in the time queue of microservice call details to map it to a high - dimensional semantic space, forming a time queue of microservice call detail semantic encoding vectors. Those of ordinary skill in the art should know that the language model based on the Transformer architecture is a natural language processing (NLP) method. Based on the self - attention mechanism, by learning the long - distance dependence relationships between various lexical units in the text sequence data, it can effectively capture and understand the deep semantic information of the text. Based on this, by using the language model based on the Transformer architecture to perform semantic encoding on each microservice call detail, the operation behaviors and execution situations of each microservice call can be fully understood, and the complex information of each microservice call can be converted into a vector representation with rich semantics, that is, a time queue of microservice call detail semantic encoding vectors, thereby providing a solid data basis for subsequent business - state recognition and anomaly detection.
[0049] In the above full - link monitoring method based on the microservice and cloud - platform architecture, in step S140, in the full - link monitoring center based on the cloud platform, a distinguishable enhancement based on essential features is performed on the time queue of the semantic encoding vectors of the microservice call detailed information to obtain a time queue of enhanced semantic encoding vectors of the microservice call detailed information. It should be understood that considering the complexity and diversity of microservice calls, there may be similarities between different microservice call information, such as similar API interface designs, repetitive service logics, shared basic libraries or frameworks, etc. This may lead to confusion in subsequent analysis and anomaly detection processes. Therefore, in order to enhance the ability to identify subtle differences between different service calls, the present application further performs feature - distinctiveness enhancement processing on the time queue of the semantic encoding vectors of the microservice call detailed information to improve the distinguishability of different microservice call information.
[0050] Figure 3 The flowchart of step S140 in the full - link monitoring method based on the microservice and cloud - platform architecture according to an embodiment of the present application is as follows. As Figure 3 shown, step S140 includes: S141, extracting the essential features of the time queue of the semantic encoding vectors of the microservice call detailed information to obtain a microservice call detailed information semantic - field essential feature vector; S142, based on the microservice call detailed information semantic - field essential feature vector, performing semantic - distinction enhancement processing on each semantic encoding vector of the microservice call detailed information in the time queue of the semantic encoding vectors of the microservice call detailed information to obtain a time queue of enhanced semantic encoding vectors of the microservice call detailed information.
[0051] Specifically, step S141 includes: First, performing field mapping on each semantic encoding vector of the microservice call detailed information in the time queue of the semantic encoding vectors of the microservice call detailed information to obtain a time queue of microservice call detailed information semantic - field feature vectors, which is expressed by the formula:
[0052] h i =W1I i W2
[0053] where, I i represents the i - th semantic encoding vector of the microservice call detailed information in the time queue of the semantic encoding vectors of the microservice call detailed information, W1 and W2 respectively represent the first linear mapping matrix and the second linear mapping matrix, and h i represents the microservice call detailed information semantic - field feature vector.
[0054] That is, first, perform a field mapping operation on the semantic encoding vectors of the detailed information of each microservice call to map them to a unified field space, thereby unifying the feature dimensions, so as to better reflect the internal semantic associations between the detailed information of each microservice call in the unified field space.
[0055] Next, calculate the field depth factor of each semantic field feature vector of the detailed information of the microservice call in the time queue of the semantic field feature vectors of the detailed information of the microservice call to obtain the time queue of the field depth factor of the semantic field of the detailed information of the microservice call, which is expressed by the formula:
[0056]
[0057] where f(·) is the field depth factor measurement function, ||·|| represents the L2 norm of the vector, log2(·) represents the logarithm function with base 2, and e i represents the h i corresponding field depth factor of the semantic field of the detailed information of the microservice call.
[0058] Then, based on the time queue of the field depth factor of the semantic field of the detailed information of the microservice call, calculate the field essential feature of the time queue of the semantic field feature vectors of the detailed information of the microservice call to obtain the semantic field essential feature vector of the detailed information of the microservice call, which is expressed by the formula:
[0059]
[0060] where exp(·) represents the natural exponential function with base e, a i represents the h i corresponding normalized field depth factor of the semantic field of the detailed information of the microservice call, and v c represents the semantic field essential feature vector of the detailed information of the microservice call. That is, after normalizing the time queue of the field depth factor of the semantic field of the detailed information of the microservice call, use it as the time queue of the weight coefficients to perform weighted aggregation on the time queue of the semantic field feature vectors of the detailed information of the microservice call to obtain the semantic field essential feature vector of the detailed information of the microservice call.
[0061] Here, by calculating the field depth factor of each microservice call detailed information semantic field feature vector obtained after mapping, the information density and feature importance in the field space are measured, a time queue of microservice call detailed information semantic field depth factors is obtained, and based on this, using the field effect normalization mechanism, it is converted into a weight coefficient, and the position-weighted aggregation of each microservice call detailed information semantic field feature vector is performed, so as to extract a semantic feature representation that comprehensively reflects the overall basic attributes of microservice call detailed information, that is, the microservice call detailed information semantic field essential feature vector.
[0062] Specifically, the step S142 includes: inputting the microservice call detailed information semantic field essential feature vector and the microservice call detailed information semantic coding vector into a semantic difference feature extraction network based on the sigmoid activation function to obtain a microservice call detailed information semantic difference feature vector; calculating the element-wise dot product between the microservice call detailed information semantic difference feature vector and the microservice call detailed information semantic coding vector to obtain the microservice call detailed information enhanced semantic coding vector, which is expressed by the formula:
[0063] O i =I i ⊙[Sigmoid(M m v c +H m I i +b m )]
[0064] where, M m and H m respectively represent the essential feature weight parameter matrix and the microservice call feature weight parameter matrix of the semantic difference feature extraction network, b m represents the bias vector, Sigmoid(·) represents the sigmoid activation function, ⊙ represents the element-wise dot product, and O i represents the microservice call detailed information enhanced semantic coding vector corresponding to the h i .
[0065] That is, taking the microservice call detailed information semantic field essential feature vector as a benchmark, a neural network model is further introduced to learn the difference information in each microservice call detailed information semantic coding vector relative to the microservice call detailed information semantic field essential feature vector, and through element-wise distinguishable enhancement processing, the unique information in each microservice call detailed information is specifically amplified, generating a time queue of microservice call detailed information enhanced semantic coding vectors, so as to ensure that in the subsequent business state analysis process, the subtle differences between different microservice call detailed information can be effectively distinguished.
[0066] In the above full-link monitoring method based on the microservice and cloud platform architecture, in step S150, in the full-link monitoring center based on the cloud platform, based on the dependency relationship between services, perform full-link semantic association encoding on the time queue of the enhanced semantic encoding vectors of the microservice call details to obtain a full-link time-series semantic encoding matrix for services. Among them, Figure 4 FIG. Figure 4 is a flowchart of step S150 in the full-link monitoring method based on the microservice and cloud platform architecture according to an embodiment of the present application. As Figure 4 shown, step S150 includes: S151, constructing a dependency relationship graph between services; S152, inputting the dependency relationship graph between services into a dependency relationship topological feature extractor based on a convolutional neural network model to obtain a dependency relationship feature matrix between services; S153, inputting the time queue of the enhanced semantic encoding vectors of the microservice call details and the dependency relationship feature matrix between services into a full-link semantic encoder of the service based on a graph neural network model to obtain the full-link time-series semantic encoding matrix for services.
[0067] Specifically, in step S151, construct a dependency relationship graph between services. It should be understood that in a microservice architecture, there are complex dependency relationships between various microservices. For example, service A may call service B through an API, and service B may depend on service C, thus forming a service dependency chain. Therefore, in order to comprehensively understand the interaction and dependency between microservices, the present application further constructs a dependency relationship graph between services to more clearly show the call relationship and dependency path between each microservice, so as to facilitate understanding the overall structure and operation mechanism of the business process.
[0068] In a specific implementation, taking the time queue of the microservice call details as input, by parsing the trace record of each call, key information such as the service name involved, call direction (upstream / downstream), call timestamp, etc. are extracted. Then, a directed acyclic graph (DAG) is established based on this information, where each vertex represents an independent microservice, and the edge represents the direct call relationship between microservices.
[0069] For example, in a microservice architecture, assuming there are N different microservices, an N×N square matrix A can be created, where each element aij represents the dependency strength or call frequency from the i-th microservice to the j-th microservice. If there is no direct dependency between microservices, the corresponding square matrix element can be set to 0.
[0070] In addition, considering that microservice calls have obvious temporal characteristics, that is, some services are always called prior to other services, a time dimension can be introduced into the matrix to form a three-dimensional tensor T. At this time, T[i][j][t] represents the dependency situation from service i to service j at time t. Doing so not only preserves the original time information but also enables the comparison of the changing trends of the dependencies between services in different time periods, which helps to discover periodic patterns or abnormal fluctuations. In this way, the dependency relationship graph between the services completely records the direct call relationships between all microservices.
[0071] Specifically, in step S152, the dependency relationship graph between the services is input into a dependency relationship topology feature extractor based on a convolutional neural network model to obtain a dependency relationship feature matrix between the services. That is, in order to more fully understand the broader dependency relationships and interaction patterns between services, the present application uses a convolutional neural network (CNN) model to extract the topological features of the dependency relationship graph between the services. It should be understood that the convolutional neural network captures local spatial correlations by moving a sliding window (or called a kernel) over the dependency relationship graph between the services, thereby being able to effectively detect the local connection patterns between different services. In addition, by stacking multiple convolutional layers, the convolutional neural network (CNN) model can understand the dependency relationships between services at different levels of abstraction, capturing from direct point-to-point connections to broader subnet or community structures, thus providing richer context information for subsequent business process analysis and anomaly detection.
[0072] Specifically, in step S153, the time queue of the enhanced semantic encoding vectors of the microservice call details and the inter-service dependency relationship feature matrix are input into the service full-link semantic encoder based on the graph neural network model to obtain the service full-link time-series semantic encoding matrix. It should be understood that the graph neural network (GNN) model is particularly suitable for processing graph-structured data, which can effectively capture complex non-Euclidean relationships between nodes, thereby modeling the full-link behavior of microservice calls. In the technical solution of this application, in order to comprehensively capture the full-link behavior of microservice calls, the time queue of the enhanced semantic encoding vectors of the microservice call details is regarded as each node in the graph structure, and the inter-service dependency relationship feature matrix is used as the edge of the graph structure. Using the message passing mechanism of the graph neural network model, self-update of node information and aggregation of neighbor node information are performed on the graph structure, so as to effectively capture the service call patterns and behaviors that change over time, enabling each microservice call detail node to comprehensively reflect its own characteristics and the characteristics of its connected neighbor nodes, forming context information that integrates inter-service dependency relationships, that is, the service full-link time-series semantic encoding matrix. In this way, the associated dependency structure and time-series relationship between microservice calls can be more accurately identified and understood, providing a more comprehensive analysis basis for subsequent business status analysis.
[0073] In the above full-link monitoring method based on the microservice and cloud platform architecture, in step S160, in the full-link monitoring center based on the cloud platform, based on the service full-link semantic association encoding to obtain the service full-link time-series semantic encoding matrix, the recognition result of the business status is determined. In a specific example of this application, the service full-link time-series semantic encoding matrix is input into the business status recognizer based on the classifier to obtain the recognition result, and the recognition result is used to indicate whether there is an abnormality in the business status. Specifically, the classifier learns through training to identify and distinguish the differences in microservice call behaviors between normal business status and abnormal business status, and establish a decision boundary. In actual applications, the classifier combines the decision rules learned during the training process, extracts features and performs pattern recognition on the input service full-link time-series semantic encoding matrix to determine whether there is an abnormality in the current business status. If an abnormality is detected, the classifier will output a corresponding abnormal signal, thereby triggering an early warning mechanism to notify the operation and maintenance personnel to check and handle in a timely manner. In this way, real-time monitoring and abnormal detection of the business status can be achieved, ensuring the stable operation of the business and rapid response to potential problems.
[0074] Here, each microservice call detail semantic encoding vector in the time queue of the microservice call detail semantic encoding vectors represents the semantic encoding features of each microservice call detail in the time queue of the microservice call details, and the inter-service dependency relationship feature matrix includes the inter-server dependency relationship topology features. After performing distinguishable reinforcement modeling on the time queue of the microservice call detail semantic encoding vectors using the feature distinguishable reinforcement module based on sequence conditions, during the process of inputting the time queue of the microservice call detail enhanced semantic encoding vectors and the inter-service dependency relationship feature matrix into the service full-link semantic encoder based on the graph neural network model, the nodes have undergone feature reinforcement but the edges have not, which will cause the semantic context association encoding weights between the edges and the nodes to be unbalanced during the graph convolution encoding process, resulting in the service full-link time-series semantic encoding matrix having an interactive fusion distribution space structure difference, thereby affecting the accuracy of the recognition result obtained by the business status recognizer based on the classifier.
[0075] Preferably, inputting the service full-link time-series semantic encoding matrix into the business status recognizer based on the classifier to obtain the recognition result includes:
[0076] Calculating the modulus cumulant of all eigenvalues of the service full-link time-series semantic encoding matrix to obtain the first service full-link time-series semantic encoding main space topology parameter, and calculating its quadratic norm basis quantity to obtain the second service full-link time-series semantic encoding main space topology parameter, that is:
[0077]
[0078] where, f i represents each eigenvalue of the service full-link time-series semantic encoding matrix, n represents the total matrix dimension of the service full-link time-series semantic encoding matrix, α represents the first service full-link time-series semantic encoding main space topology parameter, and β represents the second service full-link time-series semantic encoding main space topology parameter;
[0079] For each eigenvalue of the service full-link time-series semantic encoding matrix, calculating the first service full-link time-series semantic encoding cross-domain association factor obtained by subtracting the product of the eigenvalue and the total matrix dimension of the service full-link time-series semantic encoding matrix from the first service full-link time-series semantic encoding main space topology parameter, expressed as:
[0080] x i = α - f i ×n
[0081] where, x i represents the first service full-link time-series semantic encoding cross-domain association factor;
[0082] Calculate the product of the square root of the total matrix dimension and the eigenvalue, and subtract the second service full-link time-series semantic coding main space topology parameter to obtain the second service full-link time-series semantic coding cross-domain correlation factor, expressed as:
[0083]
[0084] where y i represents the second service full-link time-series semantic coding cross-domain correlation factor;
[0085] Mix and weight aggregate the value of the power function with the first service full-link time-series semantic coding cross-domain correlation factor as the base of the natural exponent and the reciprocal of the second service full-link time-series semantic coding cross-domain correlation factor to obtain the optimized eigenvalue corresponding to each eigenvalue, expressed as:
[0086]
[0087] where e represents the natural exponent, ε and θ represent weight hyperparameters, and f' i represents the optimized eigenvalue corresponding to each eigenvalue;
[0088] Input the optimized service full-link time-series semantic coding matrix composed of the optimized eigenvalues into the classifier-based business status recognizer to obtain the recognition result.
[0089] Based on this, aiming at the topological representation fault that may exist in the attribute set of the service full-link time-series semantic coding matrix in the multi-dimensional vector space, which leads to optimization divergence in image semantic segmentation based on the implicit derivation mechanism, a cross-dimensional feature correlation network is constructed (this network is based on the relative scale relationship between the overall representation dimension of the service full-link time-series semantic coding matrix and its spatial topology expression) to strengthen the neighborhood connectivity characteristics of the service full-link time-series semantic coding matrix. At the same time, the spatial prediction mechanism of discrete feature parameters is used to capture the dimensional fuzzy representation of the target feature parameters, thereby enhancing the spatial generalization adaptation ability of the attribute set of the service full-link time-series semantic coding matrix and improving the convergence stability of the semantic segmentation optimization process. In this way, the accuracy of the recognition result obtained by inputting the service full-link time-series semantic coding matrix into the classifier-based business status recognizer is improved.
[0090] In summary, the full - link monitoring method based on the microservice and cloud - platform architecture according to the embodiments of the present application is elucidated. It uses a distributed tracing recording system to record the detailed information of each service call involved in a business request. At the same time, a data - analysis algorithm based on deep learning is introduced in the full - link monitoring center to perform semantic parsing on the detailed information of microservice calls, so as to extract the deep semantic features of each microservice call information. Then, the dependency relationship between services is further constructed to perform full - link time - series semantic correlation analysis on each microservice call information, mining out the business - logic relationship and time - series dependency characteristics in microservice calls, thereby realizing full - link monitoring and anomaly detection of the business state. In this way, business anomalies in the microservice architecture can be effectively identified and warned, improving the accuracy and efficiency of monitoring, and thus providing strong technical support for the stable operation of the microservice architecture.
[0091] Furthermore, the present application also provides a full - link monitoring system based on the microservice and cloud - platform architecture.
[0092] Figure 5 FIG. is a block diagram of the full - link monitoring system based on the microservice and cloud - platform architecture according to the embodiments of the present application. As Figure 5 shown, the full - link monitoring system 100 based on the microservice and cloud - platform architecture according to the embodiments of the present application includes: a service - call information recording module 110, which is used to record the detailed information of each service call by using a distributed tracing recording system to obtain a time queue of microservice - call detailed information; a data - transmission module 120, which is used to upload the time queue of the microservice - call detailed information to the full - link monitoring center based on the cloud platform; a service - call information semantic - encoding module 130, which is used to perform semantic encoding on each microservice - call detailed information in the time queue of the microservice - call detailed information in the full - link monitoring center based on the cloud platform to obtain a time queue of microservice - call detailed - information semantic - encoding vectors; a semantic - discrimination strengthening module 140, which is used to perform distinguishable strengthening based on essential features on the time queue of the microservice - call detailed - information semantic - encoding vectors in the full - link monitoring center based on the cloud platform to obtain a time queue of microservice - call detailed - information enhanced - semantic - encoding vectors; a full - link semantic - correlation encoding module 150, which is used to perform service full - link semantic - correlation encoding on the time queue of the microservice - call detailed - information enhanced - semantic - encoding vectors based on the dependency relationship between services in the full - link monitoring center based on the cloud platform to obtain a service full - link time - series semantic - encoding matrix; a business - state recognition module 160, which is used to determine the recognition result of the business state based on the service full - link semantic - correlation encoding to obtain the service full - link time - series semantic - encoding matrix in the full - link monitoring center based on the cloud platform.
[0093] Here, those skilled in the art can understand that the specific operations of each module in the above-mentioned full-link monitoring system based on the microservice and cloud platform architecture have been introduced in detail in the description of the full-link monitoring method based on the microservice and cloud platform architecture above with reference to Figures 1 to 4 and thus, the repeated description thereof will be omitted.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A full-link monitoring method based on microservices and cloud platform architecture, characterized in that: include: Use the distributed tracking system to record the detailed information of each service call to obtain the time queue of microservice call details; Upload the time queue of the microservice call details to the full-link monitoring center based on the cloud platform; In a full-link monitoring center based on a cloud platform, semantic encoding is performed on each microservice call detailed information in the time queue of the microservice call detailed information to obtain a time queue of the microservice call detailed information semantic encoding vector; In a full-link monitoring center based on a cloud platform, the time queue of the microservice call detailed information semantic coding vector is distinguishably enhanced based on essential features to obtain the time queue of the microservice call detailed information enhanced semantic coding vector; In the full-link monitoring center based on the cloud platform, based on the dependency relationship between services, the time queue of the semantic coding vector of the microservice call detailed information reinforcement is coded for the full-link semantic association of the service to obtain the full-link time sequence semantic coding matrix of the service; In the full-link monitoring center based on the cloud platform, the service full-link semantic association coding is based on the service full-link temporal semantic coding matrix to obtain the service full-link, and determine the recognition result of the business status.
2. The full-link monitoring method based on microservice and cloud platform architecture according to claim 1 is characterized in that: Semantically encoding each microservice call detailed information in the time queue of the microservice call detailed information to obtain a time queue of the microservice call detailed information semantic encoding vector, including: A language model based on the Transformer architecture is used to semantically encode each microservice call detailed information in the time queue of the microservice call detailed information to obtain a time queue of the microservice call detailed information semantic encoding vector.
3. The full-link monitoring method based on microservice and cloud platform architecture according to claim 2 is characterized in that: Performing distinguishable enhancement based on essential features on the time queue of the microservice call detailed information semantic coding vector to obtain the time queue of the microservice call detailed information enhanced semantic coding vector, including: Extracting the essential features of the time queue of the microservice call detailed information semantic encoding vector to obtain the microservice call detailed information semantic field essential feature vector; Based on the microservice call detailed information semantic field essential feature vector, each microservice call detailed information semantic coding vector in the time queue of the microservice call detailed information semantic coding vector is subjected to semantic differentiation and enhancement processing to obtain the time queue of the microservice call detailed information enhanced semantic coding vector.
4. The full-link monitoring method based on microservice and cloud platform architecture according to claim 3 is characterized in that: Extracting the essential features of the time queue of the microservice call detailed information semantic encoding vector to obtain the microservice call detailed information semantic field essential feature vector, including: Performing field mapping on each microservice call detailed information semantic coding vector in the time queue of the microservice call detailed information semantic coding vector to obtain a time queue of microservice call detailed information semantic field feature vectors; Calculating the field depth factor of each microservice call detailed information semantic field feature vector in the time queue of the microservice call detailed information semantic field feature vector to obtain the time queue of the microservice call detailed information semantic field depth factor; Based on the time queue of the semantic field depth factor of the microservice call detailed information, the field essential characteristics of the time queue of the semantic field feature vector of the microservice call detailed information are calculated to obtain the semantic field essential characteristic vector of the microservice call detailed information.
5. The full-link monitoring method based on microservice and cloud platform architecture according to claim 4 is characterized in that: Based on the time queue of the semantic field depth factor of the microservice call detailed information, calculating the field essential characteristics of the time queue of the semantic field feature vector of the microservice call detailed information to obtain the semantic field essential characteristic vector of the microservice call detailed information, including: After normalizing the time queue of the microservice call detailed information semantic field depth factor, the time queue using it as a weight coefficient is weightedly aggregated to obtain the microservice call detailed information semantic field essential feature vector.
6. The full-link monitoring method based on microservice and cloud platform architecture according to claim 5 is characterized in that: Based on the microservice call detailed information semantic field essential feature vector, each microservice call detailed information semantic coding vector in the time queue of the microservice call detailed information semantic coding vector is subjected to semantic differentiation and enhancement processing to obtain the time queue of the microservice call detailed information enhanced semantic coding vector, including: Inputting the microservice call detailed information semantic field essential feature vector and the microservice call detailed information semantic encoding vector into a semantic difference feature extraction network based on a sigmoid activation function to obtain a microservice call detailed information semantic difference feature vector; The microservice call detailed information enhanced semantic coding vector is obtained by calculating the position point multiplication between the microservice call detailed information semantic difference feature vector and the microservice call detailed information semantic coding vector.
7. The full-link monitoring method based on microservice and cloud platform architecture according to claim 6 is characterized in that: Based on the dependencies between services, the time queue of the microservice call detailed information enhanced semantic coding vector is coded for service full-link semantic association to obtain a service full-link temporal semantic coding matrix, including: Build a dependency graph between services; Inputting the dependency graph between services into a dependency topology feature extractor based on a convolutional neural network model to obtain a dependency feature matrix between services; The time queue of the microservice call detailed information enhanced semantic coding vector and the inter-service dependency feature matrix are input into the service full-link semantic encoder based on the graph neural network model to obtain the service full-link temporal semantic coding matrix.
8. The full-link monitoring method based on microservice and cloud platform architecture according to claim 7 is characterized in that: Based on the service full-link semantic association coding to obtain a service full-link temporal semantic coding matrix, determining a recognition result of the business state includes: The service full-link temporal semantic coding matrix is input into a classifier-based business status identifier to obtain the identification result, and the identification result is used to indicate whether there is an abnormality in the business status.
9. A full-link monitoring system based on microservices and cloud platform architecture, characterized in that: include: The service call information recording module is used to use the distributed tracking and recording system to record the detailed information of each service call to obtain the time queue of the microservice call detailed information; A data transmission module, used to upload the time queue of the microservice call detailed information to a full-link monitoring center based on a cloud platform; A service call information semantic encoding module is used to semantically encode each microservice call detailed information in the time queue of the microservice call detailed information in the full-link monitoring center based on the cloud platform to obtain a time queue of the microservice call detailed information semantic encoding vector; A semantic differentiation and enhancement module is used to perform distinguishable enhancement based on essential features on the time queue of the microservice call detailed information semantic coding vector in a full-link monitoring center based on a cloud platform to obtain a time queue of the microservice call detailed information enhanced semantic coding vector; A full-link semantic association coding module is used to perform service full-link semantic association coding on the time queue of the microservice call detailed information reinforcement semantic coding vector based on the dependency relationship between services in a full-link monitoring center based on a cloud platform to obtain a service full-link temporal semantic coding matrix; The business status identification module is used to determine the identification result of the business status in the full-link monitoring center based on the cloud platform, based on the service full-link semantic association coding to obtain the service full-link temporal semantic coding matrix.
Citation Information
Cited By
Data center east-west flow anomaly detection method and system based on deep learning
CN121690849A
A Deep Learning-Based Method and System for Detecting East-West Traffic Anomalies in Data Centers
CN121690849B