Log management method and device of distributed system, equipment and storage medium
By assigning target tracking identifiers to target requests and storing link tracking data and tracking logs in distributed databases, the problems of high resolution complexity and high maintenance costs in existing log management methods are solved, and the standardized association and integration of logs are realized, which improves the efficiency and accuracy of log management.
Patent Information
- Application Number
- CN202510044900.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
AI Technical Summary
When facing different log formats and massive log data, the existing log management methods have high resolution complexity, high maintenance costs, and it is difficult to quickly and accurately locate and associate related log records, which affects the discovery and optimization of system problems.
By assigning a target tracking identifier to the received target request and recording the link tracking data in the execution link, the target tracking identifier is added to the log content, obtaining the tracking log, and storing the link tracking data and the tracking log in a distributed database.
The standardized association and integration of logs are realized, which reduces the analytical complexity caused by different log formats, reduces maintenance costs, improves the overall order of log management, and can quickly locate problems in the business execution link.
Smart Images

Figure CN120011422A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data processing technology, and in particular to a log management method, apparatus, device and storage medium for a distributed system. Background Art
[0002] In the field of log management in the prior art, a common approach is to collect change logs from multiple data sources through log data capture tools. The collected logs will be stored in a distributed or centralized manner.
[0003] However, this traditional log management method has many problems. Since log data capture tools come from a wide range of sources, the log formats collected by different tools are often different, and different application systems themselves also generate a variety of log formats, which causes downstream processing programs to need to parse according to different log formats. Every time a new log format is added or a log data capture tool is replaced, the downstream processing program may need to be adjusted and modified accordingly, greatly increasing the processing complexity and maintenance costs. In addition, when it comes to querying, analyzing, and mining valuable information from massive log data, the efficiency of traditional methods is relatively low. It is difficult to quickly and accurately locate and associate related log records to restore the complete business execution chain, which is not conducive to timely discovery of problems in the system and targeted optimization and decision-making. Summary of the invention
[0004] The embodiments of the present application provide a log management method, apparatus, device and storage medium for a distributed system, which can avoid the parsing complexity problem caused by different log formats, greatly reduce the maintenance cost of downstream processing programs for adjusting different formats, standardize the association and integration methods of logs, and improve the overall orderliness of log management. The technical solution is as follows.
[0005] On the one hand, a log management method for a distributed system is provided, the method comprising: Assigning a target tracking identifier to the received target request; In the execution link of the target request, recording the link tracking data of the execution link, and adding the target tracking identifier to each log content associated with the execution link to obtain each tracking log, wherein the link tracking data includes the target tracking identifier; The link tracking data and the respective tracking logs are stored in a distributed database.
[0006] In another aspect, a log management device for a distributed system is provided, the device comprising: An identification allocation module, used for allocating a target tracking identification to a received target request; An information processing module, configured to record link tracking data of the execution link in the execution link of the target request, and to add the target tracking identifier to each log content associated with the execution link to obtain each tracking log, wherein the link tracking data includes the target tracking identifier; The storage module is used to store the link tracking data and the various tracking logs in a distributed database.
[0007] In a possible implementation, the information processing module includes: An interception submodule, used for intercepting the logging method of each log framework in the execution link through a monitoring agent component; The log acquisition submodule is used to acquire the log content generated by each log framework based on the log recording method of each log framework through the monitoring agent component; The directional submodule is used to direct the log content generated by each log framework to the logging application program interface of the observability framework through the monitoring agent component; The adding submodule is used to add the target tracking identifier to the log content generated by each log framework through the log recording application program interface to obtain each tracking log.
[0008] In a possible implementation, the log content generated by each log framework includes a call chain interval identifier and call chain interval context information.
[0009] In a possible implementation, when the distributed database includes a column-based database, the storage module is used to: Exporting the link tracking data and each of the tracking logs to a collector through the monitoring agent component; Each of the tracking logs is exported to the corresponding column-based database for storage through the collector.
[0010] In a possible implementation, when the distributed database includes a vector database, the storage module is used to: Exporting the link tracking data and each of the tracking logs to a collector through the monitoring agent component; Convert each of the tracking logs into corresponding vector data through the collector, and convert the link tracking data into a vector index; The vector index and the vector data corresponding to each of the tracking logs are stored in the corresponding vector databases through the collector.
[0011] In a possible implementation manner, the device further includes: A receiving module, used for receiving a query request; A log retrieval module, used to perform log retrieval based on the query request to obtain the corresponding tracking log; The log tracking module is used to perform log tracking based on the tracking identifier contained in the tracking log and the stored link tracking data to obtain the query result of the query request.
[0012] In a possible implementation manner, the device further includes: A solution generation module is used to input the link tracking data and the query result of the query request into a large language model to obtain a problem processing solution output by the large language model.
[0013] To summarize, the log management device for a distributed system provided in an embodiment of the present application obtains a tracking log by assigning a target tracking identifier to the received target request, and adding the identifier to the log contents of each execution link association, and stores the link tracking data and each tracking log in a distributed database, so that in the entire distributed system, all logs related to the target request can be associated based on the target tracking identifier, thereby avoiding the parsing complexity problem caused by different log formats, greatly reducing the maintenance cost of downstream processing programs for adjusting different formats, standardizing the log association and integration methods, and improving the overall orderliness of log management.
[0014] In addition, the link tracking data including the target tracking identifier is recorded in the execution link of the target request, and the tracking log is stored together with the link tracking data. When it is necessary to understand the complete business execution link of the request, the corresponding data can be extracted from the stored link tracking data and tracking log with the help of the target tracking identifier, and the various links from entering the system to the final processing completion of the request, as well as the details of each link, can be clearly and completely restored. This can quickly locate the location of the problem in the business execution link, help to promptly discover faults, performance bottlenecks and other problems in the system, and provide strong support for targeted optimization of the system and making reasonable decisions.
[0015] On the other hand, a computer device is provided, the computer device comprising a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the above-mentioned log management method for a distributed system.
[0016] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned log management method for a distributed system.
[0017] On the other hand, a computer program product is provided, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute to implement the log management method of the distributed system provided in the above-mentioned various optional implementations.
[0018] The technical solution provided by this application may have the following beneficial effects: The log management method for a distributed system provided in an embodiment of the present application obtains a tracking log by assigning a target tracking identifier to a received target request, and adding the identifier to each log content of an execution link association, and stores the link tracking data and each tracking log in a distributed database, so that in the entire distributed system, all logs related to the target request can be associated based on the target tracking identifier, thereby avoiding the parsing complexity problem caused by different log formats, greatly reducing the maintenance cost of downstream processing programs for adjusting different formats, standardizing the log association and integration methods, and improving the overall orderliness of log management.
[0019] In addition, the link tracking data including the target tracking identifier is recorded in the execution link of the target request, and the tracking log is stored together with the link tracking data. When it is necessary to understand the complete business execution link of the request, the corresponding data can be extracted from the stored link tracking data and tracking log with the help of the target tracking identifier, and the various links from entering the system to the final processing completion of the request, as well as the details of each link, can be clearly and completely restored. This can quickly locate the location of the problem in the business execution link, help to promptly discover faults, performance bottlenecks and other problems in the system, and provide strong support for targeted optimization of the system and making reasonable decisions.
[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0022] Figure 1 A flow chart showing a log management method for a distributed system provided by an exemplary embodiment of the present application is shown; Figure 2 A flow chart showing a log management method for a distributed system provided by an exemplary embodiment of the present application is shown; Figure 3 A flow chart showing a log management method for a distributed system provided by an exemplary embodiment of the present application is shown; Figure 4 A schematic diagram of performing log retrieval through RAG technology provided by an exemplary embodiment of the present application is shown; Figure 5 A schematic diagram of a log management method for a distributed system provided by an exemplary embodiment of the present application is shown; Figure 6 A block diagram of a log management device for a distributed system provided by an exemplary embodiment of the present application is shown; Figure 7 A structural block diagram of a computer device shown in an exemplary embodiment of the present application is shown; Figure 8 A structural block diagram of a computer device shown in an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0024] An embodiment of the present application provides a log management method for a distributed system, which tracks the flow of target requests in the distributed system, including the services passed, the response time of each service, and the service content, by associating link tracking data with logs. This ensures that subsequent storage and query analysis links will not be affected by the diversity of log formats, reduces the parsing complexity caused by different log formats of distributed systems, reduces the maintenance cost of log management, and improves the log management effect. Figure 1 A flow chart of a log management method for a distributed system provided by an exemplary embodiment of the present application is shown. The method can be executed by a distributed system, and the distributed system can be constructed based on a server or a terminal device, such as Figure 1 As shown, the method may include the following steps.
[0025] Step 110: assign a target tracking identifier to the received target request.
[0026] After receiving a target request initiated externally or internally, the distributed system starts an identification allocation mechanism. The target tracking identification (Trace ID) is unique and can be generated by the distributed system based on a preset algorithm. Schematically, the distributed system can generate a random sequence by combining the current system timestamp, the encoding of the request source, and a random number generator. In addition, it can also combine hash functions for mixed operations to ensure the uniqueness of the generated target tracking identification, so that the request and related logs can be accurately tracked and associated later; for example, the distributed system can use a universally unique identifier (UUID, Universally Unique Identifier) generation algorithm to generate a target tracking identification.
[0027] The target tracking identifier is used to track the entire link of the target request in a distributed system environment, and to associate the relevant information generated by the target request during the processing of various services and modules, so as to facilitate the subsequent analysis of the entire request processing process, troubleshoot problems, and understand the operating status of the system.
[0028] Step 120, in the execution link of the target request, record the link tracking data of the execution link, and add the target tracking identifier to each log content associated with the execution link to obtain each tracking log, wherein the link tracking data includes the target tracking identifier.
[0029] The execution link is a series of services, components, and call relationship paths between services / components that the target request passes through from entry to final completion in the distributed system; the link tracking data is detailed record information about the execution link. In a possible implementation, multiple monitoring nodes can be deployed in the distributed system. After the target request enters the execution link, the distributed system can record the link tracking data through multiple monitoring nodes. These monitoring nodes can be distributed in key links of the execution link, such as the request initiation point, data access layer, core business logic processing components, etc. The link tracking data can include service call processes and performance indicators, etc., such as events of requests passing through each node, node processing status (including success, failure, waiting, etc.), occupied system resources (such as CPU usage, memory usage, etc.) and detailed information on interactions with upstream and downstream components.
[0030] The tracking log is the log content that is generated by each service and component in the execution link when processing the target request and added with the target tracking identifier. While recording the link tracking data of the execution link, the distributed system can also use the log injection tool to add the target tracking identifier assigned to the target request to each log content associated with the execution link; in one possible implementation, when adding the target tracking identifier, the distributed system can add the target tracking identifier to the specified position of the log content to facilitate identification. Schematically, the target tracking identifier can be added at the beginning or end of the log content, or other specific field positions, and this application does not limit this.
[0031] Step 130: Store the link tracking data and each tracking log in a distributed database.
[0032] A distributed database is a database system that stores and manages data in a dispersed manner on multiple physical nodes (such as different physical servers or virtual machines). Distributed databases can use the network to connect these physical nodes and work together to store, manage, and process data. Distributed databases can also handle a large number of concurrent read and write operations, providing high availability and scalability.
[0033] In one possible implementation, the distributed system may select an adaptive type of database table structure or data storage model for storage according to a storage strategy preset by the system for data with different characteristics; optionally, the distributed database may include a columnar database and a vector database; wherein, a columnar database is a database that stores data by column, and data in the same column is stored continuously in physical storage; a vector database is constructed based on a vector space model, and can convert various data, such as text, images, audio, etc., into vector form for storage through a specific algorithm.
[0034] To summarize, the log management method for a distributed system provided in an embodiment of the present application obtains a tracking log by assigning a target tracking identifier to the received target request, and adding the identifier to the log contents of each execution link association, and stores the link tracking data and each tracking log in a distributed database, so that in the entire distributed system, all logs related to the target request can be associated based on the target tracking identifier, avoiding the parsing complexity problem caused by different log formats, greatly reducing the maintenance cost of downstream processing programs for adjusting different formats, standardizing the log association and integration methods, and improving the overall orderliness of log management.
[0035] In addition, the link tracking data including the target tracking identifier is recorded in the execution link of the target request, and the tracking log is stored together with the link tracking data. When it is necessary to understand the complete business execution link of the request, the corresponding data can be extracted from the stored link tracking data and tracking log with the help of the target tracking identifier, and the various links from entering the system to the final processing completion of the request, as well as the details of each link, can be clearly and completely restored. This can quickly locate the location of the problem in the business execution link, help to promptly discover faults, performance bottlenecks and other problems in the system, and provide strong support for targeted optimization of the system and making reasonable decisions.
[0036] Figure 2 A flow chart of a log management method for a distributed system provided by an exemplary embodiment of the present application is shown. The method can be executed by a distributed system, and the distributed system can be constructed based on a server or a terminal device, such as Figure 2 As shown, the method may include the following steps.
[0037] Step 210: assign a target tracking identifier to the received target request.
[0038] In a distributed system, each request is assigned a unique identifier (i.e., Trace ID), which is used throughout the entire request processing process. Link tracing technology can track the entire request processing process by recording this unique identifier and its transmission path between various services or components.
[0039] Step 220 , in the execution link of the target request, record the link tracking data of the execution link.
[0040] Step 230, intercepting the logging method of each log framework in the execution link through the monitoring agent component.
[0041] In an embodiment of the present application, a distributed system can implement the addition of a target tracking identifier through a monitoring agent component combined with a logging application program interface of an observability framework; in one possible implementation, the monitoring agent component may include monitoring components corresponding to common logging frameworks such as Java Util Logging, Log4j, and Logback. Through bytecode injection technology, the monitoring agent component identifies and intercepts the logging methods of each logging framework used in the execution link, wherein the logging method is a specific function or operation method used by the logging framework to actually record logs. Different logging frameworks have their own corresponding logging methods to output log content in accordance with specific formats, levels, and other requirements. Schematically, the logging methods under different logging frameworks may be java.util.logging.Logger, org.apache.logging.log4j.Logger, and the like; the monitoring agent component intercepts the logging methods of each logging framework in the execution link, thereby implementing corresponding processing and management of the logs to be generated, so as to facilitate the subsequent association of the logs with link tracking data.
[0042] Step 240: The log content generated by each log framework is obtained through the monitoring agent component based on the log recording method of each log framework.
[0043] When the monitoring agent component successfully intercepts the logging methods of each log framework, it can obtain the log content actually generated by these log frameworks at the corresponding time. Different log frameworks will include the specific content of the log output, the log severity level, and the generation time when recording logs. Taking a logging scenario as an example, when the Log4j framework records an error log, it will record specific content such as "[ERROR] Database connection failed, error reason: xxx", and will also include the level of the log (such as ERROR level) and the timestamp of the log generation. The monitoring agent component can obtain the log content completely through the interception mechanism to prepare for further processing.
[0044] In a possible implementation, the log content generated by each log framework includes a call chain interval identifier and call chain interval context information. That is, when recording logs, the distributed system can obtain the active Sqan (call chain interval) on the current code path through the monitoring application program interface (i.e., monitoring API), and obtain the current context information, i.e., SpanContext (call chain interval context) information, from the Sqan, wherein the call chain interval context information may include some environmental information, associated information, etc. of the current Span, such as the Sqan ID of the parent Sqan of the current Sqan, Sqan metadata, etc.; when recording logs, the call chain interval context information of the Sqan and the unique identifier (Sqan) representing the Sqan are recorded in the log content, so that when viewing the logs later, the entire call chain can be restored based on this information, which is convenient for troubleshooting, performance analysis, etc. For example, when querying, the order and dependency between the call chain interval corresponding to the current log content and other operations can be known.
[0045] Optionally, the log content and call chain interval context information contained in the log content captured by the monitoring agent component from the log framework are recorded in a log structured data format, without the need for cutting and parsing. Therefore, real-time monitoring and analysis of the log can be achieved without waiting for the log cutting and loading process. Real-time monitoring and analysis enables the system to respond and process log data more quickly, reducing the need for manual intervention, thereby discovering and solving problems more promptly. Among them, the log structured data format can include the following fields: log content, log generation time, tracking identifier, call chain interval identifier, log severity level, application service name, resource attributes, and class of printing logs, etc.
[0046] Step 250, directing the log content generated by each log framework to the logging application program interface of the observability framework through the monitoring agent component.
[0047] After obtaining the log content generated by each log framework, the monitoring agent component redirects the obtained log content to the logging API of the observability framework, such as the Logging API of Opentelemetry (OTel). Opentelemetry is an open and standardized observability framework that aims to solve the problem of unified collection, processing and export of observable data in distributed systems. It is used to monitor, generate, collect and export telemetry data, such as traces, metrics and logs. Its Logging API plays a key role in this process. The monitoring agent component will redirect the log content intercepted from each log framework to the Logging API according to pre-set rules and mechanisms. For example, the monitoring agent component will pass the previously obtained complete log content including log output content, severity level, generation time, etc. to the Logging API, so that these log contents can enter the observability system built by the Otel framework.
[0048] Step 260 , adding a target tracking identifier to the log content generated by each log framework through the log recording application program interface to obtain each tracking log.
[0049] Through the logging application program interface, the target tracking identifier is added to the log content generated by each log framework that has been directed. In a distributed system, the tracking identifier is an important basis for tracking the complete execution of a specific operation. By adding the target tracking identifier to the log content, combined with the call chain interval identifier added during logging, the association between the context link information can be achieved, and the log content can be associated with the link tracking data; that is, the log content contains unique identifiers related to link tracking, namely the target tracking identifier (Trace Id) and the call chain interval identifier (Span Id). These unique identifiers can track the complete execution path of a specific operation during monitoring and troubleshooting.
[0050] In a possible implementation, the Logging API is also used to structure the received log content to better support the analysis and retrieval of log data, wherein the structured processing of the log content may include unifying the format and structure of the log level, timestamp, log message, and other metadata information.
[0051] Step 270: Store the link tracking data and each tracking log in a distributed database.
[0052] In one possible implementation, the distributed system can realize distributed storage of log content through a collector. In one possible implementation, after the monitoring agent component exports the link tracking data and each tracking log to the collector, the collector will process the received data, and the data processing may include operations such as data conversion, filtering, format conversion, and data aggregation.
[0053] When storing logs, the distributed system can store based on data. Schematically, link tracking data has the characteristics of being highly structured and frequently used for associated queries, and the link tracking data can be stored through a column-based database; and for each tracking log, on the one hand, it can be stored in a column-based database to facilitate subsequent rapid retrieval based on certain fields of the log content. On the other hand, the log conversion tool can be used to convert each tracking log into corresponding vector data, and at the same time, the link tracking data can be converted into a vector index, and the vector index and the vector data corresponding to each tracking log can be stored in their respective corresponding vector databases for storage, so as to facilitate the subsequent use of vector similarity for log retrieval and association analysis. In other words, when the distributed database includes a column-based database, the process of log storage can be implemented as follows: Export link tracking data and various tracking logs to the collector through the monitoring agent component; Each tracking log is exported to its corresponding column database for storage through the collector.
[0054] Indicatively, the columnar database may be a Clickhouse database, which is a columnar database management system for large data volume scenarios, which can efficiently process complex queries of large-scale data sets and can implement rapid aggregation, filtering and analysis operations on logs. The columnar storage characteristics and data compression function of the Clickhouse database can reduce the storage space occupied, is suitable for large-scale log storage, and supports horizontal expansion. When the data volume increases, more nodes can be added to expand performance and capacity to cope with the growing data volume and query requirements.
[0055] In the case where the distributed database includes a vector database, the process of log storage can be implemented as follows: Export link tracking data and various tracking logs to the collector through the monitoring agent component; Convert each tracking log into corresponding vector data through a collector, and convert link tracking data into vector index; The vector index and the vector data corresponding to each tracking log are stored in their respective vector databases through the collector.
[0056] Illustratively, the collector can convert the tracking logs and link tracking data stored in the log structured data format into vector data and store them in the vector database by calling the BERT-as-Service service interface.
[0057] That is to say, when performing vectorization processing, the distributed system can also process the tracking logs through a pre-trained language model, convert the original tracking logs of various forms into vectors of fixed length, and store the vector data in a vector database, wherein the vector database can be a Milvus database. The Milvus database has a powerful vector similarity search function, which can quickly find other vectors similar to a given vector in a large amount of stored vector data based on the vector similarity calculation method.
[0058] After completing the vector storage, the distributed system can also establish a mapping relationship between the vector data and the tracking log based on a non-relational database, so that when performing subsequent vector retrieval, the result of the vector retrieval can be associated with the original tracking log; optionally, the non-relational database can be a MongoDB database.
[0059] To summarize, the log management method for a distributed system provided in an embodiment of the present application obtains a tracking log by assigning a target tracking identifier to the received target request, and adding the identifier to the log contents of each execution link association, and stores the link tracking data and each tracking log in a distributed database, so that in the entire distributed system, all logs related to the target request can be associated based on the target tracking identifier, avoiding the parsing complexity problem caused by different log formats, greatly reducing the maintenance cost of downstream processing programs for adjusting different formats, standardizing the log association and integration methods, and improving the overall orderliness of log management.
[0060] In addition, the link tracking data including the target tracking identifier is recorded in the execution link of the target request, and the tracking log is stored together with the link tracking data. When it is necessary to understand the complete business execution link of the request, the corresponding data can be extracted from the stored link tracking data and tracking log with the help of the target tracking identifier, and the various links from entering the system to the final processing completion of the request, as well as the details of each link, can be clearly and completely restored. This can quickly locate the location of the problem in the business execution link, help to promptly discover faults, performance bottlenecks and other problems in the system, and provide strong support for targeted optimization of the system and making reasonable decisions.
[0061] Based on Figure 1 or Figure 2After the embodiment shown completes the storage of the log content in the execution link corresponding to the request, in the scenario of log query, the distributed system can perform log query and log tracking based on the received query request to complete the link log content. Figure 3 A flow chart of a log management method for a distributed system provided by an exemplary embodiment of the present application is shown. The method can be executed by a distributed system, and the distributed system can be constructed based on a server or a terminal device, such as Figure 3 As shown, the method may include the following steps.
[0062] Step 310: Receive a query request.
[0063] The distributed system can receive query requests from different sources through the request receiving interface, such as query requests initiated by relevant personnel through the query interface, or query requests automatically triggered by the monitoring system, etc.; the query request can include query parameters to clarify the scope of the logs to be queried, such as a specified time interval, a specific business operation type, the name of the service involved or the tracking identifier and other related content.
[0064] Step 320: perform log retrieval based on the query request to obtain the corresponding tracking log.
[0065] After receiving the query request, the distributed system can parse the query request to obtain the query parameters in the query request, and perform log retrieval in the distributed database according to the query parameters to find matching tracking logs. In the embodiment of the present application, the distributed system can perform log retrieval through RAG (Retrieval-Augmented Generation) technology. Figure 4 FIG. 1 shows a schematic diagram of performing log retrieval by using RAG technology provided by an exemplary embodiment of the present application, such as Figure 4 As shown, in the case where the distributed database includes a vector database, the process can be implemented as follows: When performing log retrieval, the distributed system can perform log retrieval through a text search engine. The text search engine can obtain query parameters in the query request and convert the query parameters into retrieval vectors; The vector database can calculate the vector similarity between the search vector and each vector data stored in the vector database; Based on the vector similarity between each vector data, the tracking log corresponding to the query request is determined.
[0066] After determining the vector data based on vector similarity, the distributed system can determine the tracking log corresponding to the vector data based on the mapping relationship between the vector data and the tracking log established in the non-relational database; optionally, the distributed system can determine the tracking log corresponding to the vector data with the highest vector similarity between the query parameters in the query request as the tracking log corresponding to the query request; or, the tracking logs corresponding to the first n vector data with the highest similarity can be determined as the tracking logs corresponding to the query request, where n is a positive integer, and this application does not impose any restrictions on this.
[0067] Step 330 , log tracing is performed based on the tracing identifier contained in the tracing log and the stored link tracing data to obtain a query result of the query request.
[0068] Since the tracking log obtained after log retrieval based on the query request may be a scattered part of the entire business execution link, in order to obtain the comprehensive log content of the business execution link, the distributed system can also extract the tracking identifier based on the retrieved tracking log. On the one hand, the tracking log containing the tracking identifier is searched in the distributed database. On the other hand, the corresponding link tracking data is searched in the distributed database with the tracking identifier as a clue. After obtaining the link tracking data, the tracking logs with the same tracking identifier are associated and integrated in the correct business execution link order according to the service node sequence, time sequence and other information recorded therein, and the complete log content of each link related to the query request is restored as the query result of the query request, thereby comprehensively presenting the flow of the corresponding business operation in the system, facilitating subsequent analysis, troubleshooting and other work.
[0069] In one possible implementation, the distributed system can use the large language model for subsequent analysis and processing. In the troubleshooting scenario, this process can be implemented as follows: The link tracking data and the query results of the query request are input into the large language model to obtain the problem processing solution output by the large language model.
[0070] That is to say, the distributed system can perform abnormal diagnosis and analysis based on link tracking data and log content through the integration of AIGC (Artificial Intelligence Generated Content) technology. In a possible implementation scenario, the distributed system can combine the prompt words built in the troubleshooting scenario, input the link tracking data and related log content into the large language model to instruct the large language model to perform troubleshooting and make suggestions based on the provided data, thereby obtaining the problem location results and problem handling solutions output by the large language model, thereby improving the fault diagnosis capability.
[0071] In one possible implementation, the large language model can sort out the complete business process of the request based on the service node sequence and timestamp in the link tracking data, match the log content in the query results with each link in the business process, and compare the timestamps, return results, etc. in the actual link and final data with the standard process, so as to perform troubleshooting and propose problem handling solutions based on the troubleshooted faults; optionally, during the comparison process, the large language model can also combine the mutual influence between different links in the business process and the troubleshooted faults to trace the source of the fault, thereby determining the source fault and proposing problem handling solutions for the source fault.
[0072] To sum up, the log management method for a distributed system provided in an embodiment of the present application, based on associating log content with link tracking data through tracking identifiers, can, upon receiving a log query request, perform log retrieval based on the query request to obtain a partial tracking log, and can then further perform log tracking based on the tracking identifier and link tracking data in the tracking log to obtain a query result corresponding to the query request, thereby improving the retrieval completeness and accuracy.
[0073] In addition, with the help of a large language model, corresponding solutions are generated based on link tracking data and related log content. Combined with RAG and artificial intelligence content generation technology, the link tracking data and log content are integrated for abnormal diagnosis and analysis, which improves the efficiency of abnormal fault diagnosis, thereby quickly locating fault logs and providing solutions, significantly improving fault diagnosis capabilities.
[0074] Indicative, Figure 5 A schematic diagram of a log management method for a distributed system provided by an exemplary embodiment of the present application is shown. The method is executed by various components in the distributed system, such as Figure 5 As shown, after the tracking identifier is assigned to the request, in the execution link corresponding to the request, the process can be implemented as follows: When storing logs, S501, the monitoring agent component monitors each log framework in the execution link.
[0075] S502, the monitoring agent component redirects the intercepted log content generated by each log framework to the Logging API of Opentelemetry.
[0076] S503, Opentelemetry's Logging API structures the log content.
[0077] S504, Opentelemetry's Logging API associates the log content with the link tracking data through the tracking identifier.
[0078] In an embodiment of the present application, the Logging API associates the log content with the link tracking data by adding a tracking identifier that is the same as the link tracking data in the log content.
[0079] S505, the monitoring agent component exports the associated log content and link tracking data to the collector for data processing.
[0080] After the monitoring agent redirects the intercepted log content to the Logging API for processing, and the Logging API completes a series of processing including association, the monitoring agent exports the log content and link tracking data that have been processed by the Logging API to the collector.
[0081] After the collector receives the log data transmitted from the monitoring agent component, it can process and convert the data, such as filtering, format conversion, data aggregation, etc. Based on actual needs and settings, the data processing operations of the collector can be different.
[0082] S506, the collector exports the processed data to a distributed database for storage.
[0083] On the one hand, the collector can forward the data to the Clickhouse database for storage; on the other hand, the collector can convert the log content into vector data by calling the BERT-as-Service service interface and combining the various fields of the log content, and store it in a vector database (for example, Milvus).
[0084] When querying the log, S507, receiving a query request.
[0085] S508: Perform log retrieval based on the query request to obtain corresponding tracking logs.
[0086] S509: Perform log tracking based on the tracking identifier contained in the tracking log and the stored link tracking data to obtain a query result of the query request.
[0087] Figure 6 A block diagram of a log management device for a distributed system provided by an exemplary embodiment of the present application is shown. The device can perform the following steps: Figures 1 to 3 All or part of the steps of any of the illustrated embodiments, such as Figure 6 As shown, the device comprises: The identifier allocation module 610 is used to allocate a target tracking identifier to the received target request; An information processing module 620 is used to record the link tracking data of the execution link in the execution link of the target request, and add the target tracking identifier to each log content associated with the execution link to obtain each tracking log, wherein the link tracking data includes the target tracking identifier; The storage module 630 is used to store the link tracking data and the various tracking logs in a distributed database.
[0088] In a possible implementation, the information processing module 620 includes: An interception submodule, used for intercepting the logging method of each log framework in the execution link through a monitoring agent component; The log acquisition submodule is used to acquire the log content generated by each log framework based on the log recording method of each log framework through the monitoring agent component; The directional submodule is used to direct the log content generated by each log framework to the logging application program interface of the observability framework through the monitoring agent component; The adding submodule is used to add the target tracking identifier to the log content generated by each log framework through the log recording application program interface to obtain each tracking log.
[0089] In a possible implementation, the log content generated by each log framework includes a call chain interval identifier and call chain interval context information.
[0090] In a possible implementation, when the distributed database includes a column-based database, the storage module 630 is configured to: Exporting the link tracking data and each of the tracking logs to a collector through the monitoring agent component; Each of the tracking logs is exported to the corresponding column-based database for storage through the collector.
[0091] In a possible implementation, when the distributed database includes a vector database, the storage module 630 is used to: Exporting the link tracking data and each of the tracking logs to a collector through the monitoring agent component; Convert each of the tracking logs into corresponding vector data through the collector, and convert the link tracking data into a vector index; The vector index and the vector data corresponding to each of the tracking logs are stored in the corresponding vector databases through the collector.
[0092] In a possible implementation manner, the device further includes: A receiving module, used for receiving a query request; A log retrieval module, used to perform log retrieval based on the query request to obtain the corresponding tracking log; The log tracking module is used to perform log tracking based on the tracking identifier contained in the tracking log and the stored link tracking data to obtain the query result of the query request.
[0093] In a possible implementation manner, the device further includes: A solution generation module is used to input the link tracking data and the query result of the query request into a large language model to obtain a problem processing solution output by the large language model.
[0094] Figure 7 The block diagram of the structure of a computer device 700 shown in an exemplary embodiment of the present application is shown. The computer device can be implemented as a server constituting a distributed system in the above-mentioned solution of the present application. The computer device 700 includes a central processing unit (CPU) 701, a system memory 704 including a random access memory (RAM) 702 and a read-only memory (ROM) 703, and a system bus 705 connecting the system memory 704 and the central processing unit 701. The computer device 700 also includes a large-capacity storage device 706 for storing an operating system 709, an application program 710 and other program modules 711. The above-mentioned system memory 704 and the large-capacity storage device 706 can be collectively referred to as a memory.
[0095] According to various embodiments of the present disclosure, the computer device 700 can also be connected to a remote computer on the network through a network such as the Internet. That is, the computer device 700 can be connected to the network 708 through the network interface unit 707 connected to the system bus 705, or the network interface unit 707 can be used to connect to other types of networks or remote computer systems (not shown).
[0096] The memory also includes at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is stored in the memory. The central processing unit 701 implements all or part of the steps in the log management method of the distributed system shown in the above-mentioned embodiments by executing the at least one instruction, at least one program, code set or instruction set.
[0097] Figure 8The following is a block diagram of a computer device 800 according to an exemplary embodiment of the present application. The computer device 800 may be implemented as a terminal constituting a distributed system in the above solution, such as a smart phone, a tablet computer, a notebook computer, a desktop computer, etc. The computer device 800 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other names.
[0098] Typically, the computer device 800 includes a processor 801 and a memory 802 .
[0099] In some embodiments, the computer device 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802 and the peripheral device interface 803 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 803 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807 and a power supply 808.
[0100] In some embodiments, the computer device 800 further includes one or more sensors 809 , including but not limited to: an acceleration sensor 810 , a gyroscope sensor 811 , a pressure sensor 812 , an optical sensor 813 , and a proximity sensor 814 .
[0101] Those skilled in the art will understand that Figure 8 The structure shown in the figure does not constitute a limitation on the computer device 800, and the computer device 800 may include more or less components than those shown in the figure, or combine some components, or adopt a different arrangement of components.
[0102] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one computer program is stored, and the computer program is loaded and executed by a processor to implement the above-mentioned log management method for a distributed system, and / or all or part of the steps in the log management method for a distributed system. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0103] In an exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, when the program instructions are executed by a computer, the computer is executed to implement the above Figure 1 , Figure 2 or Figure 3 All or part of the steps of the log management method for a distributed system shown in the embodiment.
[0104] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.
[0105] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A log management method for a distributed system, characterized in that: The method comprises: Assigning a target tracking identifier to the received target request; In the execution link of the target request, recording the link tracking data of the execution link, and adding the target tracking identifier to each log content associated with the execution link to obtain each tracking log, wherein the link tracking data includes the target tracking identifier; The link tracking data and the respective tracking logs are stored in a distributed database.
2. The method according to claim 1, characterized in that The step of adding the target tracking identifier to each log associated with the execution link to obtain each tracking log includes: Intercepting the logging method of each log framework in the execution link through a monitoring agent component; Obtain the log content generated by each log framework through the monitoring agent component based on the log recording method of each log framework; Direct the log content generated by each log framework to the logging API of the observability framework through the monitoring agent component; The target tracking identifier is added to the log content generated by each log framework through the log recording application program interface to obtain each tracking log.
3. The method according to claim 2, characterized in that The log content generated by each log framework includes the call chain interval identifier and the call chain interval context information.
4. The method according to claim 2, characterized in that: In the case where the distributed database includes a column-based database, storing the link tracking data and the respective tracking logs in the distributed database includes: Exporting the link tracking data and each of the tracking logs to a collector through the monitoring agent component; Each of the tracking logs is exported to the corresponding column-based database for storage through the collector.
5. The method according to claim 2, characterized in that: In the case where the distributed database includes a vector database, storing each of the tracking logs includes: Exporting the link tracking data and each of the tracking logs to a collector through the monitoring agent component; Convert each of the tracking logs into corresponding vector data through the collector, and convert the link tracking data into a vector index; The vector index and the vector data corresponding to each of the tracking logs are stored in the corresponding vector databases through the collector.
6. The method according to claim 1, characterized in that The method comprises: receiving a query request; Perform log retrieval based on the query request to obtain corresponding tracking logs; Log tracking is performed based on the tracking identifier contained in the tracking log and the stored link tracking data to obtain a query result of the query request.
7. The method according to claim 6, characterized in that The method further comprises: The link tracking data and the query result of the query request are input into a large language model to obtain a problem processing solution output by the large language model.
8. A log management device for a distributed system, characterized in that: The device comprises: An identification allocation module, used for allocating a target tracking identification to a received target request; An information processing module, configured to record link tracking data of the execution link in the execution link of the target request, and to add the target tracking identifier to each log content associated with the execution link to obtain each tracking log, wherein the link tracking data includes the target tracking identifier; The storage module is used to store the link tracking data and the various tracking logs in a distributed database.
9. A computer device, characterized in that: The computer device includes a processor and a memory, the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the log management method for a distributed system as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: At least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by the processor to implement the log management method for a distributed system as described in any one of claims 1 to 7.
Citation Information
Cited By
Full-stack observability method of unified platform
CN120448227A
Distributed tracking data acquisition method and device based on log analysis and computer readable medium
CN121125465A
Ceph monitoring, log and link tracking integrated method and system
CN121547375A
Inter-system data exchange log tracing system and method
CN121958216A
Database access behavior tracking method and device
CN122387992A