A method for troubleshooting user-level service problems based on service logs
Patent Information
- Application Number
- CN202311528722.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-11-16
AI Technical Summary
[0003]为了解决上述现有技术中存在的问题,本发明拟提供了一种基于业务日志实现用户级业务问题的排查方法,拟解决现有技术多人测试遇到业务异常时,测试人员无法精准获取信息,定位问题耗时多难度大的问题
[0034]1. User Information Transmission: By obtaining the request's traceid from the API gateway and the user's unique identifier (user ID) from the user center, the user ID is placed into the traceid's structure, enabling transparent transmission of the user ID across all microservices. This allows for tracing specific users based on their user IDs during call chain analysis, thus aiding in problem localization.
Smart Images

Figure CN117687982B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of testing technology, and more specifically to a method for troubleshooting user-level business issues based on business logs. Background Technology
[0002] Currently, enterprises commonly use distributed microservices to implement business systems. When conducting end-to-end testing of the user interface, the sheer number of microservices used in the business implementation creates a highly complex testing environment. A single business function may require dozens or even hundreds of microservices to achieve its specific functionality. Furthermore, the testing environment is used by multiple people, resulting in a mix of various business information generated by these users. When anomalies occur, testers struggle to distinguish which information originated from their own testing, leading to inaccurate information retrieval, time-consuming problem localization, and significant difficulty. Summary of the Invention
[0003] To address the problems existing in the prior art, this invention proposes to provide a method for troubleshooting user-level business issues based on business logs. This method aims to solve the problem that when multiple testers encounter business anomalies, testers cannot accurately obtain information, and locating the problem is time-consuming and difficult.
[0004] A method for troubleshooting user-level business issues based on business logs includes the following steps:
[0005] Step 1: Obtain the user's unique identifier (user id) and add it to the structure containing the business call chain traceid to achieve transparent transmission of user information;
[0006] Step 2: Combine the user ID and trace ID to form the flow ID and print it to the microservice log;
[0007] Step 3: Set the business code and status code for the gateway interface response and print them in the gateway interface request log along with the flowid; Step 4: Collect the core business status and caller relationships, save them to a specific database, and use them for log analysis generated on the call chain;
[0008] Step 5: Write the call trace ID back to the UI page's API response. Testers can then query the flowid field of the API response and use the log aggregation query system to locate user-level errors.
[0009] Preferably, step 1 includes:
[0010] Step 1.1: In the API gateway of the microservice system, the request information is intercepted through the encoding mechanism of the Java agent and the traceid is obtained by combining the programmable method provided by APM.
[0011] Step 1.2: Based on user information, obtain the user's unique identifier, user id, from the user center;
[0012] Step 1.3: Add a Java agent for the gateway service. Through APM, handle the specific processing logic and transparent structure of different protocols. By intercepting the application's runtime state, put the user ID into the structure where the traceid is located, so that the user ID and traceid can be transparently transmitted in all microservices.
[0013] Preferably, step 2 includes:
[0014] Step 2.1: Encapsulate the log printing method and store it in the corresponding microservice. Add a Java agent to each microservice. Intercept the call request through the Java agent and obtain the APM pass-through message body to obtain the traceid and user id.
[0015] Step 2.2: Structure the traceid and user id into a fixed-normal form of floutid;
[0016] Step 2.3: Design flowid as a mandatory output item for log printing to ensure that flowid is printed in every log print.
[0017] Preferably, step 3 includes:
[0018] Step 3.1: Design the gateway interface response and assign a service code to each service;
[0019] Step 3.2: Assign a status code for each service's different response status;
[0020] Step 3.3: After the gateway sends a request to the downstream system, it assigns different service codes and status codes to the responses of each downstream system based on the different services and response statuses.
[0021] Step 3.4: Print logs for each gateway interface request. The logs include the service code, status code, flow ID, and interface processing information.
[0022] Preferably, step 4 includes:
[0023] Step 4.1: The log visualization system monitors all logs of the gateway service and parses the logs asynchronously;
[0024] Step 4.2: Monitor the gateway service's real-time request traffic and logs to obtain requests and responses, status codes, business codes, and flow IDs, and maintain a minimal information set for rapid risk warning and retrieval. At the same time, save the additional information to a separate dataset for detailed information retrieval.
[0025] Step 4.3: Analyze the data based on the dataset saved in Step 4.2, and perform different asynchronous processing on different states.
[0026] Preferably, the asynchronous processing for different states includes: for abnormal states, immediately obtaining the flowid from the log information through a message queue mechanism, and then obtaining the user ID and traceid. Using the user ID and traceid, obtaining the relevant logs of all systems in the business chain for this business, aggregating them together, and storing them in a log visualization system for testers to visualize and query, thereby improving query efficiency; for normal processes, only the business status code and flowid are saved, and the logs are not saved. When a query is needed, the query is performed on each system.
[0027] Preferably, step 5 includes:
[0028] Step 5.1: Add a Java agent to the gateway service to intercept interface requests and obtain the request header, request content, and request response.
[0029] Step 5.2: Write the flow ID generated on the server side back into the HTTP request response header;
[0030] Step 5.3: Testers use the interface debugging mode to query the response header of the interface request, and then find the flowid of all interface requests involved in the business call.
[0031] Preferably, it also includes various information queries based on flowid;
[0032] Based on the user information input by the testers, the system queries the associated user ID, queries the associated flow ID using the user ID, and then retrieves all call chain information and log information.
[0033] The beneficial effects of this invention include:
[0034] 1. User Information Transmission: By obtaining the request's traceid from the API gateway and the user's unique identifier (user ID) from the user center, the user ID is placed into the traceid's structure, enabling transparent transmission of the user ID across all microservices. This allows for tracing specific users based on their user IDs during call chain analysis, thus aiding in problem localization.
[0035] 2. Log printing format specifications: The traceid and user id are structured into a fixed format called flood, and flood is designed as a mandatory output item in log printing, ensuring that flood is printed in every log print. This makes it possible to clearly distinguish between user logs and specific interface request logs when analyzing logs.
[0036] 3. Service Code and Status Code Standardization: Gateway interface responses are standardized, with a unique service code assigned to each service and a status code assigned to different response states for each service. This allows for rapid understanding of service execution and operational status during call chain analysis based on the service code and status code.
[0037] 4. Abnormal Process Log Aggregation: For abnormal states, the flow ID is immediately retrieved from the log information, followed by the user ID and trace ID. Using the user ID and trace ID, relevant logs from all systems along the business chain are obtained, aggregated, and stored. This allows for rapid identification of abnormal call chains during call chain analysis, improving real-time performance.
[0038] 5. Call chain traceid rewriting: The traceid generated by the request on the server side is rewritten into the HTTP request response header, allowing testers to query the request traceid through the interface debugging mode. This makes it easy to obtain the request traceid during call chain analysis, allowing users to know the log identifiers generated by their business processes.
[0039] 6. Information Query Based on Flow ID: Implement various information queries based on flow ID, including querying abnormal process logs, normal process logs, and user information. This allows for more flexible querying of required information during call chain analysis. Attached Figure Description
[0040] Figure 1 This is a flowchart of a method for troubleshooting user-level business issues based on business logs, as described in Example 1. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0042] Example 1
[0043] APM is an application performance monitoring software that provides a unified view that displays data from networks, applications, and servers monitored by a console.
[0044] Call chain: refers to the chain formed by different services calling and communicating through interfaces in a distributed system.
[0045] Trace ID (traceId): A unique identifier generated by the APM system and passed through on the call chain to distinguish a call to an interface.
[0046] User Center: A microservice that stores user-related information. It typically saves user information such as phone number, ID card number, and name, and assigns a unique user identifier, i.e., user id.
[0047] Java agent: A concept in the Java programming language used to collect and report runtime status information of an application. By using a Java agent, developers can obtain detailed information about each method executed in the application and each object, enabling them to monitor application performance and diagnose problems. Java agents are typically implemented using Java agent frameworks. These frameworks provide a set of APIs that allow developers to write their own agents and deploy them along with their applications.
[0048] The following is in conjunction with the appendix Figure 1 Specific embodiments of the present invention will be described in detail;
[0049] A method for troubleshooting user-level business issues based on business logs includes the following steps:
[0050] Step 1: Obtain the user's unique identifier (user id) and add it to the structure containing the business call chain traceid to achieve transparent transmission of user information.
[0051] Using the programmable methods provided by APM (Application Performance Monitoring System), the traceid of the call chain is obtained for each request while the microservice is running. At the same time, each microservice system obtains the user information to which the current interface request belongs and adds it to the pass-through structure used by APM.
[0052] Step 1.1: In the API gateway of the microservice system, obtain the traceid through the programmable method provided by APM;
[0053] Step 1.2: Based on user information, obtain the user's unique identifier, user id, from the user center;
[0054] Step 1.3: Add a Java agent gateway proxy. Through APM, the specific processing logic and transparent structure of different protocols are processed. By intercepting the application's runtime state, the user ID is put into the structure where the traceid is located, so that the user ID and traceid can be transparently transmitted in all microservices.
[0055] It should be noted that the pass-through structure is not consistent with different protocols. It is necessary to follow the specific protocol and implement it in a customized manner. The Java agent in this embodiment can achieve this function.
[0056] Step 2: Standardize the log printing format.
[0057] Step 2.1: Encapsulate the log printing method and store it in the corresponding microservice. Add a Java agent to each microservice. Intercept the call request through the Java agent and obtain the APM pass-through message body to obtain the traceid and user id.
[0058] Step 2.2: Structure the traceid and user id into a fixed format called floid, such as floid:userid-traceid;
[0059] Step 2.3: Design flowid as a mandatory output item for log printing to ensure that flowid is printed in every log print.
[0060] Step 3: Standardize the gateway interface response and log printing.
[0061] Step 3.1: Design the gateway interface response and assign a service code to each service;
[0062] Step 3.2: Assign a status code to the different states of each service response;
[0063] Step 3.3: After the gateway requests the downstream system, it assigns different service codes and status codes to the responses of each downstream system according to the different business and responses.
[0064] Step 3.4: Print logs for each gateway interface request. The logs include the service code, status code, flow ID, etc.
[0065] Step 4: Collect the core business status and caller relationships, and save them to a specific database for use in analyzing logs generated along the call chain;
[0066] Step 4.1: Monitor all logs of the gateway service and parse the logs asynchronously through the log visualization system;
[0067] Step 4.2: Monitor network requests and responses through a Java agent, obtain core business statuses such as status codes, business codes, and flowids, as well as caller relationships, and maintain a minimal information set for rapid risk warning and retrieval. At the same time, save the additional information to a separate dataset for detailed information retrieval.
[0068] Step 4.3: Analyze the data based on the dataset maintained in Step 4.2, and perform different asynchronous processing for different states. For abnormal states, immediately retrieve the flow ID from the log information through the message queue mechanism, and then obtain the user ID and trace ID. Using the user ID and trace ID, retrieve the relevant logs from all systems along the business chain, aggregate them, and store them in the log visualization system for testers to visualize and query, improving query efficiency. For normal processes, only the business status code and flow ID are saved, and the logs are not saved. When a query is needed, it is then performed on each system.
[0069] Step 5: Write the call trace ID back to the UI page's API response. Testers can then retrieve the call trace ID by querying the flowid field in the API response and use the log aggregation query system to locate user-level errors.
[0070] Step 5.1: Add a Java agent to the gateway service to intercept interface requests and obtain the request header, request content, and request response.
[0071] Step 5.2: Write the flow ID generated on the server side back into the HTTP request response header;
[0072] Step 5.3: Testers use the interface debugging mode to query the response header of the interface request, and then find the flow ID of all interface requests involved in the business call.
[0073] Step 6: Implement various information queries based on flowid;
[0074] Based on the user information input by testers, the system retrieves the associated user ID, then the associated flow ID, and finally obtains all call chain information and log information. Through this process, in abnormal states, testers can generally query in real-time the interfaces, call chains, interface statuses, and error logs related to their interface operations; in normal states, testers can actively and manually query the business process to obtain the interfaces, call chains, statuses, and log information involved.
[0075] The embodiments described above merely illustrate specific implementation methods of this application, and while the descriptions are detailed and specific, they should not be construed as limiting the scope of protection of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the technical solution of this application, and these modifications and improvements all fall within the scope of protection of this application.
Claims
1. A method for troubleshooting user-level business issues based on business logs, characterized in that, Includes the following steps: Step 1: Obtain the user's unique identifier (userid) and add it to the structure containing the business call chain traceid to achieve transparent transmission of user information; Step 2: Combine userid and traceid as flowid and print it to the microservice log; Step 3: Set the service code and status code for the gateway interface response and print them in the gateway interface request log along with the flowid; Step 4: Collect the core business status and caller relationships, and save them to a specific database for use in analyzing logs generated along the call chain; Step 5: Write the call traceid back to the UI page's API response. Testers can then query the flowid field of the API response and use the log aggregation query system to locate user-level errors. Step 4 includes: Step 4.1: The log visualization system monitors all logs of the gateway service and parses the logs asynchronously; Step 4.2: Monitor the gateway service's real-time request traffic and logs to obtain requests and responses, status codes, business codes, and flow IDs, maintaining a minimal information set for rapid risk warning and retrieval. Simultaneously, save this additional information to a separate dataset for detailed information retrieval. Step 4.3: Analyze the data based on the dataset saved in Step 4.2, and perform different asynchronous processing on different states; The asynchronous processing for different states includes: for abnormal states, immediately obtaining the flowid from the log information through a message queue mechanism, then obtaining the userid and traceid, and using the userid and traceid to obtain the relevant logs of all systems in the business chain for this business, aggregating them together, and storing them in a log visualization system for testers to visualize and query, thus improving query efficiency; for normal processes, only the business status code and flowid are saved, and the logs are not saved. When a query is needed, the query is performed on each system.
2. The method for troubleshooting user-level business issues based on business logs according to claim 1, characterized in that, Step 1 includes: Step 1.1: In the API gateway of the microservice system, the request information is intercepted through the encoding mechanism of the Java agent and the traceid is obtained by combining the programmable method provided by APM. Step 1.2: Based on user information, obtain the user's unique identifier, userid, from the user center; Step 1.3: Add a Java agent for the gateway service. Through APM, handle the specific processing logic and transparent structure of different protocols. By intercepting the application's runtime state, put the userid into the structure where the traceid is located, so that the userid and traceid can be transparently transmitted in all microservices.
3. The method for troubleshooting user-level business issues based on business logs according to claim 1, characterized in that, Step 2 includes: Step 2.1: Encapsulate the log printing method and store it in the corresponding microservice. Add a Java agent to each microservice. Intercept the call request through the Java agent and obtain the APM pass-through message body to obtain the traceid and userid. Step 2.2: Structure the traceid and userid into a fixed flowid paradigm; Step 2.3: Design flowid as a mandatory output item for log printing to ensure that flowid is printed in every log print.
4. The method for troubleshooting user-level business issues based on business logs according to claim 1, characterized in that, Step 3 includes: Step 3.1: Design the gateway interface response and assign a service code to each service; Step 3.2: Assign a status code for each service's different response status; Step 3.3: After the gateway sends a request to the downstream system, it assigns different service codes and status codes to the responses of each downstream system based on the different services and response statuses. Step 3.4: Print logs for each gateway interface request. The logs include the service code, status code, flowid, and interface processing information.
5. The method for troubleshooting user-level business issues based on business logs according to claim 1, characterized in that, Step 5 includes: Step 5.1: Add a Java agent to the gateway service to intercept interface requests and obtain the request header, request content, and request response. Step 5.2: Write the flowid generated on the server side back into the HTTP request response header; Step 5.3: Testers use the interface debugging mode to query the response header of the interface request, and then find the flowid of all interface requests involved in the business call.
6. The method for troubleshooting user-level business issues based on business logs according to claim 1, characterized in that, It also includes various information queries based on flowid; Based on the user information input by the testers, the associated userid is queried through the user information, the associated flowid is queried through the userid, and then all call chain information and log information are obtained.
Citation Information
Patent Citations
OpenFlow message tracking and filtering method in software defined network
CN104767720A
Asynchronous micro-service call link tracking method, device, medium and electronic equipment
CN110445643A