A distributed tracing local sampling method and system based on the sidecar model

By introducing a local sampling method of the sidecar model in the distributed tracking system, the high resource occupation problem caused by tail sampling is solved, and effective storage of errors and slow request data and resource conservation are achieved.

CN115766723BActive Publication Date: 2025-06-06HANGZHOU HARMONYCLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211465452.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-22
Publication Date
2025-06-06
Estimated Expiration
2042-11-22

AI Technical Summary

Technical Problem

When existing distributed tracking systems face high access traffic, tail sampling methods lead to excessive network bandwidth and memory resource usage and are unbearable.

Method used

The distributed local sampling method based on the sidecar model is adopted. By deploying a sidecar container in the node, the trace requests are classified, sorted and exception judgments are made. Only exception and slow trace request data are retained, and some span nodes are discarded for normal execution.

Benefits of technology

Effectively reduces network bandwidth and memory usage, saves costs, and ensures the storage of tracking data for errors and slow requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766723B_ABST
    Figure CN115766723B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed tracing local sampling method based on the sidecar model, including: deploying a sidecar container in a node; performing bytecode enhancement on a process in the node to obtain a bytecode-enhanced process; intercepting a trace request sent by the bytecode-enhanced process, and sending the trace request to the sidecar container; classifying the trace request in the sidecar container to obtain several trace request types; obtaining access information of trace requests of each trace request type, and determining whether the trace request is abnormal based on the access information; if the trace request is abnormal, performing link analysis on the trace request. The present invention also discloses a distributed tracing local sampling system based on the sidecar model. The present invention occupies less network bandwidth and memory, etc., and can save costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing technology, and in particular to a distributed tracing local sampling method and system based on a sidecar model. Background Art

[0002] In order to reduce the data processing and storage costs of distributed tracing, distributed tracing is generally sampled. There are two commonly used sampling methods for distributed tracing: head sampling and sampling rate adjustment based on head sampling, and tail sampling.

[0003] The header sampling method and the dynamic adjustment sampling method are the easiest to implement and are currently the most commonly used methods for distributed tracing. They decide whether to record the entire distributed trace at the beginning of the request access and pass it to subsequent nodes through the trace request context. Subsequent nodes decide whether to retain the span data generated by the node in this request based on the adopted identifier. Although header sampling and the dynamic adjustment method based on header sampling are easy to implement, in the face of production accidents, which are rare in the production environment, 99.99% of the trace requests retained by sampling are normal trace requests, while trace requests that developers are interested in, that is, trace requests for production accidents, are difficult to save.

[0004] Tail sampling can effectively avoid the problem of head sampling, because the method of tail sampling is to send all trace requests first, and then determine whether the trace request is related to the production accident when the request ends, and decide whether it should be sent out. Because the decision is delayed, it can ensure that all problems of production accidents are completely retained. However, tail sampling will also bring many other problems, mainly in two aspects:

[0005] 1. In terms of network bandwidth, since all tail sampling requests need to be sent to the tail sampling node in a centralized manner, the problem is that the trace request data will occupy the entire data center network bandwidth. Assuming that the trace request data of a node is 2K bytes, the average data center has 100,000 entry request accesses per second, and each request link has an average of 10 layers. Then the data center will generate millions of accesses per second, which is only the level of a medium-sized Internet company. If all trace request data is sent to one or several points, the network bandwidth borne by each point will reach the G-byte level. Although 10 Gigabit networks have been popularized in data centers, the core network of 10 Gigabit networks is used to carry trace request data, which is bound to affect the normal network requests of its original core business. Trace request data and core business will form fierce competition on network bandwidth, which is not what the distributed tracing system wants to see.

[0006] 2. In terms of memory, due to the large amount of requested data, tail sampling requires that all data be retained in the memory first, and then a decision can be made on whether to retain the trace request data. This means that if 30 seconds of data is cached, the amount of memory on each node will reach the T level. If multiple nodes are used to share the workload, each node will require hundreds of GB of memory. This is too expensive and unaffordable for many companies. Summary of the invention

[0007] The object of the present invention is to provide a distributed tracing local sampling method and system based on a sidecar model that occupies less network bandwidth and memory.

[0008] In order to solve the above technical problems, the present invention provides a distributed tracing local sampling method based on the sidecar model, comprising the following steps:

[0009] Deploy the sidecar container in the node;

[0010] Perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process;

[0011] Intercept the trace request sent by the bytecode-enhanced process and send the trace request to the sidecar container; the trace request includes the URL and access information;

[0012] According to the URL of the trace request in the sidecar container, the trace request in the sidecar container is classified to obtain several trace request types;

[0013] Sort the trace requests in each trace request type according to the access information of the trace requests;

[0014] According to the sorting results and the access information of the trace request, determine whether the trace request in each trace request type is abnormal;

[0015] If the trace request is abnormal, perform link analysis on the trace request.

[0016] Preferably, according to the URL of the trace request in the sidecar container, the trace request in the sidecar container is classified to obtain several trace request types, which specifically includes the following steps:

[0017] Decompose the URL of the race request in the sidecar container into several URL components;

[0018] Use Gibberish-Detector to identify the gibberish in the URL component, replace the gibberish with *, and get the corrected URL component;

[0019] Obtain the converged URL according to the corrected URL components;

[0020] According to the converged URL, the trace requests in the sidecar container are classified to obtain several trace request types.

[0021] Preferably, the access information includes a return code and delay information.

[0022] Preferably, judging whether a trace request in each trace request type is abnormal according to the sorting result and the access information of the trace request specifically includes the following steps:

[0023] According to the return code of the access information of the trace request, determine whether the trace request is correct;

[0024] If the trace request is wrong, the trace request is abnormal;

[0025] If the trace request is correct, determine whether the trace request is too slow based on the access information delay information and sorting results;

[0026] If the trace request is too slow, the trace request is abnormal.

[0027] Preferably, link analysis is performed on the trace request, specifically including the following steps:

[0028] Send the trace request to the distributed tracing system for tracing.

[0029] Preferably, performing bytecode enhancement on the process in the node to obtain the bytecode enhanced process specifically includes the following steps:

[0030] Use the instrument debugging tool to perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process.

[0031] Preferably, the trace request sent by the bytecode-enhanced process is intercepted and the trace request is sent to the sidecar container; specifically, the following steps are included:

[0032] The trace request sent by the bytecode-enhanced process is intercepted through ASM technology, and the trace request is sent to the sidecar container.

[0033] Preferably, the URL components include schema, authority, path, query and fragment.

[0034] The present invention also provides a distributed tracing local sampling system based on the sidecar model, including:

[0035] Deployment module, used to deploy sidecar containers in nodes;

[0036] An enhancement module is used to perform bytecode enhancement on the process in the node to obtain the process after bytecode enhancement;

[0037] An interception module is used to intercept trace requests sent by the bytecode-enhanced process and send the trace requests to the sidecar container; the trace requests include a URL and access information;

[0038] The classification module is used to classify the trace requests in the sidecar container according to the URL of the trace request in the sidecar container to obtain several trace request types;

[0039] A sorting module is used to sort the trace requests in each trace request type according to the access information of the trace requests to obtain a sorting result;

[0040] A judgment module is used to judge whether the trace request in each trace request type is abnormal according to the sorting result and the access information of the trace request;

[0041] The tracing module is used to perform link analysis on trace requests when the race request is abnormal.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] The present invention proposes a local sampling method based on the sidecar model. It still follows the idea of ​​tail sampling and tries to save wrong and slow trace requests. However, for some span nodes that execute normally, these data can be discarded because they do not affect user use. The present invention occupies less network bandwidth and memory, saving costs; the present invention can solve the problem of high resource occupancy of tail sampling when the access traffic is high. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.

[0045] Figure 1 It is a flow chart of a distributed tracing local sampling method based on the sidecar model of the present invention;

[0046] Figure 2 is a diagram of tracing requests in a node. DETAILED DESCRIPTION

[0047] Many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present invention, so the present invention is not limited to the specific implementation disclosed below.

[0048] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0049] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0050] The following is combined with Figure 1-2 The present invention is further described in detail:

[0051] like Figure 1 As shown, the present invention provides a distributed tracing local sampling method based on the sidecar model, comprising the following steps:

[0052] Deploy the sidecar container in the node;

[0053] Perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process;

[0054] Intercept the trace request sent by the bytecode-enhanced process and send the trace request to the sidecar container; the trace request includes the URL and access information;

[0055] According to the URL of the trace request in the sidecar container, the trace request in the sidecar container is classified to obtain several trace request types;

[0056] Sort the trace requests in each trace request type according to the access information of the trace requests;

[0057] According to the sorting results and the access information of the trace request, determine whether the trace request in each trace request type is abnormal;

[0058] If the trace request is abnormal, perform link analysis on the trace request.

[0059] A preferred embodiment classifies the trace requests in the sidecar container according to the URL of the trace request in the sidecar container to obtain several trace request types, specifically including the following steps:

[0060] Decompose the URL of the race request in the sidecar container into several URL components;

[0061] Use Gibberish-Detector to identify the gibberish in the URL component, replace the gibberish with *, and get the corrected URL component;

[0062] Obtain the converged URL according to the corrected URL components;

[0063] According to the converged URL, the trace requests in the sidecar container are classified to obtain several trace request types.

[0064] In a preferred embodiment, the access information includes a return code and delay information.

[0065] A preferred embodiment, according to the sorting result and the access information of the trace request, determines whether the trace request in each trace request type is abnormal, specifically comprising the following steps:

[0066] According to the return code of the access information of the trace request, determine whether the trace request is correct;

[0067] If the trace request is wrong, the trace request is abnormal;

[0068] If the trace request is correct, determine whether the trace request is too slow based on the access information delay information and sorting results;

[0069] If the trace request is too slow, the trace request is abnormal.

[0070] In a preferred embodiment, link analysis is performed on the trace request, specifically comprising the following steps:

[0071] Send the trace request to the distributed tracing system for tracing.

[0072] In a preferred embodiment, bytecode enhancement is performed on a process in a node to obtain a bytecode-enhanced process, which specifically includes the following steps:

[0073] Use the instrument debugging tool to perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process.

[0074] A preferred embodiment intercepts the trace request sent by the bytecode-enhanced process and sends the trace request to the sidecar container; specifically includes the following steps:

[0075] The trace request sent by the bytecode-enhanced process is intercepted through ASM technology, and the trace request is sent to the sidecar container.

[0076] In a preferred embodiment, the URL components include schema, authority, path, query and fragment.

[0077] In the present invention, the purpose of the distributed tracing system is to help developers find the actual execution status of nodes with problematic requests. In order to save trace request data with problematic requests as much as possible, tail sampling should be the first choice. However, the resource overhead of tail sampling under large-scale traffic is unbearable for enterprises. Therefore, in order to solve the problem of high resource occupancy of tail sampling when the access traffic is high, the present invention proposes a local sampling method based on the sidecar model. It still follows the idea of ​​tail sampling and tries to save wrong and slow trace requests. However, for the situation where some span nodes execute normally, these data can be discarded because they will not affect user use.

[0078] The present invention also provides a distributed tracing local sampling system based on the sidecar model, including:

[0079] Deployment module, used to deploy sidecar containers in nodes;

[0080] An enhancement module is used to perform bytecode enhancement on the process in the node to obtain the process after bytecode enhancement;

[0081] An interception module is used to intercept trace requests sent by the bytecode-enhanced process and send the trace requests to the sidecar container; the trace requests include a URL and access information;

[0082] The classification module is used to classify the trace requests in the sidecar container according to the URL of the trace request in the sidecar container to obtain several trace request types;

[0083] A sorting module is used to sort the trace requests in each trace request type according to the access information of the trace requests to obtain a sorting result;

[0084] A judgment module is used to judge whether the trace request in each trace request type is abnormal according to the sorting result and the access information of the trace request;

[0085] The tracing module is used to perform link analysis on trace requests when the race request is abnormal.

[0086] In order to better illustrate the technical effect of the present invention, the present invention provides the following specific example to illustrate the above technical process. The distributed tracing local sampling method based on the sidecar model is as follows:

[0087] 1. Deploy the sidecar container on each node;

[0088] 2. For the monitored processes that need to be instrumented, perform bytecode-level enhancement;

[0089] When a process sends a trace request, the trace request is intercepted by ASM technology and sent to the sidecar container.

[0090] 3. Dynamically adjust the distributed tracing program such as skywalking to achieve 100% header sampling, that is, send out all trace request data;

[0091] 4. Analyze the URLs in the trace request and converge them into a limited number of types. Otherwise, too many types of URLs will cause too much memory data in the sidecar, resulting in excessive resource usage.

[0092] 4.1 Decompose the URL into some common URL components: schema, authority, path, query and fragment, which are separated by :, / , ? and #;

[0093] For example: http: / / example.com / books / search?name=go&isbn=1234

[0094] Will be decomposed into: schema:http

[0095] authority:example.com

[0096] path:{"path0":"books","path1":"search"}

[0097] query:{"name":"go","isbn":"1234"}

[0098] 4.2 Use Gibberish-Detector (using markov chain for pattern recognition) to identify non-human natural language in URL as gibberish;

[0099] 4.3 All the nonsense words are replaced with *. http: / / example.com / books / search?name=go&isbn=1234 will be converged to http: / / example.com / books / search?name=go&isbn=*

[0100] 5. Sidecar intercepts all trace request data and counts the access volume, return code, delay information, error rate, and other data for each process URL type in Sidecar according to process ID and URL type;

[0101] 6. For the data that is not slow or good, after calculating the number of visits, delay information, and errors, cache it in the cache for 10 seconds and then discard it;

[0102] 6.1 The wrong data is determined based on the HTTP return code;

[0103] For slow data, the longest time in the delay information of a request needs to be evenly divided into 10 buckets according to the delay information. The number of requests in each bucket needs to be calculated. According to the number of data in each bucket, the bucket where the 95% position is located can be calculated. The largest bucket is removed. When the number exceeds 5%, the value of that bucket is the 95 line. Then, the interpolation method is used inside the bucket, and a relatively accurate P95 value can also be calculated by averaging within the bucket.

[0104] 6.2 When another sidecar notifies this sidecar that a trace request ID needs to be retained, the span is also sent out

[0105] 7. For slow and erroneous data, send it to the original distributed tracking system background analysis system, such as skywalking, which is sent to the OAP program

[0106] 8. Use the link analysis mechanism of the original distributed tracking system to display erroneous and slow requests.

[0107] The present invention can directly discard the Span trace request data that is neither slow nor fast. This is because data center microservice scenarios are mainly based on rpc calls, and rpc calls are synchronous calls, and there will be no asynchronous requests. This determines that when a node is abnormal, its call link will completely save the node with the problem and its upstream, and the span of the downstream node will be discarded because it does not appear. In order to ensure that when checking the node span data when a problem occurs, there will be no ambiguity due to the lack of normal downstream span data, the trace request data with the problem should contain the span data of its downstream call, such as Figure 2 shown.

[0108] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules, modules or units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units, modules or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0109] The units may or may not be physically separated, and the components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0110] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0111] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by the central processing unit (CPU), the above-mentioned functions defined in the method of the present invention are executed. It should be noted that the above-mentioned computer-readable medium of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, a system, device or device of an electrical, magnetic, optical, electromagnetic, infrared segment, or semiconductor, or any combination of the above.

[0112] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0113] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A distributed tracing local sampling method based on the sidecar model. It is characterized in that The following steps are involved: Deploy the sidecar container in the node; Perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process; Intercept the trace request sent by the bytecode-enhanced process and send the trace request to the sidecar container; the trace request includes the URL and access information; According to the URL of the trace request in the sidecar container, the trace request in the sidecar container is classified to obtain several trace request types; According to the access information of the trace request, the trace requests in each trace request type are sorted to obtain a sorting result; According to the sorting results and the access information of the trace request, determine whether the trace request in each trace request type is abnormal; If the trace request is abnormal, perform link analysis on the trace request.

2. According to the sidecar model-based distributed tracing local sampling method of claim 1, It is characterized in that According to the URL of the trace request in the sidecar container, the trace request in the sidecar container is classified to obtain several trace request types, which specifically includes the following steps: Decompose the URL of the race request in the sidecar container into several URL components; Use Gibberish-Detector to identify the gibberish in the URL component, replace the gibberish with *, and get the corrected URL component; Obtain the converged URL according to the corrected URL components; According to the converged URL, the trace requests in the sidecar container are classified to obtain several trace request types.

3. According to the distributed tracing local sampling method based on the sidecar model in claim 1, Features: The access information includes a return code and delay information.

4. According to the sidecar model-based distributed tracing local sampling method of claim 3, It is characterized in that According to the sorting results and the access information of the trace request, it is determined whether the trace request in each trace request type is abnormal, which specifically includes the following steps: According to the return code of the access information of the trace request, determine whether the trace request is correct; If the trace request is wrong, the trace request is abnormal; If the trace request is correct, determine whether the trace request is too slow based on the access information delay information and sorting results; If the trace request is too slow, the trace request is abnormal; If the trace request is not too slow, the trace request is normal.

5. According to the sidecar model-based distributed tracing local sampling method of claim 1, It is characterized in that Perform link analysis on the trace request, including the following steps: Send the trace request to the distributed tracing system for tracing.

6. According to the sidecar model-based distributed tracing local sampling method of claim 1, It is characterized in that Perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process, which specifically includes the following steps: Use the instrument debugging tool to perform bytecode enhancement on the process in the node to obtain the bytecode enhanced process.

7. According to the sidecar model-based distributed tracing local sampling method of claim 1, It is characterized in that Intercept the trace request sent by the bytecode-enhanced process and send the trace request to the sidecar container; specifically, the following steps are included: The trace request sent by the bytecode-enhanced process is intercepted through ASM technology, and the trace request is sent to the sidecar container.

8. According to the sidecar model-based distributed tracing local sampling method of claim 1, Features: The URL components include schema, authority, path, query and fragment.

9. A distributed tracing local sampling system based on a sidecar model that implements the distributed tracing local sampling method based on a sidecar model as described in any one of claims 1 to 8, It is characterized in that include: Deployment module, used to deploy sidecar containers in nodes; An enhancement module is used to perform bytecode enhancement on the process in the node to obtain the process after bytecode enhancement; An interception module is used to intercept trace requests sent by the bytecode-enhanced process and send the trace requests to the sidecar container; the trace requests include a URL and access information; The classification module is used to classify the trace requests in the sidecar container according to the URL of the trace request in the sidecar container to obtain several trace request types; A sorting module is used to sort the trace requests in each trace request type according to the access information of the trace requests to obtain a sorting result; A judgment module is used to judge whether the trace request in each trace request type is abnormal according to the sorting result and the access information of the trace request; The tracing module is used to perform link analysis on trace requests when the trace request is abnormal.

Citation Information

Patent Citations

  • Internet-based multifunctional intelligent monitoring method, device and server

    CN111782463A

  • Method, system and device for tracking call chain in micro-service environment and storage medium

    CN111984346A