A Tail Sampling Method for a Distributed Tracing System in an RPA Workflow

Through the variational graph automatic encoder model and clustering method, feature extraction and abnormal detection of the tracking data of the RPA system is solved, and the problems of high storage costs of massive tracking data and omissions of abnormal tracking are realized, efficient tail sampling is achieved, and the diversity and observability of sampling results are improved.

CN119902474BActive Publication Date: 2025-07-22安徽思高智能科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510360586.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The massive amount of tracking data generated by distributed tracking systems in existing RPA systems leads to high storage costs and abnormal tracking in sampling results is easily missed. The existing methods are insufficiently utilized for feature information and are insensitive to abnormal tracking detection.

Method used

Unsupervised characterization learning is performed using the variational graph automatic encoder model, and the abnormal tracking of loss detection is carried out through reconstruction and partial sampling is performed in combination with the clustering method, important tracking data is retained and redundancy is reduced.

Benefits of technology

Effectively reduce storage costs, improve the diversity and observability of sampling results, ensure the sensitivity of abnormal tracking, and improve sampling quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902474B_ABST
    Figure CN119902474B_ABST
Patent Text Reader

Abstract

A tail sampling method for a distributed tracking system in an RPA workflow disclosed by the present invention relates to the field of RPA workflows and includes: performing unsupervised representation learning on tracking data using a variational graph autoencoder model and saving the trained model; obtaining feature vectors based on the pre-trained variational graph autoencoder model; detecting and retaining all abnormal tracks based on the reconstruction loss of the variational graph autoencoder model; classifying normal tracks using a clustering method and preferentially sampling; and combining the sampling results of abnormal tracks and normal tracks to form a final sampled data set. By using a clustering and feature preference sampling strategy to perform biased sampling on normal tracking data, the present invention reduces the storage cost of massive tracking data, improves the sensitivity of the sampler to abnormal tracks, and ensures the diversity and observability of the sampling results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of RPA workflows, and in particular, to a tail sampling method for a distributed tracing system in an RPA workflow. Background Art

[0002] With the in-depth development of enterprise digital transformation, robotic process automation (RPA) technology has been widely applied. The RPA technology executes business processes by simulating human operations such as mouse clicks, keyboard inputs, and data processing, which can significantly improve the automation level and execution efficiency of business processes. The application scenarios of RPA technology are extensive, covering multiple fields such as data entry, document processing, information extraction, customer service, and financial statement generation. It can free employees from repetitive and cumbersome work, enabling them to invest more time and energy in creative and strategic work, and promoting the innovation and development of enterprises.

[0003] In a complex RPA system, by using a distributed tracing system, operation and maintenance personnel can monitor and analyze requests in the system, and better understand and diagnose performance bottlenecks and fault points in the system. However, as the scale of the RPA system expands, the distributed tracing system will generate a large amount of tracing data, which brings high storage costs. In addition, these data may also contain a large amount of duplicate and redundant tracings. How to effectively reduce the storage overhead of tracing data has become an important research topic.

[0004] In a production environment, it is usually chosen to sample the tracing data, only retaining a small subset of all the tracing data, so as to relieve the storage pressure. However, due to the obvious long-tail distribution of the tracing data, some common types of tracings account for the vast majority of them, while some rare or abnormal tracings appear with extremely low frequencies. However, for development and operation and maintenance personnel, these rare tracings are often more valuable, which can help them better understand the edge cases in system operation and play an important role in downstream tasks such as anomaly detection and root cause location. If only randomly sampling the tracing data (i.e., head sampling), it is difficult to cover the tracings that appear with low frequencies, which may cause rare abnormal tracings in the sampling results to be omitted, while retaining a large amount of duplicate and redundant tracings, thus reducing the observability of the system.

[0005] Therefore, it is very important to design a tail sampling method for a distributed tracing system in an RPA workflow. Such a method can execute a tail sampling method for sampling decisions in real time according to characteristics such as the call structure, latency, and status code of the tracing data, retain important and valuable tracings in a targeted manner, and reduce the number of redundant tracings, thereby optimizing the collection quality of the tracing data.

[0006] Although existing methods can reduce the number of redundant traces in the sampling results and improve the diversity of tracing to a certain extent, the characterization of tracing is usually inaccurate. They only consider call relationships and ignore attributes such as latency and status codes, resulting in the omission of the utilization of feature information of tracing. In addition, abnormal traces can collect request information that produces incorrect results or has performance bottlenecks, which is crucial for operation and maintenance personnel to perform root cause location and anomaly detection. However, existing methods are not sensitive to abnormal tracing, and abnormal traces are often omitted. Summary of the Invention

[0007] To solve the problems of large quantity of workflow logs, high storage cost, and low quality in the RPA system, the present invention proposes a tail sampling method for a distributed tracing system in an RPA workflow.

[0008] The present invention provides a tail sampling method for a distributed tracing system in an RPA workflow, mainly including:

[0009] S1: Collect the traces generated by the distributed tracing system, perform unsupervised representation learning on the trace data using a variational graph autoencoder model, and save the pre-trained variational graph autoencoder model; wherein, each piece of data represents the execution situation of a request in the RPA workflow system and consists of node information, parent node information, and operation execution time.

[0010] S2: Encode the trace data into graph structure data in the form of an adjacency matrix, and obtain feature vectors based on the encoder of the variational graph autoencoder model.

[0011] S3: Based on the decoder of the variational graph autoencoder model, obtain the reconstruction loss, perform anomaly detection on the trace data, and sample all abnormal traces.

[0012] S4: Use a clustering method to classify normal traces and perform preferential sampling.

[0013] S5: Combine the sampling results of abnormal traces and normal traces to form a final sampling data set.

[0014] A storage device stores instructions and data for implementing a tail sampling method for a distributed tracing system in an RPA workflow.

[0015] A tail sampling device for a distributed tracing system in an RPA workflow includes: a processor and a storage device; the processor loads and executes the instructions and data in the storage device for implementing a tail sampling method for a distributed tracing system in an RPA workflow.

[0016] The beneficial effects brought by the technical solution provided by the present invention are as follows: Through the representation learning based on the variational graph autoencoder model, the present invention extracts the features of the tracking data, uses the reconstruction loss for anomaly detection to ensure that all abnormal tracks are retained, and performs biased sampling on the normal tracking data through the clustering and feature preference sampling strategy. While reducing the storage cost of massive tracking data, the sensitivity of the sampler to abnormal tracking is improved, and the diversity and observability of the sampling results are ensured. Description of the Drawings

[0017] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0018] Figure 1 is a flowchart of a tail sampling method for a distributed tracking system in an RPA workflow according to an embodiment of the present invention;

[0019] Figure 2 is a structural diagram of a variational graph autoencoder model according to an embodiment of the present invention;

[0020] Figure 3 is a schematic diagram of the operation of a hardware device according to an embodiment of the present invention. Detailed Embodiments

[0021] For a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the drawings.

[0022] As Figure 1 shown, the present invention proposes a tail sampling method for a distributed tracking system in an RPA workflow, including:

[0023] S1: Collect the traces generated by the distributed tracking system, perform unsupervised representation learning on the tracking data using the variational graph autoencoder model, and save the pre-trained variational graph autoencoder model; wherein, each piece of data represents the execution situation of a request in the RPA workflow system and consists of node information, parent node information, and operation execution time; specifically:

[0024] S1.1: Continuously collect the tracking data generated by the distributed tracking system, including service call relationships, delays, and status codes, and use the tracking in batches for training the variational graph autoencoder model.

[0025] S1.2: Encode the tracking data into a graph data structure in the form of an adjacency matrix:

[0026]

[0027] where n represents the number of Spans in the trace, f is the number of attribute types in each Span, and the adjacency matrix represents the call relationship of the trace, is the node attribute matrix for tracking, which records the attribute features of each Span other than tracking (such as service name, interface name, etc.). records the latency corresponding to each Span. A Span represents a node that makes up the tracking data.

[0028] S1.3: Perform unsupervised representation learning on the tracking data based on the variational graph autoencoder model GVAE. The model structure of GVAE is as Figure 2 shown, including two parts: the structural VAE and the latency VAE. GVAE uses the graph structure in and in the structural VAE to learn the structural features of the tracking, encodes them through the GNN encoder to obtain the feature vector , and then reconstructs the reconstructed output of the tracking structure through the GNN decoder , , . In the latency VAE, both the structural features and the time features are input into the neural network, but the model weights of the structural VAE do not participate in gradient descent.

[0029] GVAE uses the variational lower bound (ELBO) as the loss function:

[0030]

[0031] where, is the reconstruction loss of the model, which evaluates the similarity between the reconstructed output and the original input through log-likelihood. The model expects the reconstructed output to be as close as possible to the original input; is the KL divergence, which is used to measure the difference between the approximate posterior distribution and the true posterior distribution . The model expects the approximate distribution to be as similar as possible to the original distribution. represents the expectation of the approximate posterior distribution , represents the log-likelihood of reconstructing the data given the feature vector z and the model parameters theta, which is used to measure the ability of the model to reconstruct data from the latent variables, represents the prior probability distribution of the feature vector z.

[0032] S1.4: Re-train the variational graph autoencoder model every time a certain amount of tracking is collected, and save the model parameters. S2: Encode the tracking data into graph structure data in the form of an adjacency matrix, and obtain the feature vector based on the encoder of the pre-trained variational graph autoencoder model. Specifically:

[0033] S2.1: Cache a certain number of tracking data, construct its graph structure, input it into the encoder of the variational graph autoencoder model for encoding, and extract the feature vectors:

[0034]

[0035] Among them, represents the feature vector, and are the hidden layer vectors of the structural VAE and the delay VAE respectively.

[0036] S3: Based on the reconstruction loss of the variational graph autoencoder model, perform anomaly detection on the tracking and sample all abnormal tracks. Specifically:

[0037] S3.1: Reconstruct the graph structure through the decoder of the variational autoencoder model to obtain the negative log-likelihood value of the reconstructed output relative to the original input ( ), that is, the reconstruction loss:

[0038]

[0039] Among them, and are sampled from the hidden layer vectors in the GVAE model and , is the graph structure, is the number of Spans in the tracking, represents three matrices in the graph structure G, represents the probability distribution learned by the encoder, represents the probability distribution learned by the decoder, and are sampled from and , and L represents the number of samples. reflects the reconstruction error of the model. For a small number of tracks with structural or delay anomalies, since they do not conform to the data distribution learned by the model, the reconstructed output usually has a large difference from the original track, resulting in a high . Therefore, can be used as an evaluation index for anomaly detection.

[0040] S3.2: Compare the output by the model with a pre-set anomaly threshold. If is lower than the threshold, the track is considered a normal track; otherwise, the track is considered an abnormal track and is retained.

[0041] S4: Use the clustering method to classify the normal tracks and perform preference sampling. Specifically:

[0042] S4.1: For the normal traces screened in step S3, use the K-means clustering algorithm to classify the traces, and sort the different classes from small to large according to the data of the traces in the classes.

[0043] S4.2: Use the max-min fairness algorithm to sequentially allocate the remaining sampling budget to different classes (a part of the sampling budget is also used in model training, detecting normal traces and abnormal traces). When the number of classes is less than the remaining average budget, sample all the traces of the classes; otherwise, evenly distribute the remaining sampling budget to all the remaining classes, and randomly sample within each class.

[0044] Preference sampling based on the max-min fairness algorithm can balance the number of sampled trace data of different classes, and improve the diversity and observability of the sampling results.

[0045] S5: Combine the sampling results of abnormal traces and normal traces to form the final sampling data set. Merge all the sampling results of abnormal trace data and normal trace data to form the final data set.

[0046] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the hardware device working in the embodiments of the present invention. The hardware device specifically includes: a tail sampling device 401 for a distributed tracing system in an RPA workflow, a processor 402, and a storage device 403.

[0047] A tail sampling device 401 for a distributed tracing system in an RPA workflow: The tail sampling device 401 for a distributed tracing system in an RPA workflow implements the tail sampling method for a distributed tracing system in an RPA workflow.

[0048] Processor 402: The processor 402 loads and executes the instructions and data in the storage device 403 to implement the tail sampling method for a distributed tracing system in an RPA workflow.

[0049] Storage device 403: The storage device 403 stores instructions and data; the storage device 403 is used to implement the tail sampling method for a distributed tracing system in an RPA workflow.

[0050] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A tail sampling method for a distributed tracing system in an RPA workflow, characterized in that: including: S1: Collect the traces generated by the distributed tracing system, perform unsupervised representation learning on the trace data using the variational graph autoencoder model, and save the trained variational graph autoencoder model; each piece of data represents the execution situation of a request in the RPA workflow system and consists of node information, parent node information, and operation execution time. S2: Encode the trace data into graph structure data in the form of an adjacency matrix, and obtain the feature vector based on the encoder of the variational graph autoencoder model. S3: Based on the decoder of the variational graph autoencoder model, obtain the reconstruction loss, perform anomaly detection on the trace data, and sample all abnormal traces; specifically: S3.1: Reconstruct the graph structure through the decoder of the variational autoencoder model to obtain the negative log-likelihood value of the reconstruction output relative to the original input, that is, the reconstruction loss. Among them, and are the hidden layer vectors of the structural VAE and the delayed VAE, is the graph structure, is the number of Spans in the tracking. A Span represents the nodes that make up the tracking data, represents three matrices in the graph structure G, represents the probability distribution learned by the encoder, represents the probability distribution learned by the decoder, and are sampled from and ; represents the number of samples;​ S3.2: Compare the negative log-likelihood value with a pre-set anomaly threshold. If the negative log-likelihood value is lower than the anomaly threshold, the trace is considered a normal trace; otherwise, the trace is considered an abnormal trace and is retained. S4: Use the clustering method to classify the normal traces and perform preferential sampling. S5: Combine the sampling results of abnormal traces and normal traces to form the final sampling data set.

2. The tail sampling method for the distributed tracking system in the RPA workflow according to claim 1, wherein: The specific steps of S1 are as follows: S1.1: Continuously collect the trace data generated by the distributed tracing system, including service call relationships, latencies, and status codes, and use the trace data in batches for training the variational graph autoencoder model. S1.2: Encode the trace data into a graph data structure in the form of an adjacency matrix. Among them, n represents the number of Spans in the trace, Span represents the nodes that make up the trace data, f is the number of types of attributes in each Span, represents the call relationship of the trace, represents the node attribute matrix of the trace, represents the latency corresponding to each Span; S1.3: Perform unsupervised representation learning on the tracking data based on the variational graph autoencoder model; the variational graph autoencoder model includes a structured VAE and a delayed VAE. In the structured VAE, the graph structure in and are used to learn the structural features of the tracking, and the feature vector is obtained through encoding by the encoder, and then the reconstructed output of the tracking structure is obtained through reconstruction by the decoder; in the delayed VAE, both the structural features and the temporal features are input into the neural network; S1.4: Every time a certain amount of trace data is collected, retrain the variational graph autoencoder model and save the model parameters.

3. The tail sampling method for the distributed tracing system in the RPA workflow according to claim 2, wherein: The variational graph autoencoder model uses the variational lower bound as the loss function. Among them, is the reconstruction loss of the model, is the KL divergence, which is used to measure the difference between the approximate posterior distribution and the true posterior distribution . represents the expectation of the approximate posterior distribution . represents the log-likelihood of the reconstructed data . represents the prior probability distribution of the feature vector z.

4. The tail sampling method for a distributed tracing system in an RPA workflow according to claim 1, characterized in that: The specific steps of S2 are as follows: Cache the trace data, construct its graph structure, input it into the encoder of the variational graph autoencoder model for encoding, and extract the feature vector. Among them, represents the feature vector, and are the hidden layer vectors of the structured VAE and the delayed VAE respectively.

5. The tail sampling method for the distributed tracing system in the RPA workflow according to claim 1, wherein: The specific steps of S4 are as follows: S4.1: For the normal traces screened in step S3, use the K-means clustering algorithm to classify the traces and sort them from small to large according to the data of the traces in the category. S4.2: Use the max-min fairness algorithm to allocate the remaining sampling budget to different categories in turn; when the number of categories is less than the remaining average budget, sample all the traces in the category; otherwise, evenly distribute the remaining sampling budget to all the remaining categories and randomly sample within each category.

6. A storage device, characterized in that: The storage device stores instructions and data for implementing the tail sampling method for the distributed tracing system in the RPA workflow according to any one of claims 1 to 5.

7. A tail sampling device for a distributed tracing system in an RPA workflow, characterized in that: including: a processor and a storage device; the processor loads and executes the instructions and data in the storage device for implementing the tail sampling method for the distributed tracing system in the RPA workflow according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method, system, medium and equipment for synthesizing abnormal RPA workflow data

    CN118521275A