Activity classification method and system supporting RPA and graph neural network
By using graph-based neural network methods in robot process automation, the process heterogeneous graph is constructed and clustered, the problem of inaccurate activity classification in multi-process trajectory interleaving execution is solved, and the accurate classification of complex trajectory tasks is achieved.
Patent Information
- Application Number
- CN202510112683.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the multi-process trajectory interleaving execution scenario, traditional activity classification methods cannot effectively identify the tasks to which activities in different trajectories hidden in the user's operation behavior log, resulting in inaccurate classification.
Using a graph neural network-based method, the active nodes and relationship edges in the user's operation behavior log are represented by constructing a process heterogeneous graph, and node features are calculated using the graph neural network layer, and clustered with spectral clustering and K-mean algorithms to obtain the tasks to which each activity belongs.
The problem of poor classification effect in user operation behavior logs is effectively dealt with. Through the attention mechanism of fusion edge and node information, more accurate activity interaction representation is obtained, and the correct classification ability of complex trajectory tasks is improved.
Smart Images

Figure CN119598352B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotic automation, and in particular to a method and system for activity classification supporting RPA (robotic process automation) and based on graph neural network. Background Art
[0002] With the advancement of enterprise digital transformation, a large number of repetitive tasks in enterprise business processes have become a bottleneck for enterprise development. Robotic automation can meet this demand and reduce manpower expenses, reduce the cost of task execution and improve task execution efficiency, thereby improving enterprise competitiveness and having a significant impact on enterprise development.
[0003] Robotic Process Automation (RPA) is a process automation technology that is suitable for scenarios that handle a large number of repetitive tasks. These scenarios usually have obvious operating rules and standardized processes, such as generating reports, entering information, and processing orders. RPA can effectively replace human tasks. Currently, the implementation of RPA is generally divided into the following four steps: (1) obtaining user operation behavior logs; (2) extracting repeatable tasks from user operation behavior logs; (3) generating task codes that can be automatically executed; (4) executing the task code and monitoring it.
[0004] Activity classification refers to classifying the various activities recorded in the user operation behavior log according to the tasks to which they belong. Among them, sharing activities with high similarity and decision-making activities with low similarity are divided into the same task, which is conducive to forming the correct trajectory. Here, task refers to a series of operations performed to achieve a certain purpose (for example, transcribing a series of user information recorded in an Excel table to a Web page), which can be represented by a process model; activity refers to a specific operation performed by the user on the application interface (for example, clicking a button, copying a piece of data, pasting a certain content, etc.), and different activities may belong to the same event (for example, click, copy, paste, etc.); and trajectory refers to an operation path formed by multiple activities in chronological order (for example, click → paste → copy → click → paste, etc.).
[0005] In practical applications, a user operation behavior log often contains multiple tracks that are executed in an interlaced manner, such as the same event appearing in different tracks; the appearance of decision points may also lead to the inability to correctly identify complex single tracks. This situation makes it impossible for traditional process mining algorithms to effectively mine clear and correct executable process models. Therefore, one of the core aspects of robotic process automation is how to correctly identify the tasks to which activities in different tracks hidden in user operation behavior logs belong, so as to help obtain the correct track later.
[0006] At present, existing activity classification research at home and abroad mainly focuses on methods such as discovering frequent patterns and constructing direct follow-up graphs. The above methods are only suitable for simple task scenarios, but still have certain limitations for scenarios where multiple process trajectories are interleaved. Summary of the invention
[0007] Aiming at the problem of inaccurate activity classification caused by the interlaced execution of multiple process trajectories, the present invention proposes an activity classification method and system that supports robotic process automation and is based on graph neural network.
[0008] A first aspect of the present invention provides an activity classification method supporting RPA and based on a graph neural network, comprising the following steps:
[0009] Collect the interaction logs between users and applications in the case of multi-trace interleaved execution;
[0010] A heterogeneous process graph is constructed based on the interaction log, where each node of the graph represents a different activity of the user and the edge represents the relationship between two activities;
[0011] Add node labels to the heterogeneous process graph. The node labels represent the tasks to which the activities belong.
[0012] Based on the heterogeneous process graph, a model is constructed that includes an input layer, multiple graph neural network layers, and an output layer;
[0013] Calculate the query vector, key vector, and value vector in the graph neural network layer, and calculate the attention score through the custom message function that fuses the edge features to measure the importance and relevance of each input node feature. The fused attention scores are aggregated to obtain the reconstructed node features.
[0014] Set a custom contrastive loss function, use the graph contrastive learning method, take the reconstructed node features after processing by the graph neural network layer as input, and calculate the contrastive loss;
[0015] The heterogeneous process graph is used as the input of the model. By comparing the loss function, the model parameters are optimized and adjusted to obtain the final model and optimized node features.
[0016] The optimized node features are clustered using spectral clustering and K-means algorithm to obtain the tasks to which each activity belongs.
[0017] A second aspect of the present invention provides an activity classification system supporting RPA and based on a graph neural network, comprising:
[0018] A log collection module, used to collect the interaction logs between users and applications in the case of multi-track interleaved execution;
[0019] Graph construction module, used to build heterogeneous graphs of processes based on interaction logs;
[0020] The label adding module is used to add node labels to the process heterogeneous graph;
[0021] The model building module is used to build a model consisting of an input layer, multiple graph neural network layers, and an output layer based on the process heterogeneous graph;
[0022] The feature calculation module is used to calculate the query vector, key vector, and value vector in the graph neural network layer, calculate the attention score through the custom message function that fuses the edge features, and aggregate the fused attention scores to obtain the reconstructed node features;
[0023] The loss calculation module is used to set a custom contrast loss function, use the graph contrast learning method, take the reconstructed node features processed by the graph neural network layer as input, and calculate the contrast loss;
[0024] The model optimization module is used to take the heterogeneous process graph as the input of the model, optimize and adjust the model parameters by comparing the loss function, and obtain the final model and optimized node features;
[0025] The clustering module is used to cluster the optimized node features using spectral clustering and K-means algorithm to obtain the tasks to which each activity belongs.
[0026] According to a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned activity classification method is implemented.
[0027] Beneficial effects of the present invention: The present invention performs clustering operations based on the output reconstruction features, which can effectively deal with the problem of poor classification effect when multiple tracks are intertwined and complex processes exist in the user operation behavior log. When constructing the model, the present invention integrates edge and node information into the attention mechanism. The obtained message feature representation comprehensively considers the relationship between nodes and edges in the context and more complex semantic information. The obtained reconstruction vector can more accurately represent the interaction between activities, which is conducive to the correct classification of complex track tasks in the user operation behavior log. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 Flowchart of the activity classification method to support robotic process automation and graph neural network.
[0029] Figure 2 This is a schematic diagram of the framework of the method proposed in the present invention. DETAILED DESCRIPTION
[0030] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings.
[0031] This embodiment provides a method for supporting robotic process automation and activity classification based on graph neural network, such as Figure 1 As shown, the following steps are included:
[0032] S1. Collect the user interaction logs between the user and the program when multiple track tasks are interleaved , as shown in Table 1.
[0033] Furthermore, activities , Represents a collection of events An event in Represents the time when the activity occurs. Representative operation user , Represents the target operation element and its value. When operating on the Web side, it represents the URL address of the corresponding web page. When operating on other sides, it is empty. The activity is based on the timestamp Sort in ascending order.
[0034] Table 1
[0035]
[0036] Data cleaning for user operation behavior logs includes the following sub-steps:
[0037] S1.1. For the blank fields of existing data in the user operation behavior log, Fill in the value. Get the activity . Each activity Corresponding to an event .event Use commonly used Event term definitions, e.g. , etc. term names.
[0038] S1.2. When generating user operation behavior logs, there will be a large number of operations that are irrelevant to the task or repeated, such as repeated clicks, accidental touches, etc. Therefore, it is necessary to denoise the data in order to remove the above operations from the user operation behavior logs.
[0039] S2. Construct a heterogeneous process graph, where each node of the graph represents a different activity of the user, and the edge represents the relationship between two activities. Depending on the relationship between the activities, the edges include follow-up edges, similarity edges, target edges, and response self-loop edges. The follow-up edge represents the relationship between two activity events; the similarity edge represents the similarity relationship between two activity events; the target edge represents the dependency relationship between two activity events; and the response self-loop edge refers to the edge connecting the node to itself, which is used to indicate that the node belongs to a repeated activity. Record the starting index that needs to be followed continuously, where the operation application is different but the content is the same, indicating that it is and This kind of context is highly relevant; the operation and application are the same and / The same workbook indicates that the operation is repeated on an application, which is a highly related context. These operations are all follow-up operations. At the same time, if there is a continuous follow-up relationship, the previous follow-up relationship is extended to the current index.
[0040] Secondly, similarity matching is used to process target relations and similar relations. of Concatenate into text strings, such as Figure 2 As shown in the heterogeneous graph construction phase, the user operation behavior log is converted , that is, in text string form, Special treatment is performed to remove prefix to reduce the similarity between different tasks. This operation is intended to normalize ,make The domain name portion of the does not affect the event representation.
[0041] use During the modeling process, each activity The description is converted into a vector representation of fixed dimension (size 5), the context window size is set to 2, and the minimum word frequency is set to 1, which helps capture the semantic relationship of each field feature in the activity. Obtain the corresponding embedding and use it as the text feature of the activity node in the process heterogeneous graph.
[0042] S3, such as Figure 2 As shown in the feature fusion stage, the upper part of the figure is a schematic diagram of the constructed heterogeneous process. Node labels are added to the heterogeneous process graph. Here, the node labels represent the tasks to which the activities belong. Node labels are not used to build graph neural networks, but are used to verify the classification effect of subsequent models. Text features are embedded as node feature information defined in the heterogeneous process graph to facilitate subsequent model training.
[0043] S4. Based on the heterogeneous process graph, a model is constructed that includes an input layer, multiple graph neural network layers, and an output layer, where multiple linear transformations are defined in the graph neural network layer for use in the multi-head attention mechanism. Calculation, such as Figure 2 As shown in the neural network construction stage, here, In the figure, it represents the query vector used by the current node (activity 1 in the figure) to match with neighboring nodes (activities 2 and 3 in the figure). It is used in the figure with the key vector to match against, is the actual information finally aggregated. Initialize the attention weights, node and edge type embeddings. is defined as follows:
[0044]
[0045]
[0046]
[0047] in, and Represents the source node and target node in the heterogeneous process graph, , , Respectively represent the query calculated by the node features of the previous layer ,key ,value Matrix-vector linear transformation function. , Respectively Node of layer (previous layer) and Here, the feature representation of the node in the 0th layer (i.e., the initial state) is the original input feature. Representative Target node in layer The query vector of the node The feature representation of the layer is obtained. In the Source node in layer The key vector of the node The feature representation of the layer is obtained. Representative Target node in layer The value vector of the node The feature representation of the layer is obtained.
[0048] The dimensions of the input and output layers are set to the dimensions of the node features. The number of graph neural network layers is set to 3, and the number of attention mechanism heads is set to 4. The attention weights are calculated by Initialization, here It is a neural network weight matrix initialization method. , are input and output feature dimensions. In the embodiment, the input layer and output layer The dimension setting is consistent with the dimension of the node feature, and the hidden layer Set to 128.
[0049] S5. Computational graph neural networks , the attention score is calculated by the custom message function of the fused edge feature to measure the importance and relevance of each input node feature. The fused attention score is aggregated to obtain the reconstructed node feature. The defined message function is defined as follows:
[0050]
[0051]
[0052]
[0053]
[0054] in, Represents the source node and the target node There are connecting edges between , is a weight matrix that depends on Edge type mapping function , the edge types are divided into three types: follower edge, similar edge, and target edge. The corresponding linear transformation function , ensuring that the model can dynamically adjust and transmit information for different types of edges in the process heterogeneous graph. Represents the source node under the action of a single attention head Passing the edge To the target node The message delivery results. Represents the splicing and merging of multiple attention heads ,in Refers to the number of multi-head attention heads. Figure 2 As shown in the neural network construction phase, the message function is finally obtained. Indicates the message passing of all attention heads, here represents the target node Through the edge From the source node Final message received.
[0055] Attention score Middle fusion edge features ,in, Represents the node feature dimension, represents the edge feature dimension, and For the query and key in S4, Represents the attention weight. Finally, the current target node The feature representation of is obtained by dot product of attention score and message function. The final output reconstructed node feature is transformed and output by the number of attention heads and output layer dimension.
[0056] S6. Model the embedded heterogeneous process graph, build a graph neural network, set a custom contrast loss function, use the graph contrast learning method, use the graph structure to add positive samples, use the reconstructed node features processed by the graph neural network layer in S5 as input, and calculate the contrast loss. The custom contrast loss function allows the model to complete self-supervised training based on the positive sample information of the graph structure without the need for node labels. Specifically:
[0057] S6.1. Calculate the similarity matrix , in , Representative Node Embedded node feature vector. The similarity matrix needs to exclude the diagonal to satisfy its own similarity for subsequent comparison loss calculation.
[0058] S6.2. Get all source nodes and target nodes of the current edge type and add positive sample masks according to the heterogeneous graph structure of the process , if there is an edge connecting the nodes and ,but ;otherwise . Positive sample similarity set Include All elements marked in . Negative sample similarity set Contains all elements except positive samples.
[0059] S6.3. For each positive sample node pair , calculate its contrast loss:
[0060]
[0061] in, is a hyperparameter that controls the distribution of similarity scores. The function converts the similarity score to a positive value, Nodes representing negative samples, Represents the positive sample pair All samples Positively correlated samples, It means taking the logarithm of the data and taking the negative value, and converting the relative probability between positive and negative samples into negative log-likelihood loss.
[0062] S7, taking the heterogeneous process graph as the input of the model constructed in S6, optimizing the model parameters by comparing the loss function, and training to obtain the final model and the optimized node feature embedding. The optimizer for training the model here adopts The algorithm is used, and the learning rate is set to 0.001 and the weight decay factor is set to 0.01. These hyperparameters contribute to the stability of the training process and the generalization ability of the model. The learning rate scheduler is introduced in the model ,The scheduler dynamically adjusts the learning rate based on the cosine annealing strategy, which helps to refine the model parameters in the later stage of training.
[0063] S8. Use spectral clustering and K-means algorithm to cluster the node features output by S7 to obtain the tasks to which each activity belongs.
[0064] First, reconstruct the node features Processing, ensure that the vector is array, because Provide efficient numerical computing support to improve the performance of subsequent clustering algorithms.
[0065] use The method was used to standardize the data and apply The method reduces the two-dimensional vector to one dimension. Spectral clustering is used to perform the first clustering of the reconstructed node features after dimensionality reduction to obtain the cluster labels of the output samples. Here, the spectral cluster labels are used in the K-means algorithm as sample weights to guide the algorithm to pay more attention to samples of specific categories.
[0066] K-means clustering is performed on the embedded data after dimensionality reduction, and the corresponding number of clusters is set to complete the activity classification. The number of classifications in the clustering algorithm is a hyperparameter, which is determined according to the detailed process. In the embodiment, the number of process classifications is 3.
[0067] After the clustering algorithm, the evaluation index is used to compare the statistical labels with the originally added labels to judge the quality of the classification results and obtain the final model classification results. The classification effect can effectively divide the labels into the corresponding three categories. The accuracy, precision, recall and The scores are calculated and the evaluation index is stable above 85%. The detailed experimental results of the embodiment are shown in Table 2.
[0068] Table 2
[0069]
[0070] Based on the same method concept as above, the embodiment of the present application also provides an activity classification system supporting robotic process automation and based on graph neural network, including:
[0071] A log collection module, used to collect the interaction logs between users and applications in the case of multi-track interleaved execution;
[0072] Graph construction module, used to build heterogeneous graphs of processes based on interaction logs;
[0073] The label adding module is used to add node labels to the process heterogeneous graph;
[0074] The model building module is used to build a model consisting of an input layer, multiple graph neural network layers, and an output layer based on the process heterogeneous graph;
[0075] The feature calculation module is used to calculate the query vector, key vector, and value vector in the graph neural network layer, calculate the attention score through the custom message function that fuses the edge features, and aggregate the fused attention scores to obtain the reconstructed node features;
[0076] The loss calculation module is used to set a custom contrast loss function, use the graph contrast learning method, take the reconstructed node features processed by the graph neural network layer as input, and calculate the contrast loss;
[0077] The model optimization module is used to take the heterogeneous process graph as the input of the model, optimize and adjust the model parameters by comparing the loss function, and obtain the final model and optimized node features;
[0078] The clustering module is used to cluster the optimized node features using spectral clustering and K-means algorithm to obtain the tasks to which each activity belongs.
[0079] Based on the same method concept as above, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the above activity classification method is implemented.
[0080] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. Support RPA and graph neural network-based activity classification method, characterized by: The following steps are involved: Collect the interaction logs between users and applications in the case of multi-trace interleaved execution; A heterogeneous process graph is constructed based on the interaction log, where each node of the graph represents a different activity of the user and the edge represents the relationship between two activities; Add node labels to the heterogeneous process graph. The node labels represent the tasks to which the activities belong. Based on the heterogeneous process graph, a model is constructed that includes an input layer, multiple graph neural network layers, and an output layer; Calculate the query vector, key vector, and value vector in the graph neural network layer, and calculate the attention score through the custom message function that fuses the edge features to measure the importance and relevance of each input node feature. The fused attention scores are aggregated to obtain the reconstructed node features. Set a custom contrastive loss function, use the graph contrastive learning method, take the reconstructed node features after being processed by the graph neural network layer as input, and calculate the contrastive loss; The heterogeneous process graph is used as the input of the model. By comparing the loss function, the model parameters are optimized and adjusted to obtain the final model and optimized node features. Use spectral clustering and K-means algorithm to cluster the optimized node features and get the tasks to which each activity belongs; In the graph neural network layer, the query vector, key vector, and value vector in the multi-head attention mechanism are defined as follows: ; in, s and t Represents the source node and target node in the heterogeneous process graph, , , Respectively represent the query calculated by the node features of the previous layer Q ,key K ,value V Matrix-vector linear transformation function, , Represents the nodes of the previous layer respectively s and t The characteristic representation of Representative l Target node in layer t The query vector is obtained by the feature representation of the previous layer of the node. Representative l Source node in layer s The key vector of is obtained by the feature representation of the previous layer of the node. Representative l Target node in layer t The value vector of is obtained by the feature representation of the previous layer of the node; The message function for calculating the attention score is defined as follows: ; in, Represents the source node under the action of a single attention head s Passing the edge e To the target node t The message delivery result is represents the linear transformation function, represents the weight matrix, It represents the effect of splicing multiple attention heads. vector, where Refers to the number of attention heads, Represents a message function, represents the attention score, W represents the attention weight, N represents the node feature dimension, L represents the edge feature dimension, Represents the fused edge feature.
2. The method according to claim 1, characterized in that The activities in the interaction log are sorted in ascending order according to timestamps, and events are defined using event terminology.
3. The method according to claim 1, characterized in that When constructing the heterogeneous process graph, the URLs in the activities are processed and their prefixes are removed to reduce the similarity between different tasks.
4. The method according to claim 3, characterized in that The Word2Vec model is used to convert the description of each activity into a vector representation of fixed dimension, the context window size is set to 2, and the minimum word frequency is set to 1.
5. The method according to any one of claims 1 to 4, characterized in that Set a custom contrast loss function. The steps to calculate the contrast loss function include: Calculate the similarity matrix , in , Representative Node i , j Embedded node feature vector; Get all source nodes and target nodes of the current edge type, and add positive sample masks according to the heterogeneous graph structure of the process , if there is an edge connecting the nodes i and j ,but ;otherwise ; For each positive sample node pair , calculate its contrast loss : ; in, represents a hyperparameter that controls the distribution of similarity scores. The exp function converts similarity scores into positive values. k Nodes representing negative samples, P( i ) represents the positive sample pair All samples in which is positively correlated with the sample.
6. The method according to claim 1, characterized in that The step of clustering the optimized node features comprises: Normalize the data of the reconstructed node features and reduce the two-dimensional vector to one dimension; Use spectral clustering to perform the first clustering on the reconstructed node features after dimensionality reduction to obtain the output sample cluster labels; Perform K-means clustering on the reduced-dimensional data, set the corresponding number of clusters, and complete the activity classification.
7. Support RPA and graph neural network-based activity classification system, characterized by: include: A log collection module, used to collect the interaction logs between users and applications in the case of multi-track interleaved execution; Graph construction module, used to build heterogeneous graphs of processes based on interaction logs; The label adding module is used to add node labels to the process heterogeneous graph; The model building module is used to build a model consisting of an input layer, multiple graph neural network layers, and an output layer based on the process heterogeneous graph; The feature calculation module is used to calculate the query vector, key vector, and value vector in the graph neural network layer, calculate the attention score through the custom message function that fuses the edge features, and aggregate the fused attention scores to obtain the reconstructed node features; The loss calculation module is used to set a custom contrast loss function, use the graph contrast learning method, take the reconstructed node features processed by the graph neural network layer as input, and calculate the contrast loss; The model optimization module is used to take the heterogeneous process graph as the input of the model, optimize and adjust the model parameters by comparing the loss function, and obtain the final model and optimized node features; The clustering module is used to cluster the optimized node features using spectral clustering and K-means algorithms to obtain the tasks to which each activity belongs; In the feature calculation module: In the graph neural network layer, the query vector, key vector, and value vector in the multi-head attention mechanism are defined as follows: ; in, s and t Represents the source node and target node in the heterogeneous process graph, , , Respectively represent the query calculated by the node features of the previous layer Q ,key K ,value V Matrix-vector linear transformation function, , Represents the nodes of the previous layer respectively s and t The characteristic representation of Representative l Target node in layer t The query vector is obtained by the feature representation of the previous layer of the node. Representative l Source node in layer s The key vector of is obtained by the feature representation of the previous layer of the node. Representative l Target node in layer t The value vector of is obtained by the feature representation of the previous layer of the node; The message function for calculating the attention score is defined as follows: ; in, Represents the source node under the action of a single attention head s Passing the edge e To the target node t The message delivery result is represents the linear transformation function, represents the weight matrix, It represents the effect of splicing multiple attention heads. vector, where Refers to the number of attention heads, Represents a message function, represents the attention score, W represents the attention weight, N represents the node feature dimension, L represents the edge feature dimension, Represents the fused edge feature.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the RPA-supporting and graph neural network-based activity classification method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Complex network node classification method based on graph attention network
CN112085124A
Graph neural network recommendation method, system and terminal based on interactive selection
CN116992099A