An RPA code generation method based on a deep learning model

Generating RPA code through deep learning models solves the problem that the RPA code syntax does not comply with the private SDK protocol and inconsistent format, and realizes an efficient automated process.

CN119376708BActive Publication Date: 2025-07-08HANGZHOU BRANCH INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411892542.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-07-08
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

In the existing fields of natural language processing and robot process automation, the generated RPA code syntax does not comply with the private SDK protocol, resulting in a large number of manual modifications and inconsistent code formats, affecting the efficiency of automation processes.

Method used

Using a deep learning model-based method, a vocabulary expansion module, a dynamic position coding mechanism and a task-related multi-head attention mechanism are generated to generate code that conforms to the RPA private SDK syntax and format.

Benefits of technology

It improves the syntax matching, format uniformity and automation of generated codes, reduces the need for manual modification, and improves the efficiency of automated processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119376708B_ABST
    Figure CN119376708B_ABST
Patent Text Reader

Abstract

The present invention discloses an RPA code generation method based on a deep learning model, comprising the following steps: Step 1: Collect and preprocess data based on a large code platform and a natural language description data source to obtain a data set; at the same time, construct and expand a vocabulary based on an existing RPA code library; Step 2: Construct a deep learning model, embed the expanded vocabulary into the deep learning model to form a vocabulary expansion module, and then use the mapped data set to train the deep learning model. During the training process, use a dynamic position encoding mechanism and a task-related multi-head attention mechanism to ensure the format consistency of the generated code and compliance with format and syntax requirements; Step 3: Use the trained deep learning model to generate RPA codes. The present invention can improve the syntax matching, format unity, accuracy and automation degree of the generated RPA private SDK codes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field of the present invention is the field of deep learning and natural language processing technology, and specifically relates to an RPA code generation method based on a deep learning model. Background Art

[0002] In the existing fields of natural language processing (NLP) and robotic process automation (RPA), generating code that meets specific syntax and format requirements has always been a challenge. Existing methods mainly focus on generating code in general programming languages (such as Python, Java, etc.), and the support for specific RPA syntax and private SDK protocols is relatively limited. Specifically, the code generated by existing methods does not conform to the RPA private SDK protocol, resulting in a large amount of manual modification; the code formats generated by existing methods are not unified, making it difficult to be parsed and processed by subsequent application programs, which affects the efficiency of the automation process. Summary of the Invention

[0003] The purpose of the present invention is to provide an RPA code generation method based on a deep learning model. The present invention can improve the syntax matching, format uniformity, accuracy, and automation degree of the generated RPA private SDK code.

[0004] The technical solution of the present invention: An RPA code generation method based on a deep learning model includes the following steps:

[0005] Step 1: Collect and preprocess data based on a large code platform and a natural language description data source to obtain a data set; at the same time, construct and expand a vocabulary based on the existing RPA code library;

[0006] Step 2: Construct a deep learning model, embed the expanded vocabulary into the deep learning model to form a word list expansion module, and then use the mapped data set to train the deep learning model. During the training process, use a dynamic position encoding mechanism and a task-related multi-head attention mechanism to ensure the format consistency of the generated code and compliance with format and syntax requirements;

[0007] Step 3: Use the trained deep learning model to generate RPA code.

[0008] In the aforementioned RPA code generation method based on a deep learning model, in Step 1, data is obtained by scraping a large code platform and a natural language description data source, the collected data is labeled and preprocessed into the JSONL format, and a mapped data set for generating RPA code from the natural language of users is constructed.

[0009] In the aforementioned RPA code generation method based on a deep learning model, in step one, relevant code snippets and documents are extracted from an existing RPA code library, and the extracted code snippets are preprocessed. The preprocessing includes removing noise and irrelevant information and retaining key syntax and function calls; then the code snippets are annotated and cleaned through regular expressions and natural language processing methods to form an RPA code library; an initial vocabulary is constructed from the RPA code library, and finally domain-specific terms and new words are added to obtain an extended vocabulary.

[0010] In the aforementioned RPA code generation method based on a deep learning model, in step two, a word vector model is used to map the words in the extended vocabulary into a high-dimensional vector and embed them into the deep learning model; the formula of the word vector model is as follows:

[0011] ;

[0012] Among them, and represent words in the extended vocabulary;, is the size of the vocabulary; is the word and the word 's co-occurrence count; and are respectively the word vectors of the word and the word ; and are the bias terms of the word and the word ; is a weighting function, defined as:

[0013] ;

[0014] Among them, and are respectively hyperparameters.

[0015] In the aforementioned RPA code generation method based on a deep learning model, in step two, the dynamic position encoding mechanism dynamically adjusts the parameters and function form of the position encoding according to the context information of the input code, and the formula is as follows:

[0016] ;

[0017] Among them, is the name of the encoding function, is the position parameter, is the index parameter, is a learnable function, is the context information.

[0018] In the above-mentioned RPA code generation method based on a deep learning model, in the task-related multi-head attention mechanism, there are a syntax attention head, a format attention head, and a context attention head;

[0019] The syntax attention head is used to capture the syntax structure in the code and extract the syntax features in the input code; the output formula of the syntax attention head is as follows:

[0020] ;

[0021] In the formula: is the query matrix, is the key matrix, is the value matrix; is the transpose operator, indicating the transpose of the matrix ; is the dimension of the key; is a normalization function used to convert the attention scores into a probability distribution;

[0022] The format attention head is used to capture the format features in the code and extract the format features in the input code; the output formula of the format attention head is as follows:

[0023] ;

[0024] The context attention head is used to enhance the ability of the attention mechanism to capture task-specific information. By encoding the context information of the input code, a context matrix is generated, and the context information is fused with the input code and used as the input of the context attention head. The formula is as follows:

[0025] ;

[0026] Among them, is the input code; is a concatenation function used to concatenate multiple tensors along a specified dimension;

[0027] The output formula of the context attention head is as follows:

[0028] .

[0029] In the above-mentioned RPA code generation method based on a deep learning model, in the task-related multi-head attention mechanism, according to the task requirements, the weight of the syntax attention head is adjusted as follows:

[0030] ;

[0031] The weight Adjust as follows:

[0032] 。

[0033] In the aforementioned RPA code generation method based on a deep learning model, the outputs of the syntax attention head, format attention head, and context attention head are weighted and summed to obtain the output of the task-related multi-head attention mechanism. The formula is as follows:

[0034] ;

[0035] where, 、 and have weight ranges in [0, 1], and 、 and have a total weight value of 1.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. The present invention introduces a variety of innovative algorithms and methods, such as task-related multi-head attention mechanism, dynamic position encoding mechanism, vocabulary embedding optimization, etc., which improves the overall performance of the model. The present invention expands the initial vocabulary according to task requirements and embeds it into the deep learning model to ensure that the generated code conforms to the RPA private SDK syntax, reducing the need for manual modification. And through the multi-head attention mechanism, it ensures that the syntax of the generated code meets the requirements. The present invention adopts a dynamic position encoding mechanism to make it applicable to specific tasks, improving the accuracy and format consistency of the generated code. In addition, the present invention also enhances the ability of the attention mechanism to capture task-specific information through context encoding and fusion mechanism, improving the accuracy and consistency of the generated code. The present invention combines the latest NLP and deep learning methods to ensure the advancement and forward-looking nature of the present invention.

[0038] 2. The present invention enables the model to adapt to different RPA task requirements and generate code that conforms to specific syntax and format through the task-related multi-head attention mechanism and domain-specific task fine-tuning. The present invention also enables the deep learning model to make dynamic adjustments according to the context information of the input code through dynamic position encoding and task-related weight learning, improving the adaptability of the generated code.

[0039] 3. The present invention generates RPA code in a unified format through specific formatting rules and optimization algorithms in the fine-tuning stage, facilitating the parsing and processing of subsequent application programs. The present invention designs an attention head specifically for capturing code format features to ensure the unity of the generated code format and improve the efficiency of the automated process.

[0040] 4. The present invention generates code in a unified format, facilitating subsequent automated parsing and processing by application programs, improving the automation and efficiency of the RPA process. It also reduces the need for manual modification and intervention by ensuring that the syntax and format of the generated code meet the requirements, thereby enhancing the overall work efficiency. Description of the Drawings

[0041] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0042] The present invention will be further described below in conjunction with the drawings and embodiments, but it shall not be used as a basis for limiting the present invention.

[0043] Embodiment: An RPA code generation method based on a deep learning model, which is applied to RPA process automation. RPA (Robot Process Automation) is an automation technology that uses software robots to simulate manual operations and process various repetitive and standardized tasks. The purpose to be achieved by the present invention is to make the generated RPA code meet the requirements of the SDK protocol. The SDK protocol refers to a series of rules and conventions in the Software Development Kit. The SDK protocol is to ensure the correct interaction between various components within the SDK, as well as the compatibility and consistency between the SDK and external systems. It usually includes regulations in aspects such as data format, communication protocol, calling convention, error handling, etc. When developers use the SDK, they need to follow these protocols in order to correctly use the functions provided by the SDK.

[0044] Specifically, as Figure 1 shown, it includes the following steps:

[0045] Step 1: Collect and preprocess data based on a large code platform and a natural language description data source to obtain a data set; at the same time, construct and expand a vocabulary based on the existing RPA code library;

[0046] In this step, a mapping data set of generating RPA code from users' natural language covering various programming tasks is constructed by scraping a large code platform and a natural language description data source through github. A semi-automated tool and domain experts are used to annotate the collected data to ensure data quality. The data is preprocessed into the JSONL format, and each record contains a set of conversations. The user describes the task in natural language, and the system generates the corresponding Python code.

[0047] Extract relevant code snippets and documentation from the existing RPA code library. Preprocess the extracted code snippets to remove noise and irrelevant information, and retain the key syntax and function calls. Then, use regular expressions and natural language processing methods (such as word segmentation and part-of-speech tagging) to annotate and clean the code snippets to form the input code and compose a new RPA code library. Next, construct an initial vocabulary from the RPA code library, including common functions, variable names, and operators. Use statistical methods (such as TF-IDF, which is a statistical method used to evaluate the importance of a word for a document set or a single document in a corpus) and deep learning models to expand the initial vocabulary, adding terms in specific domains (such as callback functions, recursive functions, and anonymous functions) and newly emerging vocabulary (such as workflows, subprocesses, and unattended), to obtain an extended vocabulary.

[0048] Step 2: Construct a deep learning model. Embed the extended vocabulary into the deep learning model to form a vocabulary expansion module. Then, use the mapped dataset to train the deep learning model. During the training process, use the dynamic positional encoding mechanism and the task-related multi-head attention mechanism to ensure the format consistency of the generated code and compliance with format and syntax requirements.

[0049] In this step, use the word vector model GloVe to map the words in the extended vocabulary to a high-dimensional vector and embed them into the deep learning model. The formula of the word vector model GloVe is as follows:

[0050] ;

[0051] where and represent words in the extended vocabulary, is the size of the vocabulary; is the word and the co-occurrence count of the word; and are the word vectors of the words and respectively; and are the bias terms of the words and respectively; is the weighting function, defined as:

[0052] ;

[0053] where and are hyperparameters respectively.

[0054] By optimizing the above GloVe objective function, the embedding vector of each word in the extended vocabulary is obtained.

[0055] Then, use the mapped dataset to train the deep learning model. During the training process, adopt a dynamic position encoding mechanism and a multi-head attention mechanism; use the mapped dataset in JSONL format for fine-tuning through the dialogue context. By fine-tuning the deep learning model, it can better adapt to specific RPA syntax and task requirements.

[0056] The design of the dynamic position encoding mechanism will focus on normalizing the generated code format, thereby ensuring that the generated code strictly follows the defined format. This can not only improve the efficiency of subsequent parsing and processing but also ensure seamless docking with the low-code platform. The following are the detailed principles and implementation steps of the dynamic position encoding mechanism, with the emphasis on ensuring the format consistency of the generated code.

[0057] According to the task requirements, the dynamic position encoding mechanism adjusts the parameters and functional forms of the position encoding to better capture the position dependencies in the input code. The dynamic position encoding mechanism dynamically adjusts the position encoding according to the context information of the input code, and the formula is as follows:

[0058] ;

[0059] where is the name of the encoding function, is the position parameter, is the index parameter, is the learnable function, is the context information.

[0060] The task-related multi-head attention mechanism processes the input code in parallel through multiple attention heads. By encoding the input code, it generates a query matrix , a key matrix and a value matrix ; Each attention head calculates a different attention distribution, and the formula is as follows:

[0061] ;

[0062] where is the dimension of the key.

[0063] According to the task requirements, adjust the parameters and calculation methods of each attention head to better capture the syntax and format information in the code, and introduce a task-specific weight matrix , and perform weighted summation on the output of each attention head;

[0064] ;

[0065] where is the The weights of the attention heads, is the number of attention heads.

[0066] In this embodiment, preferably, the task-specific multi-head attention mechanism is provided with a syntax attention head, a format attention head, and a context attention head; the syntax attention head is specifically used to capture the syntax structure in the code, such as function definitions, variable declarations, loop structures, etc., and extracts the syntax features in the code by using specific syntax rules and templates. For example, a function definition can be extracted by identifying the keyword `def` and the function name, and a variable declaration can be extracted by identifying the assignment operator `=`.

[0067] The formula for calculating the output of the syntax attention head is as follows:

[0068] ;

[0069] In the formula: is the query matrix, is the key matrix, is the value matrix; is the transpose operator, indicating that the matrix is transposed; is the dimension of the key; is the normalization function, which is used to convert the attention scores into a probability distribution so that the sum of the elements of the output is 1, and thus can be used for weight allocation.

[0070] Adjust the weights of the syntax attention head according to the task requirements , the formula is as follows:

[0071] ;

[0072] The format attention head is specifically used to capture the format features in the code, such as code indentation, comment style, variable naming convention, etc., and extracts the format features in the code by using specific format rules and templates. For example, code indentation can be extracted by identifying spaces or tabs, and the comment style can be extracted by identifying the comment symbol `#`.

[0073] The formula for calculating the output of the format attention head is as follows:

[0074] ;

[0075] Adjust the weights of the format attention head according to the task requirements , the formula is as follows:

[0076] .

[0077] Introduce task-related context information into the attention mechanism, such as the context of the code, function call relationships, etc. Through context encoding, enhance the ability of the attention mechanism to capture task-specific information.

[0078] The context attention head is used to enhance the ability of the attention mechanism to capture task-specific information, and its design can be achieved through the following steps: Encode the context information of the code snippet, such as function call relationships, variable scopes, etc., to generate a context matrix;

[0079] Fuse the task-related context information with the input code as the input of the context attention head to improve the sensitivity of the context attention head to task-specific information;

[0080] By encoding the context information of the input code, fuse the context information with the input code as the input of the attention mechanism. The formula is as follows:

[0081] ;

[0082] Among them, is the input code; A connection function for connecting multiple tensors along a specified dimension.

[0083] The formula for calculating the output of the context attention head is as follows:

[0084] .

[0085] The weight of the context attention head is optimized through deep learning model training. Finally, the context attention head is adjusted according to the weight as follows:

[0086] .

[0087] Fine-tune the deep learning model through formatting rules and optimization algorithms. In the fine-tuning stage, define specific formatting rules to ensure that the generated RPA code meets the unified format requirements. These rules can include code indentation, comment style, variable naming conventions, etc. Use optimization algorithms (such as Adam, RMSprop) to fine-tune the deep learning model to make the generated code better meet the needs of specific tasks.

[0088] At the initial stage of model training, according to the needs of specific tasks, initialize the weights of the attention heads to make them more inclined to capture specific syntax and format features.

[0089] During the model training process, through the backpropagation algorithm, dynamically adjust the weights of the attention heads to make them better adapt to the needs of specific tasks.

[0090] Design an attention head specifically for capturing the syntactic structure of code, such as function definitions, variable declarations, loop structures, etc.

[0091] Perform a weighted sum of the outputs of the syntactic attention head, format attention head, and context attention head to obtain the task-specific attention output, as shown in the following formula:

[0092] ;

[0093] Where 、 and are the weights of the syntactic attention head, format attention head, and context attention head, respectively. 、 and The weights of range from [0,1], and the sum of the weight values of , , and is 1.

[0094] Step 3: Use the trained deep learning model to generate RPA code.

[0095] Through the above design, the task-specific attention head can effectively capture the syntactic structure and format features in the code snippet, and combine the context information to improve the accuracy and consistency of the generated code.

[0096] ​​The present invention introduces a variety of innovative algorithms and methods, such as task-related multi-head attention mechanism, dynamic positional encoding mechanism, vocabulary embedding optimization, etc., which improve the overall performance of the model. The present invention expands the initial vocabulary according to task requirements and embeds it into the deep learning model to ensure that the generated code conforms to the RPA private SDK syntax, reducing the need for manual modification. And through the multi-head attention mechanism, it ensures that the syntax of the generated code meets the requirements. The present invention adopts the dynamic positional encoding mechanism to make it applicable to specific tasks, improving the accuracy and format consistency of the generated code. In addition, the present invention also enhances the ability of the attention mechanism to capture task-specific information through context encoding and fusion mechanism, improving the accuracy and consistency of the generated code. The present invention combines the latest NLP and deep learning methods to ensure the advancement and forward-looking nature of the present invention. In the fine-tuning stage, the present invention generates RPA code in a unified format through specific formatting rules and optimization algorithms, facilitating the parsing and processing of subsequent applications. The present invention enables the model to adapt to different RPA task requirements and generate code conforming to specific syntax and format through the task-related multi-head attention mechanism and domain-specific task fine-tuning. The present invention designs attention heads specifically for capturing code format features to ensure the uniformity of the generated code format and improve the efficiency of the automated process. By generating code in a unified format, the present invention facilitates the automated parsing and processing of subsequent applications, improves the automation and efficiency of the RPA process, and also reduces the need for manual modification and intervention by ensuring that the syntax and format of the generated code meet the requirements, improving the overall work efficiency. The present invention also enables the deep learning model to make dynamic adjustments according to the context information of the input code through dynamic positional encoding and task-related weight learning, improving the adaptability of the generated code.

[0097] In summary, the present invention can improve the syntax matching, format unity, accuracy and automation of the generated RPA private SDK code.

Claims

1. An RPA code generation method based on a deep learning model, characterized in that, It includes the following steps: Step 1: Obtain data by scraping large code platforms and natural language description data sources, annotate the collected data, preprocess the data into JSONL format, and construct a mapping dataset that generates RPA code from the user's natural language; at the same time, construct and expand the vocabulary based on the existing RPA code library; Step 2: Construct a deep learning model, embed the expanded vocabulary into the deep learning model to form a vocabulary expansion module, and then use the mapping dataset to train the deep learning model. During the training process, use the dynamic position encoding mechanism and the task-related multi-head attention mechanism to ensure the format consistency of the generated code and meet the format and syntax requirements; Step 3: Use the trained deep learning model to generate RPA code; In Step 2, use the word vector model to map the words in the expanded vocabulary into a high-dimensional vector and embed it into the deep learning model; the formula of the word vector model is as follows: where i and j represent the sequence numbers of words in the extended vocabulary, W is the size of the vocabulary; X ij is the co-occurrence times of word i and word j; v i and v' j are the word vectors of word i and word j respectively; b i and b' j are the bias terms of word i and word j; f(X ij ) is a weighting function, defined as: where X max and α are hyperparameters respectively; In Step 2, the dynamic position encoding mechanism dynamically adjusts the parameters and function forms of the position encoding according to the context information of the mapping dataset, and the formula is as follows: QE(loc,k)=h(loc,k,env): Where, QE is the name of the encoding function, loc is the position parameter, k is the index parameter, h is the learnable function, and env is the context information.

2. The RPA code generation method based on a deep learning model according to claim 1, wherein: In Step 1, extract relevant code snippets and documents from the existing RPA code library, preprocess the extracted code snippets, and the preprocessing includes removing noise and irrelevant information and retaining key syntax and function calls; then annotate and clean the code snippets through regular expressions and natural language processing methods to form an RPA code library; then construct an initial vocabulary based on the RPA code library, and finally add domain-specific terms and new words to obtain the expanded vocabulary.

3. The RPA code generation method based on a deep learning model according to claim 1, characterized in that: The task-related multi-head attention mechanism is provided with a syntax attention head, a format attention head, and a context attention head; The syntax attention head is used to capture the syntax structure in the mapping dataset and extract the syntax features in the mapping dataset; the output formula of the syntax attention head is as follows: Where: Q is the query matrix, K is the key matrix, and V is the value matrix; T is the transpose operator, indicating the transpose of matrix K; d k is the dimension of the key; softmax is a normalization function used to convert the attention scores into a probability distribution; The format attention head is used to capture the format features in the mapping dataset and extract the format features in the mapping dataset; The output formula of the format attention head is as follows: The context attention head is used to enhance the ability of the attention mechanism to capture task-specific information. By encoding the context information of the mapping dataset, a context matrix C is generated, and the context information is fused with the mapping dataset as the input of the context attention head. The formula is as follows: ContextFusion(X,C)=concat(X,C); Where, X is the mapping dataset; concat is a concatenation function used to concatenate multiple tensors along the specified dimension; The output formula of the context attention head is as follows:

4. The RPA code generation method based on a deep learning model according to claim 3, wherein: In the task-related multi-head attention mechanism, according to the task requirements, the syntactic attention head is adjusted according to the weight W syntax as follows: TaskSyntax Attention(Q, K, V, W syntax ) = W syntax ·SyntaxAttention(Q, K, V); The format attention head adjusts according to the weight W fomat as follows: TaskFormat Attention(Q, K, V, W format ) = W format ·FormatAttention(Q, K, V); The context attention head is adjusted according to the weight W format as follows: ContextAttention(Q, KV, C, W context ) = W context ·ContextAttention(Q, K, V, C).

5. The RPA code generation method based on a deep learning model according to claim 4, characterized in that: Perform a weighted sum of the outputs of the syntax attention head, the format attention head, and the context attention head to obtain the output of the task-related multi-head attention mechanism. The formula is as follows: TaskAttention(Q, K, V, W) = W syntax ·SyntaxAttention(Q, K, V) + W format ·FormatAttention(Q, K, V) + W context ·ContextAttention(Q, K, V, C); Among them, W syntax , W format and W context have weight ranges within [0, 1], and the sum of the weight values of W syntax , W format and W context is 1.

Citation Information

Patent Citations

  • Lightweight code generation method based on prompt learning

    CN116301893A

  • Python code automatic generation method and system

    CN116400901A