Robot process automation console control method and device
Through natural language interaction and large language model processing, combined with historical dialogue recording and vector similarity calculation, the problem of RPA console operation complexity and inefficiency is solved, and efficient and accurate user interaction and task execution is achieved.
Patent Information
- Application Number
- CN202510332181.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
The human-computer interactive experience of the existing RPA console relies on a graphical operation interface, which causes operators with non-technical backgrounds to cross a steep learning curve, which is inefficient in operation and error-prone.
Obtain user instructions through natural language interaction, use large language model segmentation and label generation instructions, combine historical dialogue records to judge potential conflicts, and call the console interface through vector similarity calculation to provide feedback information.
It reduces operational difficulty and learning costs, improves operation efficiency, avoids mistakes in complex tasks, and improves user interaction experience and task execution accuracy.
Smart Images

Figure CN120338699A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotic process automation, and particularly to a method and device for controlling a robotic process automation console. Background Art
[0002] Robotic Process Automation (RPA) technology is reshaping the enterprise-level business process management model. By deploying intelligent software agents to simulate manual operation modes, this technology enables the automated execution of standardized business processes, significantly improving operation efficiency, reducing the human error rate, and optimizing the operation cost structure.
[0003] Although mainstream RPA management platforms have integrated diverse functional components, there is still significant room for improvement in their human-computer interaction experience. Traditional interaction paradigms overly rely on Graphical User Interfaces (GUIs). Although they have the advantage of intuitiveness, they are difficult to meet the requirements of professional scenarios. Specifically, non-technical background operators need to overcome a steep learning curve to master the operation methods of the system, and for RPA management platforms relying on graphical user interfaces, the operation efficiency of operators is low and prone to errors when performing complex tasks.
[0004] Therefore, how to simplify the operation steps of the RPA console and improve the work efficiency of operators is a technical problem that needs to be solved currently. Summary of the Invention
[0005] This application provides a method and device for controlling a robotic process automation console to solve the technical problem of simplifying the operation steps of the RPA console and improving the work efficiency of operators.
[0006] To solve the above technical problem, an embodiment of this application provides a method for controlling a robotic process automation console, including:
[0007] Obtaining a natural language command sent by a user, and extracting a first instruction from the natural language command;
[0008] Judging whether there is a planned event conflicting with the first instruction according to the historical conversation record maintained currently;
[0009] If there is, returning a first feedback message to the user; otherwise, calling a first interface corresponding to the console according to the first instruction, returning a second feedback message to the user, and adding the first instruction and the second feedback message to the historical conversation record.
[0010] Compared with the prior art, the embodiments of the present application have the following beneficial effects: By replacing the traditional graphical interface operation with natural language interaction, the operation experience requirements of non-technical personnel are effectively reduced, thus significantly reducing the learning cost; at the same time, there is no need to operate the graphical interface, and for complex execution tasks, the operation process is effectively simplified, improving the work efficiency of the staff; finally, by maintaining the historical conversation records and dynamically detecting potential conflicts between the currently issued first instruction and the existing planned events based on the historical conversation records, it is possible to avoid work mistakes caused by the staff in handling complex tasks or work negligence, effectively avoid the execution of invalid tasks, and improve the work efficiency.
[0011] In some embodiments of the first aspect of the present application, obtaining the natural language command sent by the user and extracting the first instruction from the natural language command includes:
[0012] When the natural language command is voice data, recognizing the voice data to obtain text input;
[0013] When the natural language command is text data, using the text data as text input;
[0014] Segmenting the text input through a first large language model to obtain a number of text entities;
[0015] Determining the labels of the text entities through a second large language model and generating the first instruction according to the labels.
[0016] Compared with the prior art, the above embodiments have the following beneficial effects: By judging the presentation form of the user input data, it is possible to obtain accurate text input through corresponding processing methods whether it is voice or text input, reducing the restrictions and thresholds for users to use. Further, through different large language models, text segmentation and label determination are respectively realized, so as to accurately extract the user's intention, improve the accuracy of the instruction, and at the same time, without the user learning complex graphical interface operations, significantly reducing the operation difficulty and learning cost of the RPA console.
[0017] In some embodiments of the first aspect of the present application, judging whether there is a planned event conflicting with the first instruction according to the currently maintained historical conversation record includes:
[0018] When at least one second instruction in the historical conversation record meets the first condition, it is determined that there is a planned event conflicting with the first instruction; the first condition is set according to the first instruction.
[0019] Compared with the prior art, the above embodiments have the following beneficial effects: By maintaining the historical conversation records, before executing the first instruction, the second instruction with potential conflicts is searched from the currently maintained historical conversation records according to the first instruction to be executed, avoiding system errors or resource waste caused by repeated or conflicting operations.
[0020] In some embodiments of the first aspect of the present application, the first condition is set according to the first instruction, including:
[0021] Wherein, each instruction includes an operation type, an operation target, and an operation range parameter;
[0022] The operation range parameter of the second instruction overlaps with the operation range parameter of the first instruction;
[0023] The operation target of the second instruction is the same as the operation target of the first instruction;
[0024] Under the preset business rules, the operation type in the second instruction conflicts with the operation type of the first instruction.
[0025] Compared with the prior art, the above embodiments have the following beneficial effects: According to the specific parameter settings in the first instruction, the first condition used for conflict judgment is dynamically set, avoiding misjudgment caused by fuzzy judgment, ensuring that only when there is a real potential conflict is it feedback to the user, thereby improving the accuracy and overall efficiency of task execution.
[0026] In some embodiments of the first aspect of the present application, the calling the first interface corresponding to the console according to the first instruction includes:
[0027] Converting the first instruction into a first vector through a third large language model;
[0028] Calculating the similarity between the first vector and each second vector in the preset vector database, and screening the third vector with the highest similarity to the first vector from the second vectors; wherein, each of the second vectors corresponds to an interface;
[0029] Taking the interface corresponding to the third vector as the first interface, and after filling the parameters of the first interface according to the first instruction, sending a call request for the first interface to the console.
[0030] Compared with the prior art, the above embodiments have the following beneficial effects: By adopting the vector similarity calculation technology, the user instruction can be efficiently and accurately mapped to the most matching console interface, ensuring that the call request is accurate and error-free; at the same time, through vector similarity matching, the interface is automatically selected and called, avoiding human operations and the risk of mistakes caused by human operations, and improving the instruction execution speed.
[0031] In some embodiments of the first aspect of the present application, the returning of the second feedback information to the user includes:
[0032] Obtaining the operation result returned after the console performs an operation according to the call request of the first interface;
[0033] Converting the operation result into the second feedback information through a fourth large language model in combination with prompt engineering.
[0034] Compared with the prior art, the above embodiments have the following beneficial effects: Through the fourth large language model and prompt engineering technology, complex structured data is converted into feedback information that is easy to understand, improving the user interaction experience. At the same time, the operation result is fed back in a timely and accurate manner, enabling the user to clearly understand the instruction execution situation, and then making more reasonable operation decisions, overall improving the responsiveness and transparency of operating the RPA console.
[0035] In a second aspect, an embodiment of the present application further provides a robot process automation console control device, including: a first instruction acquisition module, a dialogue maintenance module, and an instruction execution module;
[0036] Among them, the first instruction acquisition module is used to obtain a natural language command sent by the user and extract a first instruction from the natural language command;
[0037] The dialogue maintenance module is used to determine whether there is a planned event conflicting with the first instruction according to the currently maintained historical dialogue record;
[0038] The instruction execution module is used to, if there is, return the first feedback information to the user; otherwise, call the corresponding first interface of the console according to the first instruction, return the second feedback information to the user, and add the first instruction and the second feedback information to the historical dialogue record.
[0039] In some embodiments of the second aspect of the present application, the obtaining of the natural language command sent by the user and extracting the first instruction from the natural language command includes:
[0040] When the natural language command is voice data, recognizing the voice data to obtain text input;
[0041] When the natural language command is text data, using the text data as text input;
[0042] Segmenting the text input through a first large language model to obtain a number of text entities;
[0043] Determine the tags of each of the text entities through a second large language model, and generate the first instruction according to the tags.
[0044] In some embodiments of the second aspect of the present application, the determining whether there is a planned event conflicting with the first instruction according to the currently maintained historical conversation record includes:
[0045] When there is at least one second instruction in the historical conversation record that satisfies the first condition, it is determined that there is a planned event conflicting with the first instruction; the first condition is set according to the first instruction.
[0046] In some embodiments of the second aspect of the present application, the invoking the corresponding first interface of the console according to the first instruction includes:
[0047] Convert the first instruction into a first vector through a third large language model;
[0048] Calculate the similarity between the first vector and each second vector in the preset vector database, and screen out the third vector with the highest similarity to the first vector from the second vectors; where each second vector corresponds to an interface respectively;
[0049] Use the interface corresponding to the third vector as the first interface, and after filling the parameters of the first interface according to the first instruction, send a call request for the first interface to the console. Description of the Drawings
[0050] Figure 1 It is a flowchart of a method for controlling a robotic process automation console provided in some embodiments of the present application;
[0051] Figure 2 It is a schematic diagram of the architecture of a robotic process automation console control system provided in some embodiments of the present application;
[0052] Figure 3 It is an interactive flowchart of adjusting a robot work schedule through a method for controlling a robotic process automation console provided in some embodiments of the present application;
[0053] Figure 4 It is a schematic diagram of the structure of a robotic process automation console control device provided in some embodiments of the present application. Detailed Embodiments
[0054] Traditional interaction paradigms rely too much on the Graphical User Interface (GUI). Although it has the advantage of intuitiveness, it is difficult to meet the needs of professional scenarios. Specifically, non-technical operators need to overcome a steep learning curve to master the operation methods of the system, and the RPA management platform relying on the graphical user interface has low operation efficiency and is prone to errors when performing complex tasks.
[0055] To solve the above technical problems, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0056] Before describing the embodiments of the present application, the professional terms involved in the present application will be explained first.
[0057] RPA: RPA refers to robotic process automation, specifically a technology that uses software robots or "bots" to automatically execute repetitive and standardized tasks. These tasks are usually involved in business processes, such as data entry, file operations, and communication. The goal of RPA is to improve efficiency, reduce errors, and free up human resources by reducing manual labor, enabling them to focus more on creative and strategic work. In many organizations, RPA has become one of the key tools to improve productivity.
[0058] RPA process: An RPA process refers to the steps of a specific task or a series of tasks executed in robotic process automation. These processes are usually designed to automatically complete a specific business process to reduce manual intervention and improve efficiency.
[0059] RPA console: An RPA console refers to a centralized platform or interface used to monitor, manage, and control robotic process automation (RPA). This console is usually provided by RPA tools to help administrators, developers, and other relevant users effectively manage RPA processes.
[0060] Semantic Role Labeling (SRL): It is an important task in the field of natural language processing, whose purpose is to identify and label the arguments (i.e., participants or involved objects) of each predicate (usually a verb or some adjectives) in a sentence, as well as the relationships between them. In other words, SRL aims to capture information such as "who did what to whom" and "under what circumstances" in a sentence.
[0061] Named Entity Recognition (NER): Also known as entity recognition or entity extraction, it is a subtask in the field of information extraction that mainly focuses on identifying specific categories of entities from unstructured text and classifying them. These entities can include person names, place names, organization names, time expressions, amounts, percentages, etc.
[0062] Embodiment 1
[0063] Please refer to Figure 1 , a method for controlling a robotic process automation console provided by an embodiment of the present application. Figure 2 In some embodiments of the present application, a robotic process automation console control system that can implement the method for controlling a robotic process automation console. Among them, the method for controlling a robotic process automation console includes S101 to S103, and the robotic process automation console control system includes a natural language processing module, an intent parsing engine, a dialogue management system, an execution engine, a feedback generator, and an RPA console.
[0064] Next, in combination with S101 to S103 described in the method for controlling a robotic process automation console and the robotic process automation console control system, the solution of the present application will be explained.
[0065] S101: Obtain the natural language command sent by the user, and extract the first instruction from the natural language command.
[0066] Refer to Figure 2 , a robotic process automation console control system. In some embodiments of the present application, the method for controlling a robotic process automation console described in the present application can be implemented through this system.
[0067] Furthermore, in some embodiments of the present application, S101 includes S1011 to S1012, specifically:
[0068] S1011: When the natural language command is voice data, recognize the voice data to obtain text input.
[0069] Furthermore, in some embodiments of the present application, when the natural language command is voice data, through Automatic Speech Recognition (ASR) technology, the voice data is converted into text data, and the present application does not make specific limitations on the adopted speech recognition scheme.
[0070] S1012: When the natural language command is text data, use the text data as text input.
[0071] S1013: Segment the text input through a first large language model to obtain a number of text entities.
[0072] Exemplarily, in some embodiments of the present application, when the text input entered by the user is obtained and S1013 is executed, it can be passed through Figure 2 the natural language processing module of the system shown to process the text input to obtain a number of text bodies. This module uses a pre-trained first large language model (LLM) that can accurately capture and understand the nuances of human language. Through continuous updating and customized training, the understanding ability of specific industry terms can be continuously improved, so as to better serve users in different fields.
[0073] Further, in some embodiments of the present application, Figure 2 the first large language model adopted by the natural language processing module selects GPT-4 or BERT as the base model and fine-tunes it according to the operation instructions of the RPA console. The fine-tuning methods include but are not limited to: extracting operation logs and user instructions from the RPA system to ensure that sufficient context information is included, such as the system or the state of the target being operated before and after the operation, the operation type (such as click, input text), and the operation object; further, according to the extracted operation logs and user instructions, using transfer learning, freezing some of the underlying parameters of the first large language model and only training the top-level parameters for fine-tuning; finally, accelerating the training process through Mixed Precision Training.
[0074] Further, in some embodiments of the present application, after the first large language model adopted by the natural language processing module is fine-tuned, the input text is tokenized, such as using Word Piece of BERT or Byte Pair Encoding (BPE) in GPT-4. For long texts, truncation or chunking is first performed to ensure that the text input entered into the first large language model later does not exceed the maximum limit of the model (such as 512 tokens).
[0075] S1014: Determine the labels of each text entity through a second large language model and generate a first instruction according to the labels.
[0076] Exemplarily, in some embodiments of the present application, when the text input entered by the user is completed with text segmentation and S1014 is executed, it can be passed through Figure 2The intent parsing engine of the system shown determines the tags corresponding to each text entity. The intent parsing engine uses a pre-trained second large language model (such as BERT) to perform SRL and NER tasks, so that key information can be extracted from complex sentences. For example, when the input text is "I want to view the task completion status of all financial departments last week", the intent parsing engine will identify the time range (last week), department (financial department), and query type (task completion status) from it, and construct a corresponding query structured instruction (i.e., the first instruction) according to the tags of each text entity and send it to the backend service.
[0077] Furthermore, in some embodiments of the present application, Figure 2 The architecture of the second large language model adopted by the intent parsing engine is specifically: adopting a multi-task learning framework, specifically: using a shared BERT encoder, which is respectively connected to the heads of the SRL and NER tasks. Among them, the SRL task uses a bidirectional LSTM (BiLSTM) and a CRF layer for sequence labeling; the NER task uses a linear classifier and a CRF layer for entity recognition. Further, the SRL task uses the instruction dataset in the RPA operation log as the training data; the NER task uses the entity annotation dataset in the RPA console (such as robot name, task type, etc.) as the training data.
[0078] It can be seen from the above embodiments that the present application realizes that whether it is voice or text input, accurate text input can be obtained through corresponding processing methods by judging the presentation form of the user input data, reducing the restrictions and thresholds when the user uses it. Further, through different large language models, text segmentation and label determination are respectively realized, so as to accurately extract the user's intention, improve the accuracy of the instruction, and at the same time, the user does not need to learn complex graphical interface operations, significantly reducing the operation difficulty and learning cost of the RPA console.
[0079] S102: According to the currently maintained historical conversation record, determine whether there is a planned event conflicting with the first instruction.
[0080] Furthermore, in some embodiments of the present application, determining whether there is a planned event conflicting with the first instruction according to the currently maintained historical conversation record includes:
[0081] When there is at least one second instruction in the historical conversation record that meets the first condition, it is determined that there is a planned event conflicting with the first instruction; the first condition is set according to the first instruction.
[0082] Exemplarily, in some embodiments of the present application, when executing S102, it can be passed through Figure 2The dialogue management system of the shown system dynamically maintains the historical dialogue record. Specifically, the dialogue management system remembers the previous dialogue content with the user and the natural language content returned by the feedback generator, and uses this information in subsequent exchanges to construct more relevant and coherent responses. The above process requires maintaining a dynamic and real-time updated dialogue state or context, which contains the dialogue information of previous rounds, similar to the function of short-term memory, and can store persistent information such as user preferences and historical interactions.
[0083] Furthermore, Figure 2 The shown dialogue management system, to solve the problem of dynamically maintaining the historical dialogue record, adopts an efficient data storage solution, such as an in-memory database (e.g., Redis) or a relational database (e.g., PostgreSQL), to persistently store the natural language input of the user; further through Figure 2 the shown intent parsing engine analyzes to obtain the second instruction, and the second feedback information corresponding to the second instruction after being processed by the feedback generator (to improve the data processing speed, the second instruction can be directly stored, and at the same time, only the input text of the user can be stored, and when judging conflicts each time, the second instruction is obtained again from the recorded historical input text through S1013 to S1014). This not only ensures the rapid retrieval and update of data, but also supports that when an exception occurs at any link, the system can intelligently guide the user to further provide the necessary natural language input. Combining with the existing dialogue history, the natural language processing module and the intent parsing engine can more accurately construct the structured instruction required for query, so as to provide a more satisfactory response for the user.
[0084] The data processing process of the dialogue management system after receiving the first instruction is as follows:
[0085] 1. Based on the obtained first instruction and the currently maintained dialogue state (including historical interaction records, user preferences, etc.), the dialogue management system dynamically updates the dialogue state or context environment. Specifically, it includes: persistently storing the new dialogue information in the database (such as Redis or PostgreSQL) for subsequent query and update. Whenever a new dialogue occurs or the existing dialogue state changes, the dialogue management system real-time updates the relevant records in the database to ensure the up-to-dateness and accuracy of the dialogue state. By maintaining a continuously updated dialogue data to track the history of the dialogue, the system can understand and respond to new requests based on the previous dialogue content.
[0086] 2. Using the updated dialogue state and the first instruction, the dialogue management system retrieves the historical dialogue record, through Figure 2With the cooperation of the natural language processing module and the intent parsing engine in it, several instructions are obtained from the historical conversation records. When at least one second instruction among these instructions conflicts with the first instruction, corresponding replies are further generated. Here, it may be necessary to combine the natural language content provided by the feedback generator to enhance the relevance and coherence of the replies.
[0087] As can be seen from the above embodiments, through the maintenance of the historical conversation records in this application, before executing the first instruction, according to the first instruction to be executed, the second instruction with potential conflicts is searched from the currently maintained historical conversation records, avoiding system errors or resource waste caused by repeated or conflicting operations.
[0088] Furthermore, in some embodiments of this application, the first condition is set according to the first instruction, including:
[0089] Among them, each instruction includes an operation type, an operation target, and an operation range parameter;
[0090] The operation range parameter of the second instruction overlaps with the operation range parameter of the first instruction;
[0091] The operation target of the second instruction is the same as the operation target of the first instruction;
[0092] Under the preset business rules, the operation type in the second instruction conflicts with the operation type in the first instruction.
[0093] Exemplarily, referring to Figure 3 , when the first instruction corresponding to the text input by the user is "Operation type: Update working time configuration; Operation target: Robot A; Operation range parameter: Start time: 09:00 End time: 18:00", first, it is judged whether there is an instruction whose operation range parameter conflicts with the operation range parameter of the above first instruction. If so, then it is judged whether the operation target of this instruction and the operation target of the first instruction are both Robot A. If they are the same, then according to the preset business rules, it is judged whether the operation type of this instruction conflicts with the operation type of the first instruction (for example, if the operation type of this instruction is also update, it means there is no conflict, but if the operation type is add, it will cause this instruction to conflict with the first instruction within the overlapping operation range parameters). This application does not specifically limit the preset business rules.
[0094] As can be seen from the above embodiments, this application dynamically sets the first condition used for conflict judgment according to the specific parameters in the first instruction, avoiding misjudgment caused by fuzzy judgment, ensuring that only when there is indeed a potential conflict is it feedback to the user, thereby improving the accuracy and overall efficiency of task execution.
[0095] S103: If it exists, return the first feedback message to the user; otherwise, call the corresponding first interface of the console according to the first instruction, return the second feedback message to the user, and add the first instruction and the second feedback message to the historical conversation record.
[0096] Further, in some embodiments of the present application, the calling of the corresponding first interface of the console according to the first instruction includes:
[0097] Convert the first instruction into a first vector through the third large language model;
[0098] Calculate the similarity between the first vector and each second vector in the preset vector database, and screen the third vector with the highest similarity to the first vector from the second vectors; where each second vector corresponds to an interface respectively.
[0099] Use the interface corresponding to the third vector as the first interface, and after filling the parameters of the first interface according to the first instruction, send a call request for the first interface to the console.
[0100] Exemplarily, in some embodiments of the present application, when calling the corresponding first interface of the console according to the first instruction, relevant operations can be executed through Figure 2 the execution engine of the system shown. Among them, Figure 2 the execution engine is a bridge connecting the intelligent control device and the RPA console. The execution engine calls various services provided by the RPA platform by using RESTful API, SOAP protocol or other adapters. Exemplarily, if a specific robot instance needs to be started, the execution engine will construct a correct HTTP POST request, carry the necessary parameters (such as robot ID, input parameters, etc., which are extracted from the first instruction), and initiate a call to the RPA server.
[0101] Further, in some embodiments of the present application, in order to accurately identify the first interface, it is first necessary to store the vectorized representation of the corresponding mapping rules of each interface (Application Programming Interface, API) in a vector database (such as Pinecone, Weaviate or FAISS) to facilitate subsequent rapid positioning of the first interface through similarity retrieval. Specifically, it includes the following three steps:
[0102] 1. Convert the mapping rules corresponding to each API (including information such as API endpoints, methods, parameters, etc.) into text descriptions;
[0103] 2. Use a pre-trained third large language model (such as BERT or Sentence-BERT) to convert the above text description into a vector form (i.e., the second vector);
[0104] 3. Store the transformed second vector in a selected vector database for subsequent retrieval.
[0105] Further, in some embodiments of the present application, when Figure 2 After the execution engine of the system shown receives the first instruction, the first instruction is structured text data (usually in JSON format), and the first instruction details the type of operation to be performed and its related parameters (such as the robot ID and input parameters required to start a specific robot instance, etc.). Therefore, for the received first instruction, the execution engine realizes the invocation of the first interface through the following three steps:
[0106] 1. Use the vector database for similarity retrieval to find the API mapping rule that best matches the current first instruction;
[0107] 2. Extract the action parameters required for the first interface corresponding to the API mapping rule from the first instruction, map the extracted action parameters to the corresponding RPA platform API endpoint, and determine the corresponding HTTP request method (GET, POST, etc.);
[0108] 3. Extract all necessary parameters from the first instruction and ensure that these parameters are correctly filled into the upcoming API request to ensure the integrity and accuracy of the request. Finally, the execution engine sends this request to the RPA server according to the HTTP request determined in step 2 (for example, for a request to start a specific robot, this may be an HTTP POST request carrying the necessary parameters) to trigger the corresponding automated process or operation.
[0109] It can be seen from the above embodiments that the present application can efficiently and accurately map user instructions to the best-matching console interface by adopting the vector similarity calculation technology, ensuring that the call request is accurate and error-free; at the same time, through vector similarity matching, the interface is automatically selected and called, avoiding manual operations and the risk of errors caused by manual operations, and improving the instruction execution speed.
[0110] Further, in some embodiments of the present application, the returning of the second feedback information to the user includes:
[0111] Obtain the operation result returned after the console performs the operation according to the call request of the first interface;
[0112] Through the fourth large language model, combined with prompt engineering, transform the operation result into the second feedback information.
[0113] Exemplarily, in some embodiments of the present application, it can be through Figure 2The feedback generator of the system shown obtains the second feedback information. The feedback generator converts the operation results from the RPA console into an easy-to-understand text description through a pre-trained fourth large language model. In addition, the feedback generator can also enrich the presentation form of the second feedback information through data visualization techniques, such as chart generation or simple summary reports. In addition, for some abnormal situations or error messages, the feedback generator will also give clear guidance to help users solve problems.
[0114] Furthermore, in some embodiments of the present application, Figure 2 The steps for the feedback generator of the system shown to generate the second feedback information include:
[0115] 1. The RPA (Robotic Process Automation) console sends the data of the operation results to the execution engine, and the execution engine provides the received operation results and the data of the intent analysis engine, the dialogue management system, and the natural language processing module to the feedback generator;
[0116] 2. According to the output result of the intent analysis engine, identify the key information points of the user for the operation results, and select appropriate chart types (such as bar charts, pie charts, line charts, etc.) so that the second feedback information can be intuitively displayed later;
[0117] 3. Extract the user's preference description from the dialogue management system and the natural language processing module as the prompt words for the fourth large language model, and customize the personalized second feedback information according to the operation results. For example, provide more detailed technical analysis to users with a stronger technical background; while for ordinary users, emphasize the significance and impact of the results.
[0118] 4. Synthesize the second feedback information in different presentation forms obtained in steps 2 and 3, and automatically generate a concise and clear text summary report.
[0119] Wherein:
[0120] Returning the RPA console as structured data (such as JSON format), the feedback generator can convert it into a natural language description or a visual chart.
[0121] For the data that needs to be converted into a natural language description, use the fourth large language model (such as GPT-3.5) combined with prompt engineering (prompt) to generate a natural language description and return it to the user.
[0122] For the data that needs to be converted into a visual chart, use Matplotlib or Plotly to generate a chart (such as a pie chart, bar chart, etc.) picture, and convert the picture into a Base64 encoding and return it to the user.
[0123] For the case where users customize charts according to their own needs, the feedback generator supports receiving operation instructions from the intent parsing engine, extracting information such as chart types and data dimensions in the instructions, generating chart images that meet the specific requirements of users using Matplotlib or Plotly, and converting the images into Base64 encoding and returning them to users.
[0124] The Base64-encoded images can be directly embedded into HTML for convenient front-end display.
[0125] For cases of abnormal situations or error messages, the feedback generator utilizes the fourth largest language model to classify the exceptions into different types, such as system errors, network errors, user input errors, etc., and provides solution suggestions for each type of exception to help users quickly understand the problem and take corresponding solutions.
[0126] As can be seen from the above embodiments, through the fourth largest language model and prompt engineering technology, this application converts complex structured data into easily understandable feedback information, improves the user interaction experience, and at the same time timely and accurately feedbacks the operation results, enabling users to clearly understand the instruction execution situation, and then make more reasonable operation decisions, overall improving the responsiveness and transparency of operating the RPA console.
[0127] To better understand the solution of this application, next, taking the example of a user adjusting the robot work schedule, combined with Figure 2 the system shown, and Figure 3 the interaction process shown, the method for controlling the robot process automation console in some embodiments of this application will be illustrated by examples:
[0128] Suppose the user wants to adjust the work schedule of Robot A to 9:00 to 18:00 through a voice command. First, the natural language processing module captures the user's voice signal, transcribes it into text, and extracts useful information from it to organize it into a text input that can be understood by a computer, that is: "Action: Adjust; Object: Robot A; Attribute: Working hours; Value: 9:00 to 18:00"; Then, the intent parsing engine analyzes which robot the user wants to modify and its new working period. In this example, "adjust the working hours of Robot A" actually means modifying a specific setting, and an API needs to be called or the database needs to be updated. Therefore, the first instruction output by the intent parsing engine directly points to the operation that can be executed, that is: "Operation type: Update working hours configuration; Target entity: Robot A; Parameters: Start time: 09:00 End time: 18:00"; Subsequently, after the dialogue management system confirms that the first instruction will not conflict with other scheduled plans; After that, the execution engine updates the corresponding settings in the RPA console through an API call according to the set time period parameters; The RPA console feeds back the result of successful operation to the execution engine, and the execution engine issues instructions and relevant data to the feedback generator to make it prepare the feedback information; Finally, the feedback generator generates the second feedback information to inform the user that the change has taken effect successfully, and the latest schedule can be viewed in real time on the console.
[0129] In summary, a method for controlling a robot process automation console provided in an embodiment of the present application has the following beneficial effects: By replacing traditional graphical interface operations with natural language interaction, the operation experience requirements of non-technical personnel are effectively reduced, thereby significantly reducing the learning cost; At the same time, there is no need to operate on the graphical interface. For complex execution tasks, the operation process is effectively simplified, and the work efficiency of staff is improved; Finally, by maintaining the historical conversation record and dynamically detecting potential conflicts between the currently issued first instruction and existing planned events based on the historical conversation record, the work mistakes caused by staff in dealing with complex tasks or work negligence are avoided, the execution of invalid tasks is effectively avoided, and the work efficiency is improved.
[0130] Embodiment 2
[0131] Reference Figure 4 , a device for controlling a robot process automation console provided in an embodiment of the present application, includes: a first instruction acquisition module 201, a dialogue maintenance module 202, and an instruction execution module 203.
[0132] Further, in some embodiments of the present application, the first instruction acquisition module 201 is configured to acquire a natural language command sent by a user and extract a first instruction from the natural language command; the dialogue maintenance module 202 is configured to determine whether there is a planned event conflicting with the first instruction according to the currently maintained historical dialogue record; the instruction execution module 203 is configured to, if so, return a first feedback message to the user; otherwise, call a first interface corresponding to the console according to the first instruction, return a second feedback message to the user, and add the first instruction and the second feedback message to the historical dialogue record.
[0133] Further, in some embodiments of the present application, the acquiring a natural language command sent by a user and extracting a first instruction from the natural language command includes: when the natural language command is voice data, recognizing the voice data to obtain text input; when the natural language command is text data, using the text data as text input; segmenting the text input through a first large language model to obtain a plurality of text entities; and determining labels of each text entity through a second large language model and generating the first instruction according to the labels.
[0134] Further, in some embodiments of the present application, the determining whether there is a planned event conflicting with the first instruction according to the currently maintained historical dialogue record includes: when at least one second instruction in the historical dialogue record satisfies a first condition, determining that there is a planned event conflicting with the first instruction; the first condition is set according to the first instruction.
[0135] Further, in some embodiments of the present application, the first condition being set according to the first instruction includes: wherein each instruction includes an operation type, an operation target, and an operation range parameter; the operation range parameter of the second instruction overlaps with the operation range parameter of the first instruction; the operation target of the second instruction is the same as the operation target of the first instruction; and under a preset business rule, the operation type in the second instruction conflicts with the operation type in the first instruction.
[0136] Further, in some embodiments of the present application, the calling a first interface corresponding to the console according to the first instruction includes: converting the first instruction into a first vector through a third large language model; calculating the similarity between the first vector and each second vector in a preset vector database, and screening a third vector with the highest similarity to the first vector from the second vectors; wherein each second vector corresponds to an interface; using the interface corresponding to the third vector as the first interface, and after filling the parameters of the first interface according to the first instruction, sending a call request for the first interface to the console.
[0137] Further, in some embodiments of the present application, the returning of the second feedback information to the user includes: obtaining an operation result returned by the console after performing an operation according to the call request of the first interface; and converting the operation result into the second feedback information through a fourth large language model in combination with prompt engineering.
[0138] It can be understood that the above device item embodiments correspond to the method item embodiments of the present invention. An active distribution network capacity configuration device provided by the embodiments of the present invention can implement any method item embodiment of the present invention, that is, the active distribution network capacity configuration method provided in Embodiment 1.
[0139] In summary, a robot process automation console control device provided in the embodiments of the present application has the following beneficial effects: replacing traditional graphical interface operations with natural language interaction effectively reduces the operation experience requirements for non-technical personnel, thereby significantly reducing the learning cost; at the same time, there is no need to operate on the graphical interface, which effectively simplifies the operation process for complex execution tasks and improves the work efficiency of staff; finally, by maintaining historical conversation records and dynamically detecting potential conflicts between the currently issued first instruction and existing planned events based on the historical conversation records, it avoids work mistakes caused by staff in handling complex tasks or work negligence, effectively avoids the execution of invalid tasks, and improves work efficiency.
[0140] Embodiment 3
[0141] Based on the above embodiments of the active distribution network capacity configuration method, another embodiment of the present application provides an active distribution network capacity configuration terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the active distribution network capacity configuration method of any embodiment of the present application.
[0142] Exemplarily, in this embodiment, the computer program can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the active distribution network capacity configuration device.
[0143] The active distribution network capacity configuration device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The active distribution network capacity configuration terminal device may include, but is not limited to, a processor and a memory.
[0144] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the active distribution network capacity configuration device, and connects various parts of the entire active distribution network capacity configuration device through various interfaces and lines. The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the active distribution network capacity configuration device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0145] Embodiment 4
[0146] Based on the embodiments of the above active distribution network capacity configuration method, another embodiment of the present application provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the active distribution network capacity configuration method of any embodiment of the present application.
[0147] In this embodiment, the above storage medium is a computer-readable storage medium, and the computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0148] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only specific embodiments of the present application and is not used to limit the protection scope of the present application. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for controlling a robotic process automation console, characterized in that, Including: Obtain a natural language command sent by a user, and extract a first instruction from the natural language command; According to the currently maintained historical conversation record, determine whether there is a planned event that conflicts with the first instruction; If there is, return a first feedback message to the user; otherwise, call the corresponding first interface of the console according to the first instruction, return a second feedback message to the user, and add the first instruction and the second feedback message to the historical conversation record.
2. The method for controlling a robotic process automation console according to claim 1, wherein, The obtaining the natural language command sent by the user and extracting the first instruction from the natural language command includes: When the natural language command is voice data, recognize the voice data to obtain text input; When the natural language command is text data, use the text data as text input; Segment the text input through a first large language model to obtain several text entities; Through a second large language model, determine the labels of each text entity, and generate the first instruction according to the labels.
3. A method for controlling a robotic process automation console according to claim 1, characterized in that, The determining whether there is a planned event that conflicts with the first instruction according to the currently maintained historical conversation record includes: When at least one second instruction in the historical conversation record meets the first condition, it is determined that there is a planned event that conflicts with the first instruction; the first condition is set according to the first instruction.
4. The method for controlling a robotic process automation console according to claim 3, wherein The first condition is set according to the first instruction, Including: Wherein, each instruction includes an operation type, an operation target, and an operation range parameter; The operation range parameter of the second instruction overlaps with the operation range parameter of the first instruction; The operation target of the second instruction is the same as the operation target of the first instruction; Under the preset business rules, the operation type in the second instruction conflicts with the operation type in the first instruction.
5. A method for controlling a robotic process automation console according to claim 1, characterized in that, The calling the corresponding first interface of the console according to the first instruction includes: Convert the first instruction into a first vector through a third large language model; Calculate the similarity between the first vector and each second vector in the preset vector database, and screen the third vector with the highest similarity to the first vector from the second vectors; wherein, each second vector corresponds to an interface; Use the interface corresponding to the third vector as the first interface, and after filling the parameters of the first interface according to the first instruction, send a call request for the first interface to the console.
6. The method for controlling a robotic process automation console according to claim 1, wherein The returning the second feedback message to the user includes: Obtain the operation result returned by the console after performing the operation according to the call request of the first interface; Through a fourth large language model, combined with prompt engineering, convert the operation result into the second feedback message.
7. A robot process automation console control device, characterized in that, Including: A first instruction acquisition module, a conversation maintenance module, and an instruction execution module; Wherein, the first instruction acquisition module is used to obtain a natural language command sent by a user and extract a first instruction from the natural language command; The conversation maintenance module is used to determine whether there is a planned event that conflicts with the first instruction according to the currently maintained historical conversation record; The instruction execution module is used to return the first feedback information to the user if it exists; otherwise, call the corresponding first interface of the console according to the first instruction, return the second feedback information to the user, and add the first instruction and the second feedback information to the historical conversation record.
8. The robotic process automation console control device according to claim 7, wherein, Obtaining the natural language command sent by the user and extracting the first instruction from the natural language command includes: When the natural language command is voice data, recognizing the voice data to obtain text input; When the natural language command is text data, using the text data as text input; Dividing the text input through a first large language model to obtain several text entities; Determining the labels of each text entity through a second large language model and generating the first instruction according to the labels.
9. The robot process automation console control device according to claim 7, characterized in that, Judging whether there is a planned event conflicting with the first instruction according to the currently maintained historical conversation record includes: When at least one second instruction in the historical conversation record meets the first condition, it is determined that there is a planned event conflicting with the first instruction; the first condition is set according to the first instruction.
10. A robot process automation console control device according to claim 7, characterized in that, Calling the corresponding first interface of the console according to the first instruction includes: Converting the first instruction into a first vector through a third large language model; Calculating the similarity between the first vector and each second vector in the preset vector database, and screening the third vector with the highest similarity to the first vector from the second vectors; each second vector corresponds to an interface respectively; Taking the interface corresponding to the third vector as the first interface, filling the parameters of the first interface according to the first instruction, and sending a call request for the first interface to the console.