Robot process automation design method, device and equipment based on large model, medium and product

By combining large models and multimodal scenario understanding algorithms, executable RPA operation processes are generated, which solves the problems of insufficient generalization capabilities and insufficient understanding of complex interfaces in existing technologies, and realizes efficient process design and execution in untrained scenarios.

CN120805970APending Publication Date: 2025-10-17CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510899102.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing robotic process automation design solutions rely too much on preset question-and-answer library training models, resulting in insufficient generalization capabilities and an inability to flexibly respond to untrained business scenarios. In addition, when traditional RPA is combined with large models, it lacks a deep multimodal understanding of complex interface environments, making it difficult to generate operational processes with strong adaptability and high execution success rates.

Method used

By obtaining the business requirement information input by the user, combining the large model and prompt word template, generating subtask sequences and business domain types, and extracting text data through multimodal scene understanding and conversion algorithms, using thinking chain technology to generate executable operation step sequences, combined with RPA components for automated operations.

Benefits of technology

It achieves accurate understanding of user intentions in untrained business scenarios, generates RPA operation processes with strong adaptability and high execution success rate, lowers the threshold for user understanding and operation, and expands the application scope of RPA technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805970A_ABST
    Figure CN120805970A_ABST
Patent Text Reader

Abstract

The invention provides a robot process automation design method and device based on a large model, equipment, a medium and a product, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining business demand information inputted by a user, and obtaining a first cue word template based on the business demand information, the first cue word template and a first large model; the subtask sequence comprises at least one subtask, and each subtask corresponds to a different service scene interface; obtaining text state data based on each business scene interface; obtaining an executable operation step sequence based on the subtask sequence, the text state data corresponding to each business scene interface, the type, the second cue word template and the second large model; according to the sequence of the operation steps in the operation step sequence, robot process automation assemblies corresponding to the operation steps are called and executed in sequence. Therefore, the user demand can be automatically analyzed, the accurate automatic process can be generated, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a robot process automation design method and device based on a large model, equipment, medium and product. BACKGROUND

[0002] In today's digital wave, the rapid development of artificial intelligence technology is reshaping enterprise operation mode with unprecedented strength, leading a new era of efficiency and cost optimization. Among them, RPA (Robotic Process Automation) as an innovative solution combining desktop automation and advanced artificial intelligence technology, can simulate human computer terminal operation tasks, and let software robots automatically process a large number of repetitive and complex workflow tasks, thereby significantly improving employee work efficiency and reducing operating costs, so it has become a key driving force for digital transformation and upgrading of many enterprises.

[0003] When users use RPA to perform tasks, they need to use design tools to customize RPA execution scripts according to their own business processes. This process roughly covers two core links: RPA process design: users need to fully understand the task requirements and intentions, and convert them into RPA process task descriptions; RPA process generation: users flexibly combine the ability modules provided by the RPA platform according to the operation steps disassembled in the previous step, and build RPA automation processes that meet their own business needs and are executable.

[0004] However, in the RPA process design stage, beginners and non-IT (Information Technology) professional users often face many challenges. First, the RPA platform integrates a large number of highly specialized ability components, each component has its unique attributes and configuration requirements, for users who are not familiar with programming and automation technology, it takes a high learning curve to understand and use these components, and this process is often time-consuming and tedious. Second, it is difficult for non-professional users to decompose complex business processes into RPA robot executable steps, users need to deeply understand the business process and master the ability to convert abstract logic into specific automation tasks, if the decomposition is not detailed or accurate enough, it may cause automation execution deviation, affecting the accuracy and efficiency of the business.

[0005] The existing RPA process design scheme has the following disadvantages: limited generalization ability, and unable to accurately respond to demands outside the training scene. Currently, generative RPA process design technology mainly relies on a question and answer database constructed by user consultation and operation sequence to train a natural language understanding model, so as to identify each link of the business process and the required capability component. However, this method has significant limitations, that is, the model can only accurately analyze and respond to demands within the preset scene, and it is difficult to output accurate operation sequences for business scenarios beyond the existing question and answer data range, and the generalization ability is limited; the success rate of process execution in complex scenarios is low. The traditional RPA products on the current market face certain limitations when integrating large models, especially in processing complex interface process design scenarios. It is often difficult for large models to flexibly adapt to the variability of the environment due to the lack of sufficient context information, which affects the success rate of the designed RPA process execution.

[0006] In summary, the existing RPA process design scheme relies too much on preset question and answer library training models, resulting in insufficient generalization ability, and being unable to flexibly respond to untrained business scenarios. When traditional RPA is combined with large models, it is difficult to generate operation processes with strong adaptability and high execution success rate due to the lack of multi-modal deep understanding of complex interface environments. SUMMARY

[0007] Embodiments of the present application provide a robot process automation design method, device, equipment, medium and product based on a large model, to solve the technical problem that the existing robot process automation design scheme relies too much on preset question and answer library training models, resulting in insufficient generalization ability, and being unable to flexibly respond to untrained business scenarios. When traditional robot process automation is combined with large models, it is difficult to generate operation processes with strong adaptability and high execution success rate due to the lack of multi-modal deep understanding of complex interface environments.

[0008] To solve the above technical problems, the present application is implemented as follows:

[0009] In a first aspect, embodiments of the present application provide a robot process automation design method based on a large model, the method comprising:

[0010] obtaining business demand information input by a user, based on the business demand information, a first prompt word template and a first large model, obtaining a subtask sequence and a type of a business domain to which the business demand information belongs, the subtask sequence comprising at least one subtask, each subtask in the at least one subtask corresponding to a different business scenario interface;

[0011] based on each of the business scenario interfaces, obtaining text state data corresponding to each of the business scenario interfaces;

[0012] obtaining an executable operation step sequence based on the subtask sequence, the text state data corresponding to each of the business scenario interfaces, the type, the second prompt word template, and the second large model;

[0013] sequentially invoking and executing the robot process automation components corresponding to the operation steps in the operation step sequence based on the operation step sequence.

[0014] Optionally, before obtaining the business requirement information input by the user, obtaining a subtask sequence and a type of a business domain to which the business requirement information belongs based on the business requirement information, a first prompt word template, and a first large model, the method further comprises:

[0015] based on sample business requirement information, a type of a business domain to which the sample business requirement information belongs, and a sample subtask sequence split based on the sample business requirement information, fine-tuning an initial large model to obtain the first large model by using a model fine-tuning technology.

[0016] Optionally, based on each of the business scenario interfaces, obtaining text state data corresponding to each of the business scenario interfaces comprises:

[0017] performing scene analysis on each of the business scenario interfaces to obtain multi-modal information corresponding to each of the business scenario interfaces;

[0018] performing compression processing and / or filtering processing on the multi-modal information corresponding to each of the business scenario interfaces to obtain processed multi-modal information corresponding to each of the business scenario interfaces;

[0019] respectively converting the processed multi-modal information corresponding to each of the business scenario interfaces to obtain text state data corresponding to each of the business scenario interfaces.

[0020] Optionally, performing scene analysis on each of the business scenario interfaces to obtain multi-modal information corresponding to each of the business scenario interfaces comprises at least two of the following:

[0021] performing text detection on each of the business scenario interfaces by using an OCR algorithm to obtain text content;

[0022] performing target detection on each of the business scenario interfaces by using a target detection algorithm to obtain a category of a first UI element and coordinate information of the first UI element;

[0023] performing target detection on each of the business scenario interfaces by using a visual language model to obtain a category of a second UI element and coordinate information of the second UI element; wherein the complexity of the second UI element is greater than the complexity of the first UI element.

[0024] obtain desktop information and taskbar information where each of the business scenario interfaces is located.

[0025] Optionally, based on the sub-task sequence, the text state data corresponding to each of the business scenario interfaces, the type, the second prompt word template, and the second large model, an executable operation step sequence is obtained.

[0026] Based on the sub-task sequence, the text state data corresponding to each of the business scenario interfaces, the type, the second prompt word template, and the second large model, an executable operation step sequence is obtained in combination with a thinking chain technology.

[0027] Optionally, before sequentially calling and executing the RPA components corresponding to the operation steps in the operation step sequence according to the order of the operation steps in the operation step sequence based on the operation step sequence, the method further comprises:

[0028] predefining a correspondence table between the RPA components and the operation steps;

[0029] Based on the operation step sequence, sequentially calling and executing the RPA components corresponding to the operation steps in the operation step sequence according to the order of the operation steps in the operation step sequence comprises:

[0030] Based on the operation step sequence and the correspondence table, sequentially calling and executing the RPA components corresponding to the operation steps in the operation step sequence according to the order of the operation steps in the operation step sequence.

[0031] In a second aspect, an embodiment of the present application provides a robot process automation design device based on a large model, the device comprising:

[0032] An acquisition module is configured to acquire business requirement information input by a user, obtain a sub-task sequence and a type of a business field to which the business requirement information belongs based on the business requirement information, a first prompt word template, and a first large model, and the sub-task sequence comprises at least one sub-task, and each of the at least one sub-task corresponds to a different business scenario interface.

[0033] An execution module is configured to obtain text state data corresponding to each of the business scenario interfaces based on each of the business scenario interfaces.

[0034] Based on the sub-task sequence, the text state data corresponding to each of the business scenario interfaces, the type, the second prompt word template, and the second large model, an executable operation step sequence is obtained.

[0035] Based on the sequence of operation steps, the RPA components corresponding to the operation steps in the sequence of operation steps are invoked and executed in sequence according to the order of the operation steps.

[0036] In a third aspect, an embodiment of the present application provides a network device, comprising a processor, a memory, and a program stored in the memory and executable on the processor, and when the program is executed by the processor, the steps of the method for designing a robot process automation based on a large model according to the first aspect are implemented.

[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for designing a robot process automation based on a large model according to the first aspect are implemented.

[0038] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, and when the computer instructions are executed by a processor, the steps of the method for designing a robot process automation based on a large model according to the first aspect are implemented.

[0039] In the embodiment of the present application, the system performs semantic analysis and task decomposition by combining the first prompt word template and the first large model based on the business requirement information input by the user, generates a structured subtask sequence and a corresponding domain type determination result. Each subtask in the subtask sequence is explicitly directed to a specific business scenario interface, ensuring accurate matching of task granularity and business scenarios. Subsequently, for each business scenario interface, the system extracts and generates normalized text data, which carries the parameters and context information required for interface operation. Next, the system inputs the subtask sequence, text data, domain type, and second prompt word template into the second large model, and generates an executable operation step sequence through multi-dimensional information fusion, which defines the calling logic and execution order of the RPA component in detail. Finally, the system triggers the corresponding RPA component to complete the automation operation in strict accordance with the arrangement of the operation step sequence. In this way, through the multi-stage reasoning capability of the large model and the precise execution capability of the RPA, the technical penetration from natural language requirements to automation operation is realized, effectively improving the efficiency and accuracy of business process processing. BRIEF DESCRIPTION OF DRAWINGS

[0040] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not intended to limit the scope of the present application. Moreover, the same reference numerals are used throughout the same figures. In the drawings:

[0041] Figure 1A flowchart of a robot process automation design method based on a large model provided for an embodiment of the present application is shown in FIG. 1.

[0042] Figure 2 An architecture block diagram of an intelligent design system provided for an embodiment of the present application is shown in FIG. 3.

[0043] Figure 3 An input-output diagram of a first large model provided for an embodiment of the present application is shown in FIG. 4.

[0044] Figure 4 A flowchart of a multi-modal scene understanding and conversion method provided for an embodiment of the present application is shown in FIG. 5.

[0045] Figure 5 A screen image understanding information diagram provided for an embodiment of the present application is shown in FIG. 6.

[0046] Figure 6 A desktop and taskbar icon information diagram provided for an embodiment of the present application is shown in FIG. 7.

[0047] Figure 7 A JSON format operation object diagram provided for an embodiment of the present application is shown in FIG. 8.

[0048] Figure 8A In a flow design module, an input of a large model and a prompt word completion diagram provided for an embodiment of the present application is shown in FIG. 9.

[0049] Figure 8B In a flow design module, an output of a large model diagram provided for an embodiment of the present application is shown in FIG. 10.

[0050] Figure 9 A structure block diagram of a robot process automation design device based on a large model provided for an embodiment of the present application is shown in FIG. 11.

[0051] Figure 10 A structure block diagram of a network device provided for an embodiment of the present application is shown in FIG. 12. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.

[0053] Figure 1 A robot process automation design method based on a large model provided for an embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the method comprises:

[0054] Step S101, obtaining the service requirement information input by the user, and based on the service requirement information, a first prompt word template and a first large model, obtaining a subtask sequence and a type of a business field to which the service requirement information belongs;

[0055] The subtask sequence includes at least one subtask, and each subtask corresponds to a different business scenario interface.

[0056] Step S102, based on each business scenario interface, obtaining text state data corresponding to each business scenario interface respectively;

[0057] Step S103, based on the subtask sequence, the text state data corresponding to each business scenario interface respectively, the type, a second prompt word template and a second large model, obtaining an executable operation step sequence;

[0058] Step S104, based on the operation step sequence, sequentially calling and executing RPA components corresponding to the operation steps in the operation step sequence according to the order of the operation steps in the operation step sequence.

[0059] It should be noted that the technical process constructs a full-link processing mechanism from business requirement input to automatic operation execution. The intelligent design system first receives the service requirement information input by the user, performs semantic analysis and structured processing by guiding the first large model through the first prompt word template, outputs a subtask sequence including multiple independent subtasks, and simultaneously determines the type of the field to which the service requirement information belongs. Each subtask strictly corresponds to a specific business scenario interface, forming an atomized task unit. Then, the system extracts and generates text state data containing field definitions, operation rules and other elements according to the interaction characteristics and data specifications of each business scenario interface, providing rich context information for subsequent operation instruction generation. In the operation step generation stage, the system integrates the process logic of the subtask sequence, the interface characteristics of the text state data, the domain type knowledge base and the constraint conditions of the second prompt word template, and drives the second large model to generate an executable operation step sequence that conforms to the actual system operation specification. Finally, through the standardized interface of the RPA component, the system strictly follows the action sequence and parameter configuration defined by the operation step sequence to complete the automatic operation execution across business systems.

[0060] In summary, the method shown in the embodiments of the present application deeply combines RPA technology with large models, computer vision and other AI technologies, and proposes an intelligent RPA process design system. The system has the ability of task understanding and intelligent decision-making, and can select the most suitable agent for process design according to the professional field to which the user's demand belongs. Even in the face of business scenarios that have never appeared in the training process, it can accurately understand user intent, complete complex task splitting, and have stronger expandability. At the same time, by introducing multi-modal scene understanding and conversion algorithms, various complex modal data on the screen are unified into text modal that is easy to process, which is input to the large model, enriches its context vision, and improves the depth and breadth of the large model's understanding of the current business scenario, making it easier to understand business requirements, identify environmental changes, and generate more accurate and effective RPA operation processes.

[0061] Moreover, the user does not need to invest a lot of time and effort to deeply study the various capability components of the RPA platform, but only needs to clearly state his own requirements to the system through intuitive dialogue interaction, so as to complete the intelligent design of the RPA process. This greatly simplifies the operation process, reduces the technical use threshold, and expands the application range of RPA technology.

[0062] In a possible implementation manner, before the step S101 of acquiring the business requirement information input by the user, and based on the business requirement information, the first prompt word template and the first large model, obtaining the sub-task sequence and the type of the business field to which the business requirement information belongs, the method further comprises: based on the sample business requirement information, the type of the field to which the sample business requirement information belongs, and the sample sub-task sequence split based on the sample business requirement information, combining model fine-tuning technology, fine-tuning the initial large model to obtain the first large model.

[0063] It should be noted that in this possible implementation, a model pre-training optimization stage is added before the core flow is executed, and the business adaptation ability of the first large model is improved through a domain knowledge injection mechanism. The specific implementation logic is as follows: before the formal deployment of the system, based on the sample business demand information accumulated by the enterprise, the type of the sample business annotated by the artificial sample business domain, and the sample subtask sequence disassembled by the expert, the model fine-tuning technology such as LoRA (Low-Rank Adaptation, low-rank adaptation) is used to optimize the general first large model. This process combines the demand description text in the actual business scenario, the corresponding domain classification label, and the standardized subtask path into training data, so that the first large model learns the task disassembly mode and domain feature association rules specific to the business domain. For example, in the field of supply chain finance, the model can accurately identify the association logic of the business scenario interface implied in the demand such as "accounts receivable financing application" and "goods warehouse verification" and the like through a large number of sample demands and their corresponding subtask chains. The fine-tuned first large model has stronger domain task analysis capability. In the subsequent formal operation stage, when the user input business demand information is obtained, the system calls the fine-tuned first large model instead of the original model, combines the first prompt word template for demand analysis, and outputs a subtask sequence and accurate domain type determination result that is highly consistent with the actual business scenario, providing a reliable basis for the generation of subsequent operation steps.

[0064] In a possible implementation, the step S102, obtaining the text state data corresponding to each business scenario interface respectively based on each business scenario interface comprises: performing scene analysis on each business scenario interface respectively to obtain multi-modal information corresponding to each business scenario interface respectively; performing compression processing and / or filtering processing on the multi-modal information corresponding to each business scenario interface respectively to obtain processed multi-modal information corresponding to each business scenario interface respectively; and converting the processed multi-modal information corresponding to each business scenario interface respectively to obtain the text state data corresponding to each business scenario interface respectively.

[0065] The scene analysis is performed on each business scenario interface respectively to obtain the respective multi-modal information corresponding to each business scenario interface, including at least two of the following: text content is obtained by performing text detection on each business scenario interface respectively using an OCR (Optical Character Recognition) algorithm; the category and coordinate information of a first UI (User Interface) element are obtained by performing target detection on each business scenario interface respectively using a target detection algorithm; the category and coordinate information of a second UI element are obtained by performing target detection on each business scenario interface respectively using a visual language model; the complexity of the second UI element is greater than that of the first UI element; and desktop information and taskbar information of each business scenario interface are obtained.

[0066] It should be noted that the technical process can realize high-precision extraction and standardized expression of business scenario interface features through multi-modal information fusion and structured conversion mechanism. Specifically, the system performs deep scene analysis on each business scenario interface, adopts multiple technologies such as OCR algorithm, target detection algorithm, visual language model for parallel processing, and obtains original multi-modal information including text content, UI element attribute, and interface environment information. In specific implementation, the OCR algorithm is responsible for extracting all visible text content in the interface; the target detection algorithm identifies the category and coordinate position of the basic UI element (such as button, input box); the visual language model performs semantic analysis on complex second UI elements such as combined controls (such as dynamic table, nested menu); and the running environment parameters such as desktop resolution and taskbar state are captured. Subsequently, the original multi-modal information is compressed (such as removing duplicate text paragraphs) and filtered (such as removing non-operable decorative elements) through data cleaning rules, to eliminate information redundancy and retain core interaction features. Finally, the optimized multi-modal information is converted into a standard text data format, thereby providing high-fidelity, low-noise input data for subsequent operation steps, providing more comprehensive context information, and effectively avoiding the problem of inaccurate RPA operation steps caused by missed interface element detection.

[0067] In a possible implementation, based on the sub-task sequence, the text data corresponding to each business scenario interface, the type, the second prompt word template, and the second large model, the executable operation step sequence is obtained by combining the thinking chain technology.

[0068] It should be noted that the technical process strengthens the logical coherence and system operation compliance of the operation step generation stage by introducing the thinking chain technology. In the operation step sequence generation process, the second large model, based on the execution order of the subtask sequence, the interface feature description in the business scenario interface text state data, the domain type constraint rule, and the guidance framework of the second prompt word template, combines the thinking chain technology and adopts a step-by-step reasoning mechanism to make multi-level decisions: first, analyze the field definition and operation constraints in the business scenario interface text state data corresponding to the current subtask (such as "the maximum value of the payment amount input box is limited to 1 million"), then activate the corresponding business rule knowledge base according to the domain type, and then gradually generate atomic operation instructions containing operation object positioning, parameter filling, and exception handling elements through the causal deduction path of the thinking chain, and finally assemble these instructions into a complete operation step sequence according to the logical dependency relationship. This mechanism can make the generated RPA operation steps more accurate.

[0069] In a possible implementation, before sequentially calling and executing the RPA components corresponding to the operation steps in the operation step sequence according to the order of the operation steps in the operation step sequence, the method further includes: defining a correspondence table between the RPA components and the operation steps in advance; and sequentially calling and executing the RPA components corresponding to the operation steps according to the order of the operation steps in the operation step sequence based on the operation step sequence includes: sequentially calling and executing the RPA components corresponding to the operation steps according to the order of the operation steps in the operation step sequence based on the operation step sequence and the correspondence table.

[0070] It should be noted that the precise matching of operation steps and RPA components can be achieved through the component mapping predefinition mechanism to ensure the reliability and maintainability of the automation execution process. In the system deployment stage, a correspondence table between RPA components and operation steps is established in advance, which stores the unique identifier of each operation step and its associated RPA component execution entry in the form of key-value pairs. Therefore, the reliability and maintainability of the automation execution process can be ensured.

[0071] In summary, the method provided in the embodiments of the present application has the following advantages: the service scenario understanding capability is improved. The traditional RPA products on the current market face certain limitations when integrating large models, especially in processing complex interface process design scenarios. Often, due to the lack of sufficient context information input into the large model, it is difficult to flexibly adapt to the variability of the environment. This directly leads to many obstacles in the actual execution of the designed RPA process, and the executability is greatly reduced. The method shown in the embodiments of the present application introduces a multi-modal scenario understanding and conversion algorithm to capture and analyze rich information in the current interface, including but not limited to text content, image elements and other multi-modal data. The complex content on the screen is uniformly converted into text modal that is easy to process, and is provided to the large model in the form of context for RPA process design, so that the large model can better understand the business requirements, recognize environmental changes, and generate more accurate and effective RPA operation sequences.

[0072] The adaptability of the model is improved. The existing generative RPA process design technology mainly relies on a question and answer database constructed by user questioning and operation sequence construction to train a natural language understanding model, so as to identify each link of the business process and the required capability components. However, this method has significant limitations, that is, the model can only accurately analyze and respond to the requirements within the preset scenario. For business scenarios beyond the existing question and answer data range, it is difficult to output accurate operation sequences, and the generalization ability is limited. To solve this problem, the embodiments of the present application propose an intelligent RPA process design system by deeply combining RPA technology with large models and computer vision AI technology. The system has the ability of task understanding and intelligent decision-making, and can select the most suitable intelligent agent for process design according to the professional field to which the user demand belongs. Even in the face of business scenarios that have never appeared in the training process, the system can accurately understand the user's intention, complete the splitting of complex tasks, and has stronger generalization ability.

[0073] Now, the intelligent design system shown in the embodiments of the present application will be explained from the specific application scenario and the virtual module.

[0074] The intelligent design system provided in the embodiments of the present application supports users to express their business requirements in natural language form, and gradually generates the operation sequence of each business process through multiple rounds of interaction. Finally, the system can autonomously complete the whole chain conversion from business requirements to RPA process design, realize the intelligentization and high efficiency of process design, and the system architecture is as shown in Figure 2 The system can be divided into three key parts: requirement analysis module, scenario understanding module and process design module.

[0075] 1. Requirement analysis module

[0076] The requirement analysis module in the embodiments of the present application is responsible for analyzing the business requirements input by the user. Based on the powerful understanding and generation capability of the large model, the intention of the user is identified, and the business requirements of the user are further subdivided into multiple sub-tasks in combination with the prompt words. Through semantic understanding, expert knowledge, model fine-tuning and other technologies, the business field to which the current task belongs is determined, and the subsequent process design module in a specific field is further guided.

[0077] Specifically, after the user inputs the business requirements in a natural language manner, the module performs semantic analysis on the user input based on the large model, combines the prompt word technology, splits the tasks according to different operation interfaces, and generates a coarse-grained task sequence. The process design interface between each sub-task is different, ensuring that the process design of each sub-task can obtain accurate scene information in the corresponding scene understanding module, and adapt to different interface environments. At the same time, based on expert knowledge and model fine-tuning technology, the module can intelligently identify and determine the business type to which the current business requirement belongs, and hand over to the process design module of the corresponding business field for subsequent step processing.

[0078] The specific implementation process is taken as an example of financial statement export:

[0079] <1> The user describes the business process requirement of financial statement export in a natural language form: what is the login account information, which interface of the financial system to enter, and what format to export the report data;

[0080] <2> Convert the requirement description into a coarse-grained task sequence: based on the prompt word, use the fine-tuned large model to understand and split the user requirements, take the operation interface as the division dimension, and convert it into a coarse-grained task sequence, including opening the financial system, logging into the financial system, entering the report interface, etc., to achieve task splitting of the financial statement export business requirement.

[0081] <3> Determine the business type according to the requirement description: based on expert knowledge, use the fine-tuned large model to determine the business type to which the current task belongs. The user input involves financial statements and financial systems, and the module determines that the current business requirement is related to the financial field, and needs to call the financial-process design module for subsequent processing, such as Figure 3 as shown.

[0082] The fine-tuned large model refers to fine-tuning a large language model by artificially constructing a high-quality process design dataset. The fine-tuning method is not limited to LoRA (Low-Rank Adaptation), P-Tuning (Prompt Tuning), Adapter Tuning, etc., thereby optimizing the performance of the model on specific tasks, so that the model can better adapt to and complete specific tasks. The process design dataset contains multiple coarse-grained process design data described in natural language. Each piece of data is divided into input and output. The input is the user's business requirements, and the output includes task sequences split by interface dimensions and the business type of the current task. The specific fine-tuning sample is shown in Table 1.

[0083] Table 1

[0084]

[0085] 2. Scene understanding module

[0086] The scene understanding module uses a multi-modal scene understanding and conversion algorithm to identify and detect the multi-modal information of the current business scene interface. The identified objects include text, interactive UI elements, desktop icons, etc. The detection results are converted into text modal data and provided as context information of the current interface environment to the subsequent process design module, enriching the context view of the large model. The algorithm integrates a visual understanding model and a graphical application automation operation component. The specific analysis process is shown in Figure 4

[0087] First, the system obtains the current screen image by taking a screenshot, detects and recognizes the text using an OCR algorithm, obtains the text content of the current interface, detects general UI interactive components using a target detection algorithm, obtains UI element categories and coordinate information, and uses a visual language model to understand complex screen images to obtain special UI element information. In addition, the module calls the operating system automation library to obtain the icon information and text information of the desktop and taskbar. The obtained multi-modal data is filtered and compressed, and the compressed multi-modal data is converted into text data as context information provided to the subsequent process design module.

[0088] ​There is a lot of redundancy in the multi-modal information obtained from the screen image initially, and if it is directly input into a large model without processing, it will consume a lot of computing resources. The algorithm uses a multi-modal information compression and filtering algorithm to remove invalid context information. Specifically, for component and icon information, geometric rules are used for screening, including length, width, intersection ratio, area, confidence, etc. attributes, and only interactive element icon descriptions and their coordinate information are retained; for text information, text similarity detection is used for screening, the text information is input into the embedding model to obtain the text vector, the similarity between the text vectors is calculated, only the text content with a certain threshold of similarity to the business scenario requirement is retained, and the compressed and filtered information is converted into a text modality and input into the subsequent process design module.

[0089] The specific implementation process takes the scene understanding of the system login interface as an example:

[0090] <1> Screen image understanding: obtain the login interface screenshot of the business scenario, input into OCR, target detection, and multi-modal visual language model for screen understanding, and obtain the element and text information related to the login function, such as Figure 5 as shown.

[0091] <2> Icon information retrieval: obtain window information through an automatic component library, retrieve desktop and taskbar icon information based on window information, and finally extract the coordinates of the icons and the text they represent to obtain the desktop and taskbar icon information of the user's operating scenario, such as Figure 6 as shown.

[0092] <3> Multi-modal information compression and filtering: filter the multi-modal information of the login interface obtained in steps <1> and <2>, including text information such as "register VIP", "register new account", "forget password", and some misidentified or meaningless UI element information.

[0093] <4> Text modality conversion: convert the filtered multi-modal scene information into natural language description, including text coordinates and text content, etc. for text content; and for UI element icons, including element type, coordinate information, and extended description, etc.

[0094] 3. Process design module

[0095] The process design module obtains the task requirements of each interface and the context information of the current scenario, dynamically analyzes the current environment intelligently based on large language models and prompt engineering, generates a system understandable and standardized process description sequence, and can adapt and adjust the business process design in complex and variable environments.

[0096] The whole flow design module utilizes the thinking chain technology to interact with the large model in multiple rounds, and according to the historical information and the current context information, the next task is deeply thought and reasoned. The whole thinking chain interaction mode mainly includes the following four parts: ① question: the question input by the user; ② task record: the result of the last interaction and the action taken; ③ observation: thinking about the next action according to the task record; ④ action: the action taken by the agent after decision-making.

[0097] In addition, in order to ensure that the system can analyze the process generated by the large model, and enable the RPA designer to find the corresponding RPA capability component according to the operation sequence, the module defines a standardized operation sequence of the process for the RPA component capability, and each operation corresponds to a capability component in the RPA designer. When defining the operation object, the corresponding capability component type and key attributes need to be explicitly indicated, and the operation granularity cannot be further divided. In order to standardize the output of the large model, the following takes JSON (JavaScript Object Notation) as an example to construct the operation sequence rule. Each element corresponds to an operation object in JSON format, and the structure is as shown in Figure 7

[0098] The JSON format operation object includes the following attribute fields: id, action, component, type, value, ref, desc.

[0099] ① The id field is used to identify the order of operation execution, which is represented in an increasing manner in the sequence.

[0100] ② The action field is used to identify the operation type, such as moving, clicking, and inputting, etc. Each type of atomic operation corresponds to a capability component in the RPA designer.

[0101] ③ The component field is used to identify the capability component of the RPA designer corresponding to the current operation.

[0102] ④ The type field is used to represent the business sub-type of the same type of operation, so as to be able to handle more complex business scenarios.

[0103] ⑤ The value field represents the parameter value required by the current operation. For example, keyboard input requires content parameters, and waiting operation requires a waiting time parameter, etc.

[0104] ⑥ The ref field references the ID of a previous operation, indicating that the value attribute value of the current operation depends on the atomic operation result corresponding to the ID. If the parameter value of the current atomic operation does not depend on other operations, the field is set to -1.

[0105] ​The desc field provides a description of the current operation, so that the large model can understand the context and perform system debugging.

[0106] The specific implementation takes logging into a financial system in financial statement export as an example, as shown in FIGS. Figure 8A and Figure 8B

[0107] <1> Input the current task and scene context information: in the login interface, input the task "log in to the financial system, account is abc@xx.com, password is 123", and the element information of the current interface.

[0108] <2> Large model prompt word completion: complete the prompt word according to the template based on the thinking chain set by the financial field agent, and the obtained text content is used as the input of the large model.

[0109] <3> Large model flow design: based on the completed prompt word, interact with the flow design agent to obtain a system parsable flow operation sequence.

[0110] In summary, the method shown in the embodiments of the present application proposes the following points:

[0111] 1. Task decomposition and intelligent decision based on large model. Based on the powerful understanding and generation ability of the large model, combined with semantic understanding, expert knowledge, model fine-tuning and other technologies, the input business requirements are understood and analyzed, the business requirements are converted into multiple coarse-grained task sequences according to different operation interfaces, and the business type is determined, realizing the full-automatic demand analysis and reducing the user understanding cost.

[0112] 2. Multi-modal scene understanding and conversion. The application proposes a multi-modal scene understanding and conversion algorithm, which generates text description information, UI interaction element components and other multi-modal information by analyzing the business process scene, helps the large model to better understand the current business scene, and generates more accurate and effective RPA operation sequence according to the user demand of each operation interface.

[0113] 3. RPA flow design agent based on large language model, prompt engineering and thinking chain technology. The agent can use thinking chain technology for in-depth thinking and reasoning, and generate accurate atomic operation sequence after multiple rounds of interaction. In addition, the application proposes a set of specification flow description for RPA component capability, which can effectively support the operation of complex business processes.

[0114] Figure 9 A robot process automation design device based on a large model is shown, and the device 90 includes:

[0115] ​The acquisition module 901 is configured to acquire service requirement information input by a user, obtain a subtask sequence and a type of a business field to which the service requirement information belongs based on the service requirement information, the first prompt word template, and the first large model, and the subtask sequence includes at least one subtask, and each subtask in the at least one subtask corresponds to a different business scenario interface respectively;

[0116] The execution module 902 is configured to obtain text state data corresponding to each business scenario interface respectively based on each business scenario interface.

[0117] Based on the subtask sequence, the text state data corresponding to each business scenario interface respectively, the type, the second prompt word template, and the second large model, an executable operation step sequence is obtained.

[0118] Based on the operation step sequence, RPA components corresponding to operation steps in the operation step sequence are sequentially called and executed in an order of the operation steps.

[0119] In a possible implementation, the execution module 902 is further configured to, before acquiring the service requirement information input by the user, obtaining the subtask sequence and the type of the business field to which the service requirement information belongs based on the service requirement information, the first prompt word template, and the first large model, fine-tune an initial large model based on sample service requirement information, a type of a field to which the sample service requirement information belongs, and a sample subtask sequence split based on the sample service requirement information, to obtain the first large model by using a model fine-tuning technology.

[0120] In a possible implementation, the execution module 902 is further configured to perform scene analysis on each business scenario interface respectively to obtain multi-modal information corresponding to each business scenario interface respectively.

[0121] The multi-modal information corresponding to each business scenario interface respectively is compressed and / or filtered to obtain processed multi-modal information corresponding to each business scenario interface respectively.

[0122] The processed multi-modal information corresponding to each business scenario interface respectively is converted to obtain text state data corresponding to each business scenario interface respectively.

[0123] In a possible implementation, the scene analysis on each business scenario interface respectively to obtain multi-modal information corresponding to each business scenario interface respectively includes at least two of the following:

[0124] Each business scenario interface is subjected to text detection by using an OCR algorithm to obtain text content;

[0125] Each business scenario interface is subjected to target detection by using a target detection algorithm to obtain a category of a first UI element and coordinate information of the first UI element.

[0126] The visual language model is used for target detection on each business scenario interface respectively to obtain the category of the second UI element and the coordinate information of the second UI element; wherein, the complexity of the second UI element is greater than the complexity of the first UI element.

[0127] The desktop information and the taskbar information of each business scenario interface are obtained.

[0128] In a possible implementation, the execution module 902 is further configured to obtain an executable operation step sequence based on the sub-task sequence, the text state data corresponding to each business scenario interface respectively, the type, the second prompt word template and the second large model, and in combination with the thinking chain technology.

[0129] In a possible implementation, the execution module 902 is further configured to define a correspondence table between the RPA component and the operation step in advance before sequentially calling and executing the RPA component corresponding to the operation step in the operation step sequence based on the operation step sequence and in the order of the operation steps in the operation step sequence.

[0130] The execution module 902 is further configured to sequentially call and execute the RPA component corresponding to the operation step in the operation step sequence based on the operation step sequence and the correspondence table and in the order of the operation steps in the operation step sequence.

[0131] In summary, the method shown in the embodiments of the present application deeply combines the RPA technology with the large model (such as the Llama, GPT, etc.) and the computer vision AI technology, and proposes an intelligent RPA process design system. The system has the ability of task understanding and intelligent decision-making, and can select the most suitable intelligent agent for process design according to the professional field to which the user demand belongs. Even in the face of business scenarios that have never appeared in the training process, the system can accurately understand the user's intention, complete the splitting of complex tasks, and has stronger expansibility. At the same time, by introducing the multi-modal scene understanding and conversion algorithm, the system can unify the data of various complex modalities on the screen into an easy-to-handle text modality, input it into the large model, enrich its context view, and improve the depth and breadth of the large model's understanding of the current business scenario, so as to better understand the business requirements, identify environmental changes, and generate more accurate and effective RPA operation processes.

[0132] The embodiments of the present application provide a network device 100, as shown in the figure. Figure 10 The network device 100 includes a processor 1001, a memory 1002, and a program stored in the memory 1002 and executable on the processor 1001, and the program is executed by the processor 1001 to implement the steps of the large model-based robot process automation design method shown in the above embodiments.

[0133] The embodiment of the application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0134] The embodiment of the application further provides a computer program product, which comprises computer instructions, and the computer instructions are executed by a processor to implement the steps of the above-mentioned large model-based robotic process automation design method and achieve the same technical effects.

[0135] It should be noted that in this document, the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of another identical element in the process, method, article, or device comprising the element.

[0136] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the methods described in various embodiments of the application.

[0137] The embodiments of the application are described above in combination with the drawings, but the application is not limited to the above specific embodiments, and the above specific embodiments are only illustrative, not restrictive, and those skilled in the art can make many forms under the inspiration of the application without departing from the scope of the application and the protection scope of the claims.

Claims

1. A robotic process automation design method based on a large model, characterized in that: The method comprises: Obtaining business requirement information input by a user, and obtaining a subtask sequence and a type of business domain to which the business requirement information belongs based on the business requirement information, a first prompt word template, and a first large model, wherein the subtask sequence includes at least one subtask, and each subtask in the at least one subtask corresponds to a different business scenario interface; Based on each of the business scenario interfaces, obtaining text data corresponding to each of the business scenario interfaces; Obtaining an executable operation step sequence based on the subtask sequence, the text data corresponding to each of the business scenario interfaces, the type, the second prompt word template, and the second large model; Based on the operation step sequence, the robotic process automation components corresponding to the operation steps are called and executed in sequence according to the order of the operation steps in the operation step sequence.

2. The method according to claim 1, characterized in that Before obtaining the business requirement information input by the user and obtaining the subtask sequence and the type of the business field to which the business requirement information belongs based on the business requirement information, the first prompt word template and the first large model, the method further includes: Based on the sample business demand information, the type of business field to which the sample business demand information belongs, and the sample subtask sequence split based on the sample business demand information, combined with model fine-tuning technology, the initial large model is fine-tuned to obtain the first large model.

3. The method according to claim 1, characterized in that Based on each of the business scenario interfaces, obtaining the text data corresponding to each of the business scenario interfaces includes: Performing scenario analysis on each of the business scenario interfaces to obtain multimodal information corresponding to each of the business scenario interfaces; Performing compression processing and / or filtering processing on the multimodal information corresponding to each of the business scenario interfaces to obtain processed multimodal information corresponding to each of the business scenario interfaces; The processed multimodal information corresponding to each of the business scenario interfaces is converted respectively to obtain textual data corresponding to each of the business scenario interfaces.

4. The method according to claim 3, characterized in that Scenario analysis is performed on each of the business scenario interfaces to obtain multimodal information corresponding to each of the business scenario interfaces, including at least two of the following: Using an optical character recognition (OCR) algorithm to perform text detection on each of the business scenario interfaces to obtain text content; Performing target detection on each of the business scenario interfaces using a target detection algorithm to obtain the category of the first user interface UI element and the coordinate information of the first UI element; Performing target detection on each of the business scenario interfaces using a visual language model to obtain a category of a second UI element and coordinate information of the second UI element; wherein the complexity of the second UI element is greater than the complexity of the first UI element; Obtain the desktop information and taskbar information of each business scenario interface.

5. The method according to claim 1, wherein Based on the subtask sequence, the text data corresponding to each of the business scenario interfaces, the type, the second prompt word template and the second large model, an executable operation step sequence is obtained, including: Based on the subtask sequence, the text data corresponding to each of the business scenario interfaces, the type, the second prompt word template and the second large model, combined with the thinking chain technology, an executable operation step sequence is obtained.

6. The method according to any one of claims 1 to 5, characterized in that Before sequentially calling and executing the robotic process automation components corresponding to the operation steps in the order of the operation steps in the operation step sequence based on the operation step sequence, the method further includes: Predefine the correspondence table between RPA components and operation steps; Based on the operation step sequence, in the order of the operation steps in the operation step sequence, sequentially calling and executing the robotic process automation components corresponding to the operation steps includes: Based on the operation step sequence and the correspondence table, the robotic process automation components corresponding to the operation steps are called and executed in sequence according to the order of the operation steps in the operation step sequence.

7. A large-scale model-based robotic process automation design device, characterized in that: The device comprises: an acquisition module configured to acquire business requirement information input by a user and, based on the business requirement information, a first prompt word template, and a first large model, obtain a subtask sequence and the type of business domain to which the business requirement information belongs, wherein the subtask sequence includes at least one subtask, and each subtask in the at least one subtask corresponds to a different business scenario interface; An execution module, configured to obtain text data corresponding to each of the business scenario interfaces based on each of the business scenario interfaces; Obtaining an executable operation step sequence based on the subtask sequence, the text data corresponding to each of the business scenario interfaces, the type, the second prompt word template, and the second large model; Based on the operation step sequence, the robotic process automation components corresponding to the operation steps are called and executed in sequence according to the order of the operation steps in the operation step sequence.

8. A network device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of a large model-based robotic process automation design method as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a large-model-based robotic process automation design method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of a large model-based robotic process automation design method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Information determination method and device and computer readable storage medium

    CN122045524A