RPA adaptive repair system based on bayesian inference and ai code generation
Patent Information
- Application Number
- CN202611153346.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-25
AI Technical Summary
[0003](1)人工维护成本高:目标网页发生任何UI(User Interface,用户界面)变化,如弹窗、DOM(Document Object Model,文档对象模型,将HTML/XML文档表示为树形结构的标准编程接口)结构调整和选择器失效等,RPA脚本即中断,运维人员需手动定位问题、修改脚本和重新部署
[0016]本申请过原子单元执行引擎实现 RPA 步骤解耦扩展,配合故障检测与上下文捕获模块收集异常消息、异常类型、异常发生时的页面截图和DOM快照等多维故障数据;依托贝叶斯诊断引擎融合异常类型与 DOM 弹窗特征开展双证据概率式根因判定,有效降低单一异常判定的误判概率、提升故障诊断效率,省去人工逐条排查日志的繁琐工作;再由 AI修复代码生成模块构造结构化提示词调用大模型输出修复代码,经安全沙箱受限权限热加载运行规避生成代码带来的生产环境安全隐患,最后通过先验概率在线更新模块利用指数移动平均迭代优化贝叶斯模型先验概率,让系统具备故障记忆自迭代能力,相较传统无记忆式 RPA 修复方案可实现诊断准确率持续迭代提升。本申请既实现了系统化的自动诊断、修复,无需依赖人工,降低人工成本,还整体提升了 RPA 流程异常自动恢复成功率,保障了修复执行全过程安全性,同时还具备良好流程拓展性。
Smart Images

Figure CN122816602A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary fields of robotic process automation and artificial intelligence, and in particular to an RPA adaptive repair system based on Bayesian inference and AI code generation. Background Technology
[0002] The operation and maintenance of current industrial RPA (Robotic Process Automation) systems face the following pain points:
[0003] (1) High manual maintenance costs: Any UI (User Interface) change to the target webpage, such as pop-ups, DOM (Document Object Model, a standard programming interface for representing HTML / XML documents as a tree structure) structure adjustments, or selector failures, will interrupt the RPA script. Maintenance personnel need to manually locate the problem, modify the script, and redeploy. A medium to large enterprise's RPA cluster often has hundreds of processes, and each process may fail weekly due to page changes, resulting in extremely high manual intervention costs.
[0004] (2) Fault diagnosis relies on experience: There are various types of anomalies (elements are obscured, selectors fail, page load times out, DOM structure changes, network errors), and there is a lack of systematic automatic diagnosis methods. Operation and maintenance personnel need to check the logs one by one, which results in low diagnostic efficiency.
[0005] (3) Security risks of repair code: Even if AI is introduced to generate repair scripts, there are security risks in directly executing the AI-generated code - AI may generate malicious or erroneous code containing file reading and writing, network calls, and system commands. Running it directly in the RPA execution environment may cause data leakage or system damage.
[0006] (4) Lack of adaptive learning mechanism: Traditional solutions handle each fault independently and do not accumulate experience. Even if the same type of pop-up appears repeatedly, the system cannot learn from historical repair experience, resulting in repeated consumption of AI call costs, and the diagnostic accuracy does not improve with running time.
[0007] The aforementioned technical problems urgently need to be solved. Summary of the Invention
[0008] To address the aforementioned technical problems, the purpose of this application is to provide an RPA adaptive repair system based on Bayesian inference and AI code generation, aiming to solve at least one of the technical problems mentioned above.
[0009] This application provides an RPA adaptive repair system based on Bayesian inference and AI code generation. The system includes:
[0010] The atomic unit execution engine is used to schedule the steps of the RPA process in sequence, encapsulating each step into an independent atomic unit, and each atomic unit executes independently and catches exceptions.
[0011] The fault detection and context capture module is used to collect fault context when an exception occurs; the fault context includes the current step name, exception message, exception type, page screenshot and DOM snapshot when the exception occurs;
[0012] The Bayesian diagnostic engine is used to determine DOM pop-up features based on the DOM snapshot; based on the Naive Bayes model, it integrates the anomaly type and DOM pop-up features, calculates the posterior probability of each preset root cause, and takes the root cause corresponding to the highest posterior probability as the diagnostic result.
[0013] The AI-powered repair code generation module generates structured prompts based on the fault context and diagnostic results, and then uses a large language model to generate repair code based on these structured prompts.
[0014] The security sandbox hot-load execution module is used to hot-load and execute the repair code in a security sandbox with restricted permissions to obtain the repair results;
[0015] The prior probability online update module is used to update the corresponding root based on the repair results using an exponential moving average.
[0016] This application utilizes an atomic unit execution engine to decouple and extend RPA steps. It works in conjunction with a fault detection and context capture module to collect multi-dimensional fault data, including exception messages, exception types, page screenshots at the time of exceptions, and DOM snapshots. A Bayesian diagnostic engine integrates exception types and DOM pop-up features to perform dual-evidence probabilistic root cause determination, effectively reducing the false positive probability of single-exception judgments, improving fault diagnosis efficiency, and eliminating the tedious work of manually checking logs one by one. An AI-generated repair code module constructs structured prompts to call a large model to output repair code. This code is then hot-loaded and run with restricted permissions in a secure sandbox to avoid security risks in the production environment. Finally, an online prior probability update module uses exponential moving averages to iteratively optimize the Bayesian model's prior probability, giving the system fault memory and self-iterative capabilities. Compared to traditional memoryless RPA repair solutions, this allows for continuous iterative improvement in diagnostic accuracy. This application achieves systematic automatic diagnosis and repair without relying on manual intervention, reducing labor costs, and improving the overall success rate of automatic recovery from RPA process anomalies, ensuring the security of the entire repair execution process, while also possessing good process scalability. Attached Figure Description
[0017] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of the structure of the RPA adaptive repair system based on Bayesian inference and AI code generation provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0020] Those skilled in the art will understand that, unless explicitly stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in the specification of this application means the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any modules and all combinations of one or more associated listed items.
[0021] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0022] Please see Figure 1 This application provides an RPA adaptive repair system based on Bayesian inference and AI code generation, the system comprising:
[0023] Atomic Unit Execution Engine 1 is used to schedule the steps of the RPA process in sequence, encapsulate each step into an independent atomic unit, and each atomic unit executes independently and catches exceptions.
[0024] Fault detection and context capture module 2 is used to collect fault context when an exception occurs; wherein, the fault context includes the current step name, exception message, exception type, page screenshot and DOM snapshot when the exception occurs;
[0025] Bayesian diagnostic engine 3 is used to determine DOM pop-up features based on the DOM snapshot; based on the Naive Bayes model, it integrates the anomaly type and DOM pop-up features, calculates the posterior probability of each preset root cause, and takes the root cause corresponding to the highest posterior probability as the diagnostic result.
[0026] AI repair code generation module 4 generates structured prompt words based on the fault context and diagnostic results, and generates repair code by calling a large language model based on the structured prompt words;
[0027] The security sandbox hot-loading execution module 5 is used to hot-load and execute the repair code in the security sandbox with restricted permissions to obtain the repair results;
[0028] The prior probability online update module 6 is used to update the prior probability of the corresponding root cause based on the repair results using an exponential moving average.
[0029] The atomic unit execution engine 1, fault detection and context capture module 2, Bayesian diagnostic engine 3, AI repair code generation module 4, safe sandbox hot loading execution module 5, and prior probability online update module 6 are sequentially connected to the atomic unit execution engine 1.
[0030] In Atomic Unit Execution Engine 1, each step is encapsulated as an independent atomic unit, and each atomic unit only needs to focus on the implementation of its own normal path. Adding a new RPA step requires no modification to the fault detection and context capture module, Bayesian diagnostic engine, AI repair code generation module, or prior probability online update module, thus improving the recovery success rate.
[0031] In the fault detection and context capture module 2, the exception type is identified by keyword matching based on the exception message. DOM (Document Object Model) is a standard programming interface for representing HTML / XML documents as a tree structure.
[0032] In the Bayesian diagnostic engine 3, it should be noted that RPA fault root causes are diverse (pop-up occlusion, selector failure, page not loading, DOM structure changes, network errors), making accurate judgment difficult based solely on anomaly information. This application's embodiment, based on a Naive Bayes model, integrates two types of evidence: anomaly type and DOM pop-up features. It calculates the posterior probability of each preset root cause and takes the root cause corresponding to the highest posterior probability as the diagnostic result, achieving probabilistic automatic root cause diagnosis. This solves the problems of existing methods that lack systematic automatic diagnosis due to diverse anomaly types, requiring maintenance personnel to check logs one by one, resulting in low diagnostic efficiency. Furthermore, the cross-validation of dual evidence (anomaly type and pop-up features) in this application reduces misjudgments caused by relying solely on anomaly type, improving diagnostic accuracy.
[0033] In the AI-based code generation module 4, the generation of structured prompts based on the fault context and diagnostic results includes: filling the fault context and diagnostic results into a preset prompt template to obtain the filled prompt template, which is the structured prompt. The prompt is the structured input text sent to the large language model to guide the model in generating the desired output.
[0034] Specifically, the diagnostic result is the root cause, which can be understood as the cause of the fault. The large language model mentioned is such as the DeepSeek application. The output repair code specifically outputs the repair code from the browser automation testing framework (Playwright). Playwright is an open-source browser automation testing framework from Microsoft that supports multiple browser engines such as Chromium, Firefox, and WebKit.
[0035] In the secure sandbox hot-loading execution module 5, the AI-generated fix code may contain malicious or unexpected behaviors (such as file reading and writing, network calls, and system command execution), posing significant security risks if executed directly in a production environment. This application reduces security risks such as data leakage or system damage by hot-loading and executing the fix code with restricted permissions in a secure sandbox.
[0036] In the prior probability online update module 6, based on the repair results, the prior probability of the corresponding root cause is updated using an exponential moving average, thereby continuously improving the diagnostic accuracy with the number of runs. It should be noted that traditional RPA adaptive repair schemes have no memory capability; each diagnosis starts from zero and does not improve accuracy with the number of runs. In contrast, this application, after each repair attempt is completed—that is, after the repair code has finished executing and the repair result is obtained—updates the prior probability of the corresponding root cause using an exponential moving average based on the repair result, thus continuously improving the diagnostic accuracy with the number of runs.
[0037] In summary, this application achieves decoupling and extension of RPA steps through an atomic unit execution engine, and collects multi-dimensional fault data such as abnormal messages, abnormal types, page screenshots and DOM snapshots at the time of the abnormality in conjunction with a fault detection and context capture module. It relies on a Bayesian diagnostic engine to integrate abnormal types and DOM pop-up features to conduct dual-evidence probabilistic root cause determination, effectively reducing the false positive probability of single abnormality determination, improving fault diagnosis efficiency, and eliminating the tedious work of manually checking logs one by one. Furthermore, an AI-powered repair code generation module constructs structured prompts to call a large model to output repair code. This code is then hot-loaded and run with restricted permissions in a secure sandbox to avoid security risks in the production environment caused by the generated code. Finally, an online prior probability update module uses exponential moving average to iteratively optimize the prior probability of the Bayesian model, giving the system fault memory and self-iterative capabilities. Compared to traditional memoryless RPA repair solutions, this allows for continuous iterative improvement in diagnostic accuracy. This application achieves systematic automatic diagnosis and repair without relying on manual labor, reducing labor costs, and improving the overall success rate of automatic recovery from RPA process anomalies, ensuring the security of the entire repair execution process, while also possessing good process scalability.
[0038] In one embodiment, a standard two-layer adaptive repair loop is implemented using a PRA executor: the outer loop sequentially traverses the list of atomic units, initializing the inner loop upon entering each unit, and executing the `execute` method of the current unit in each inner iteration; if the result is successful, the output data is stored in the context's accumulated data container and the inner loop is exited to proceed to the next unit; if the result is unsuccessful, four recovery sub-stages are sequentially entered, namely, the fault detection and context capture module, the Bayesian diagnostic engine, the AI repair code generation module, and the secure sandbox hot-loading execution module. After the repair is completed, the prior probability update method is called to update the prior probability of the corresponding root cause based on the returned success / failure boolean value (i.e., entering the prior probability online update module). If the repair is successful, the process returns to the starting point of the inner loop to retry the current unit; if the repair fails, the retry process of the current unit is terminated.
[0039] The process of encapsulating each step into an independent atomic unit includes:
[0040] Define an abstract base class for atomic units as a unified contract template for all RPA steps, forcing each step to implement a unified execution method signature. The unified execution method signature is: receive a browser page object and a unit execution context, and return a unified execution result. The execution result is designed using an algebraic data type and has two mutually exclusive states: a success state carrying a business data dictionary, and a failure state carrying the aforementioned fault context. The unit execution context is a thread-safe data container that sequentially transmits three parts of information between RPA steps, including the cumulative business data container produced by the executed steps, the zero-based index of the current step, and the step operation timeout threshold.
[0041] In this embodiment, each atomic unit only needs to focus on implementing its own normal business path (locating elements, filling data, clicking buttons, and waiting for redirection). Exceptions are uniformly captured by the architecture layer and packaged as execution results in a failed state (the success flag in the result object is false, and the fault context field is not null). Exception handling, diagnosis, repair, and retry logic are all centrally managed in the executor layer. Adding a new RPA step only requires writing a new class that inherits from the atomic unit abstract base class and implements the execution method. Each step is retried a maximum of K times (controlled by the maximum retry count configuration item). Due to the randomness of the large language model, each round of retry may generate different repair strategies (such as clicking the close button on the first attempt and pressing the Escape key on the second attempt), improving the recovery success rate.
[0042] In one embodiment, before calculating the posterior probability of each preset root cause based on a Naive Bayes model, by fusing anomaly type and DOM pop-up features, the Bayesian diagnostic engine is further configured to:
[0043] Predefine N mutually exclusive root causes;
[0044] Two types of observable evidence are defined, including anomaly types and DOM pop-up features; DOM pop-up features are obtained by performing regular expression pattern matching on DOM snapshots to detect whether they contain target keywords; pop-up features are represented by Boolean values.
[0045] Construct and store a prior probability configuration file; the prior probability configuration file includes a root cause prior probability table, an anomaly type conditional probability table, and a DOM pop-up conditional probability table; the root cause prior probability table is used to characterize the prior probability of each root cause; the anomaly type conditional probability table is used to characterize the probability of different anomaly types occurring under various root causes; the DOM pop-up conditional probability table is used to characterize the probability of DOM pop-up features being true under various root causes.
[0046] In this embodiment, specifically, N=5, and the five root causes are pop-up occlusion, selector failure, page not loaded, DOM structure change, and network error; the five anomaly types are: element not interactive, timeout, element not found, element separated from document, and unknown. Target keywords include modal, pop-up, overlay, dialog box, cookie banner, and banner; the anomaly type conditional probability table is a 5x5 matrix, with rows corresponding to the five anomaly types and columns corresponding to the five root causes. The DOM pop-up conditional probability table is a 5x1 vector, with rows corresponding to the five root causes.
[0047] This application provides a data foundation for calculating the posterior probability of each preset root cause based on the Naive Bayes model, by predefining N mutually exclusive root causes, defining two types of observable evidence, and constructing and storing a prior probability configuration file.
[0048] In one embodiment, the step of calculating the posterior probability of each preset root cause based on the Naive Bayes model, fusing anomaly type and DOM pop-up features, includes:
[0049] When an anomaly occurs, based on the anomaly type and DOM pop-up characteristics, the scores of each preset root cause are calculated according to formula (1):
[0050] (1)
[0051] Where C is the score of a single root cause. It is the prior probability of the root cause. It is the conditional probability of the anomaly type, P(E1|C). It is the conditional probability of the DOM pop-up; where, when E2 is false, P(E2|C) is 1 minus the conditional probability value;
[0052] The scores of each preset root cause are added together to obtain the total score;
[0053] Divide the score of each preset root cause by the total score to obtain the posterior probability of each preset root cause.
[0054] In this embodiment, it should be noted that E2 is false (the boolean value of false is 0), meaning that the target keyword was not detected in the DOM snapshot. The probability corresponding to the "non-pop-up" condition is obtained by subtracting this conditional probability value from 1 in P(E2|C).
[0055] For example, if an element is not interactive and a modal keyword appears in the DOM, the posterior probability of the root cause of pop-up occlusion is 0.35×0.70×0.85÷sum of scores of each root cause≈0.80, which is significantly higher than other hypotheses. Therefore, pop-up occlusion is the diagnostic root cause, and the diagnosis result is pop-up occlusion.
[0056] The probability calculation in this embodiment only involves multiplication and division, and the inference time is <1ms, which does not affect the performance of RPA execution.
[0057] In this embodiment, if a certain type of pop-up appears repeatedly and the corresponding repair method remains effective (such as closing the pop-up), its prior probability will gradually increase, and the system will be able to locate the root cause more quickly. If a repair method fails frequently, the prior probability of its corresponding root cause will decrease, and the system will automatically reduce the confidence level of the diagnosis result. This mechanism enables the system to automatically adapt to different target websites and different usage scenarios without manual parameter tuning.
[0058] In one embodiment, during the process of generating repair code by calling the large language model based on the structured prompts, the system prompts in the structured prompts restrict the large language model to use only a specified whitelist of methods of the browser's automated testing framework synchronous application programming interface. The specified method whitelist includes one or more of the following: element location methods, click methods, fill methods, visibility judgment methods, selector wait methods, timeout wait methods, keyboard key methods, page navigation methods, and loading status wait methods. It also explicitly prohibits import statements, file input / output, network calls, class / function definitions, and restricts the output format to plain code text wrapped in Python code block tags, without any explanatory text.
[0059] In Python, the `import` statement is used to import code from other modules (i.e., files containing Python code) into the current script or interactive interpreter. It also restricts the output format, forcing it to be plain code text wrapped in Python code block tags, without any explanatory text, making it easier to extract.
[0060] It should be noted that AI-generated repair code may contain malicious or unexpected behaviors (file reading / writing, network calls, system command execution). To address this issue, this application's embodiments reduce security risks during the generation of repair code using a large language model (i.e., AI) by employing structured prompts. Specifically, by pre-defining a whitelist of browser automation testing framework synchronous API (Application Programming Interface) calls, disabling high-risk syntax and violations, and constraining code output format within the constructed structured prompts, the output behavior of the large language model can be constrained from the source of code generation. This proactively avoids AI-generated code snippets containing security vulnerabilities such as module imports, file reading / writing, external network requests, and custom class functions, thus preemptively curbing potential security risks such as unauthorized access to local files, leakage of sensitive page data, and execution of illegal system calls. This achieves pre-emptive security control over AI-generated repair code, improving the overall security and controllability of the RPA automatic fault repair process.
[0061] The prior probability of updating the corresponding root cause using exponential moving average based on the repair results includes:
[0062] Update the prior probability of diagnosing the root cause using formula (2).
[0063] (2)
[0064] Use formulas (3), (4) and (5) to update the prior probabilities of other root causes;
[0065] (3)
[0066] in, 1; (4)
[0067] (5)
[0068] in, To diagnose the root cause, that is, the diagnostic result, To diagnose the current prior probability of the root cause, The prior probability after the root cause update is used for diagnosis. The outcome is a binary result indicating whether the repair was successful, i.e., the repair result, with 1.0 for success and 0.0 for failure. The learning rate; Let i be the updated prior probability of root cause i. Let be the current prior probability of root cause i; n is the total number of other root causes.
[0069] The updated prior probabilities of each root cause are persisted to the prior probability configuration file.
[0070] The embodiments of this application are based on the prior probability update mechanism of exponential moving average, which enables the system to automatically adjust the prior weight of each root cause according to the repair success rate during continuous operation, and the diagnostic accuracy converges and improves with the number of runs.
[0071] In the embodiments of this application, The default value is 0.1.
[0072] In one embodiment, the step of hot-loading and executing the fix code with restricted permissions in a secure sandbox includes:
[0073] The code for fixing is subjected to syntax compilation verification, restricted global namespace injection, and thread pool timeout control.
[0074] In this application's embodiments, the AI-generated repair code may contain malicious or unexpected behaviors, such as file reading and writing, network calls, or system command execution, posing significant security risks if executed directly in a production environment. To address this technical problem, this application performs syntax compilation verification, restricted global namespace injection, and thread pool timeout control on the AI-generated code to ensure its secure execution. The syntax compilation verification of the repair code includes: using a compilation function to perform a syntax validity check on the repair code; if the repair code contains syntax errors or is invalid, it is rejected from entering the execution phase; wherein, the compilation function receives a code string, a source identifier, and an execution mode parameter. Specifically, the built-in compilation function of the Python language is used to perform a syntax validity check on the code generated for the repair code.
[0075] In one embodiment, the steps of performing syntax compilation verification, restricted global namespace injection, and thread pool timeout control on the repair code include:
[0076] Construct a restricted global namespace; wherein the restricted global namespace does not contain dangerous built-in functions, any standard library and third-party module references, and exposes basic data type constructors, safe built-in functions, exception types and browser page objects in a whitelist manner;
[0077] The repair code is wrapped into a repair function;
[0078] The repair function is executed in an independent thread within the restricted global namespace. If the repair function does not return a result within a preset time, the control thread is interrupted and the repair is deemed to have failed.
[0079] In this embodiment, basic data type functions include constructors for integers, strings, floating-point numbers, booleans, lists, dictionaries, and tuples. Safe built-in functions include the length calculation function `len`, the range generation function `rang`, and the printing function `print`. Exception types include basic exception classes and timeout exception classes. Dangerous built-in functions include the block module import built-in function `__import__`, the file opening function `open`, and the code execution functions `eval`, `exec`, and `compile`. Specifically, by setting a restricted global namespace to exclude the block module import built-in function `__import__`, import statements can be blocked; by setting a restricted global namespace to exclude the file opening function `open`, file access can be blocked; and by setting a namespace to exclude the code execution functions `eval`, `exec`, and `compile`, code injection chains can be blocked. The repair function is executed in a separate thread pool. In one embodiment, a default timeout of 15 seconds is set. Specifically, the asynchronous task is submitted through the thread pool executor (ThreadPoolExecutor) provided by the Python standard library's `concurrent.futures` module. The `result()` method is called on the returned `Future` object with a 15-second timeout parameter. If the repair function does not return a result within 15 seconds, the thread is automatically interrupted and the repair is considered to have failed. In the `concurrent.futures` module, the `Future` object is the object used in process pools and thread pools to implement asynchronous operations. The `result()` method in the `concurrent.futures` module is used to retrieve the result of the asynchronous execution from the `Future` object.
[0080] In one test, the AI-generated code "import the time module and sleep for 1 second" compiled successfully, but during the sandbox runtime, a NameError exception was triggered because the built-in function __import__ for module import was removed. The repair function returned failure, and the entire repair process terminated normally without blocking the system. This test result indicates that the sandbox only grants the minimum permissions to operate on the browser page, limiting the attack surface to the current page within the browser sandbox and preventing penetration to the operating system layer.
[0081] This application employs a three-tiered defense system: structured prompt constraints, syntax compilation verification, and restricted global namespaces. This system leverages the large-scale code generation capabilities of the model while eliminating security risks such as code injection, file leakage, and system corruption.
[0082] In one embodiment, the system further includes a chaos engineering injection module, used for:
[0083] Register the Chaos Injector component as a request interception hook in the Flask framework. Before the Hypertext Transfer Protocol (HTTP) response is returned, insert a random pop-up structure before the closing tag of the Hypertext Markup Language (HML) response body with a configurable probability. The random pop-up structure includes N types of pop-ups, which are randomly selected each time. Both the pop-ups and their close buttons are defined using semantic CSS class names.
[0084] In this embodiment, the ChaosInjector component is registered as a request-after-request hook in the Flask framework, and a random pop-up structure is inserted before the closing tag ( ) of the Hypertext Transfer Protocol (HTTP) response body with a configurable probability (default 30%) before the HTTP response is returned. The random pop-up structure contains 5 pop-up variants (equivalent to pop-ups), which are randomly selected each time: (1) Full-screen ad overlay (CSS class name is ad-modal, containing a close ad button, whose CSS class name is close-ad); (2) Bottom Cookie consent banner (CSS class name is cookie-banner, containing an accept cookie button, whose CSS class name is accept-cookies); (3) Centered questionnaire pop-up (CSS class name is survey-popup, containing a reminder button, whose CSS class name is dismiss-survey); (4) Newsletter subscription modal (CSS class name is newsletter-modal, containing a "No, thank you" button, whose CSS class name is btn-no-thanks); (5) Full-screen semi-transparent overlay (CSS class name is fullscreen-overlay, containing a close overlay button, whose CSS class name is dismiss-overlay). Each popup uses a combination of Cascading Style Sheets (CSS) properties: fixed positioning (position value is fixed) and top stacking order (z-index value is 9999; the z-index property sets the position of a positioned element along the z-axis, which is defined as the axis extending vertically to the display area. A positive number means closer to the user, and a negative number means farther away). This ensures the popup completely obscures the form elements. The close buttons for the popups use semantic CSS class names, allowing AI to read these class names from the DOM snapshot and generate targeted fix code. It should be understood that CSS (Cascading Style Sheets) is a style language used to describe the appearance and formatting of HTML documents. Flask is a lightweight web application framework written in Python. A full-screen ad overlay is a modal popup used to display advertising content. A cookie consent banner is a notification bar on a webpage prompting users to accept or manage cookie usage policies. A survey popup is a popup used to invite users to participate in a survey. A newsletter subscription modal is a modal window used to invite users to subscribe to an email newsletter. A full-screen semi-transparent overlay is a semi-transparent mask that occupies the entire viewport.
[0085] It should be noted that testing RPA adaptive repair systems requires simulating various real-world page anomalies. Traditional methods necessitate manually writing a large amount of stub code or modifying the target website's source code. However, the embodiments in this application do not require modification of the target website's source code or deployment of additional interception proxies; fault injection can be achieved solely through Flask middleware. Configurable probabilities (controlled by the popup_inject_probability configuration parameter) can simulate page interference at different frequencies. The semantic design of popup class names reduces the difficulty of AI inference and improves the accuracy of repair code generation. Furthermore, configurable probabilities of random popup injection are achieved through a Flask middleware-level request-after-request interception hook, providing a low-cost and high-coverage testing method for RPA adaptive repair systems.
[0086] In one embodiment, the system further includes a full-link observability module for:
[0087] For each large model call, a separate JSON file is generated; the JSON file includes a summary of system prompt words, a summary of user prompt words, the model name, the response length, the time taken, and the success or failure status of the fix code generation.
[0088] Store the posterior probabilities of all root causes obtained from each Bayesian diagnostic operation;
[0089] When the prior probability configuration file is updated, retain the audit trail of each round of prior probability iteration modification;
[0090] Output a structured operation log, which includes the step name, success or failure status, step time, fault diagnosis conclusion, repair summary, and repair result.
[0091] In this embodiment, the user prompt word summary refers to the summary of prompt words other than system prompt words in the structured prompt words. The audit trajectory of the prior probability iterative modification is the file-level version record of the prior probability configuration file. This embodiment facilitates post-event analysis of AI repair quality and diagnostic accuracy, tracks the convergence process of prior probabilities, and constructs feedback data for future supervised fine-tuning of the model.
[0092] To facilitate understanding of the differences between this invention and traditional RPA systems, Table 1 below shows the differences between traditional RPA systems and this invention:
[0093] Table 1
[0094]
[0095] The following is an in-depth comparison between traditional RPA systems and the present invention:
[0096] (1) The essential difference between this invention and traditional RPA systems:
[0097] Traditional RPA hardcodes automation logic into scripts. Taking a 5-step form submission process as an example, traditional RPA requires explicitly writing a linear sequence of instructions in each step: navigating to the target URL (Uniform Resource Locator), finding and filling in the form fields, clicking the submit button, and asserting successful page redirection. After step 1 is completed, step 2 is executed, and so on. All operation paths are fixed at compile time.
[0098] When an unexpected pop-up appears on the page, the click operation on the "Next" button in step 2 throws an exception because the button is obscured, interrupting the entire process. All subsequent steps (3, 4, 5) are rendered useless, and all filled form data is lost. Maintenance personnel must manually reproduce the problem, insert pop-up closing logic into the script, redeploy, and re-execute. An exception in one step causes the entire process to be interrupted, wasting time and resources.
[0099] The core difference of this invention is that the normal path is executed with zero overhead, and the abnormal path is repaired by calling AI as needed. After the repair is successful, the process continues from the current step.
[0100] During normal execution (without pop-ups), atomic units directly manipulate the DOM, resulting in performance consistent with traditional RPA and zero additional latency. Fault diagnosis → AI repair → retry process is only triggered when the browser's automated testing framework throws an exception.
[0101] After a successful repair, only the currently failed step will be retried, without having to start over from step 1. The already executed business context (accumulated business data container accumulated_data) will be fully preserved.
[0102] (2) Comparison of Token Costs between the Invention and the “RPA + AI Stepwise Verification” Scheme
[0103] In recent years, a solution has emerged that embeds AI validation into each step of RPA: after each step, a multimodal large model is called to perform visual validation on the page screenshot (to determine whether the page has been redirected to the correct page, whether the form fields have been filled in correctly, etc.), and only after confirming that there are no errors will the next step be performed.
[0104] The token consumption pattern for this type of scheme is as follows: token consumption is proportional to the product of the number of process steps and the token cost for each step verification.
[0105] Taking a 5-step process as an example, assuming each visual verification consumes approximately 500-2000 tokens (screenshot description + judgment prompts + response), and a complete execution requires at least 5 AI calls, consuming a total of 2500-10000 tokens. For RPA tasks running hundreds of times daily, daily token consumption can reach millions.
[0106] The token consumption pattern of this invention is as follows: token consumption is proportional to the product of the number of failures and the token cost of a single repair.
[0107] The core optimization involves using a Bayesian diagnostic engine to calculate the posterior probability of each preset root cause when a fault occurs, thereby obtaining a diagnostic result (i.e., the root cause diagnosis). This calculation is purely mathematical, with an inference time of <1ms and 0 tokens. Subsequently, the AI-based code generation module is called, and the prompts carry structured diagnostic results (such as the root cause diagnosis being obscured by a pop-up window) instead of the original screenshot file, significantly reducing the size of the prompts.
[0108] Table 2
[0109]
[0110] 5A.3 Token Cost Quantification Analysis
[0111] Take a typical 5-step RPA process, executing 100 rounds per day, and encountering pop-up windows in 30% of the rounds as an example:
[0112] RPA+AI step-by-step verification solution: AI verification calls: 100 rounds × 5 steps = 500 times / day. Token consumption: 500 × 1000 (median) = 500,000 tokens / day. Even with 100% successful execution, an equal amount of tokens will still be consumed (verification even without failures).
[0113] This invention's solution: Only 30 rounds of fault triggering (30% probability) × an average of 1.2 AI calls (first-time repair success rate approximately 80%, some requiring secondary repairs) ≈ 36 AI calls / day. Each repair consumes approximately 500 tokens. Token consumption: 36 × 500 = 18,000 tokens / day - Token saving rate ≈ 96.4%
[0114] If the failure probability is reduced to 10% (the target website is relatively stable), the daily token consumption of this invention is reduced to about 3,000 tokens, while the gradual verification scheme still consumes a fixed 500,000 tokens / day, resulting in a saving rate of 99.4%.
[0115] Key findings: The token cost of this invention is directly proportional to the failure rate, while the token cost of the step-by-step verification scheme is directly proportional to the number of steps multiplied by the number of execution rounds. In a real-world RPA production environment, page anomalies are low-probability events (typically occurring in <10% of rounds). This invention enables on-demand triggering of token consumption, offering an order-of-magnitude cost advantage compared to the step-by-step verification scheme.
[0116] The following example illustrates how the system works using a complete execution process:
[0117] Step 1 (System Startup): The main entry module starts the demo site (including the 5-step form page) built with the Flask lightweight web framework, and registers the chaos injector component to Flask's request after-request hook, configuring the pop-up injection probability to 30%. The demo site runs in a background thread.
[0118] Step 2 (Initialize Browser and Execution Engine): The RPA executor starts the Chromium browser kernel, creates a browser page object (Playwright Page), loads four atomic units in sequence (name entry, email entry, address entry, and confirmation submission), and prepares the execution context for each step (including the cumulative business data container, step index, and timeout configuration).
[0119] Step 3 (Pop-up Trigger): The executor proceeds to Step 2 (Email Entry Page). When the HTTP request returns the page HTML, the Chaos Injector component hits with a 30% probability, injecting a Cookie consent banner pop-up (CSS class name: cookie-banner) before the closing tag of the response body. This pop-up uses fixed positioning (position:fixed) and a stacking order (z-index) of 9999, completely covering the form submit button.
[0120] Step 4 (Exception Handling): The atomic unit execution method in Step 2 attempts to perform a click operation on the "Next" button (Cascading Style Sheet selector "#next-btn"). The browser automation engine throws a timeout exception with the message "Element is not interactive". The atomic unit's exception handling mechanism wraps this exception into a unified execution result, where the success flag is false, and the failure context field contains the exception type (initially marked "unknown") and the exception message.
[0121] Step 5 (Fault Detection and Context Collection): Take a screenshot of the current page to obtain Base64 (a method of encoding binary data into ASCII text characters, commonly used to transmit binary data in text protocols) encoded PNG (Portable Network Graphics, a lossless compressed bitmap image format); simultaneously, obtain a DOM snapshot and extract the first 5000 characters as an analysis sample. Perform keyword matching on the abnormal message, identify keywords meaning "interactive," and correct the abnormal type from "unknown" to "non-interactive element."
[0122] Step 6 (Bayesian Diagnosis): The Bayesian diagnostic engine calls the pop-up indicator detection method of the prior probability storage object, performs a regular expression scan on the DOM snapshot, detects the keyword "cookie-banner", and returns that the pop-up feature is true. Then, the posterior probability calculation method is invoked, with the input evidence being (anomaly type = non-interactive element, pop-up feature = true). The unnormalized scores for the five root causes are calculated: the pop-up occlusion root cause score is calculated as: prior 0.35 multiplied by the conditional probability of non-interactive element 0.70, then multiplied by the conditional probability of pop-up occlusion 0.85, equaling 0.2083; the selector failure root cause score is 0.25 × 0.15 × 0.05 = 0.0019; the page not loading root cause score is 0.20 × 0.10 × 0.05 = 0.0010; the DOM structure change root cause score is 0.15 × 0.03 × 0.03 ≈ 0.0001; and the network error root cause score is 0.05 × 0.02 × 0.02 ≈ 0.00002. The sum of the five scores is approximately 0.2113, and the normalized posterior probability of the pop-up occlusion root cause is approximately 0.985. The Bayesian diagnostic engine returns the diagnostic result (i.e., the root cause of the diagnosis). The root cause of the diagnosis is obscured by the pop-up window, and the corresponding posterior probability is 0.985.
[0123] Step 7 (AI Repair Code Generation): Populate the fault context (including the current step name, page screenshot, exception type, exception message, and DOM snapshot) and diagnostic root cause into the preset Prompt template, and call the DeepSeek large language model interface (temperature parameter set to 0.1, low temperature ensures deterministic output). AI returns a piece of Playwright repair code, the logic of which is: locate the "Accept Cookies" button (class name accept-cookies) through the Cascading Style Sheets class selector, determine whether the button is visible on the page, if visible, perform the click operation and return success (indicating that the repair has been applied), otherwise return failure.
[0124] Step 8 (Syntax Compilation and Verification): Extract the pure code content from the repair code. The code wrapping method wraps this code into a standard repair function. The system's compilation function performs a syntax check on the complete code, and only after passing the verification can it enter the execution phase.
[0125] Step 9 (Secure Sandbox Execution): The wrapped code is executed within a restricted global namespace. This restricted namespace does not contain dangerous built-in functions such as file opening, module import, or code execution; it only exposes the browser page object and necessary secure built-in functions. During execution, the "Accept Cookies" button (class name accept-cookies) is indeed present and visible in the Document Object Model (DOM). Clicking it successfully triggers the action, the cookie banner pop-up closes, and the repair function returns a success. The entire process is executed in a separate thread and is protected by a 15-second timeout.
[0126] Step 10 (Online Update of Prior Probabilities): The update operation targets the root cause occlusion via a pop-up window, achieves successful repair, and sets the learning rate α to 0.1. The prior probability of this root cause is updated from 0.35 to 0.35×(1-0.1)+1.0×0.1=0.415. The prior values of the other four root causes are scaled proportionally (multiplied by a scaling factor (1-0.415) / (1-0.35)≈0.90) to ensure the sum of the five priors remains 1. The updated complete prior probability table is persistently written to the prior probability configuration file.
[0127] Step 11 (Retry Recovery): The RPA executor returns to the beginning of the inner loop and retryes Step 2. The chaos injector component was not hit when the HTTP response returned in this round (the injection of the after_request hook is a probabilistic event), the page renders normally, and there are no pop-up obstructions. The atomic unit of Step 2 is re-executed: the email input box is located, test data is entered, the next button is clicked, and the page is waited to redirect to Step 3. Execution is successful, and a success result and form data are returned. The executor writes the output data of Step 2 into the cumulative data container of the context and continues to execute Step 3.
[0128] Step 12 (Execution Complete): After all 4 atomic units have been executed, the browser is closed, and the main entry module prints a summary of the run (such as the number of successful steps, the number of failed steps, and the exception type of each failed step) to the standard output (the program's default output target stream, usually pointing to the terminal or console).
[0129] It should be noted that in the examples above, Chromium is a cross-platform open-source web browser developed under the leadership of Google. The demo site is an example site for verifying the system.
[0130] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0132] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An RPA adaptive repair system based on Bayesian inference and AI code generation, characterized in that, The system includes: The atomic unit execution engine is used to schedule the steps of the RPA process in sequence, encapsulating each step into an independent atomic unit, and each atomic unit executes independently and catches exceptions. The fault detection and context capture module is used to collect fault context when an exception occurs; the fault context includes the current step name, exception message, exception type, page screenshot and DOM snapshot when the exception occurs; The Bayesian diagnostic engine is used to determine DOM pop-up features based on the DOM snapshot; based on the Naive Bayes model, it integrates the anomaly type and DOM pop-up features, calculates the posterior probability of each preset root cause, and takes the root cause corresponding to the highest posterior probability as the diagnostic result. The AI-powered repair code generation module generates structured prompts based on the fault context and diagnostic results, and then uses a large language model to generate repair code based on these structured prompts. The security sandbox hot-load execution module is used to hot-load and execute the repair code in a security sandbox with restricted permissions to obtain the repair results; The prior probability online update module is used to update the prior probability of the corresponding root cause based on the repair results using an exponential moving average.
2. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, The process of encapsulating each step into an independent atomic unit includes: Define an abstract base class for atomic units as a unified contract template for all RPA steps, forcing each step to implement a unified execution method signature. The unified execution method signature is: receive a browser page object and a unit execution context, and return a unified execution result. The execution result is designed using an algebraic data type and has two mutually exclusive states: a success state carrying a business data dictionary, and a failure state carrying the aforementioned fault context. The unit execution context is a thread-safe data container that sequentially transmits three parts of information between RPA steps, including the cumulative business data container produced by the executed steps, the zero-based index of the current step, and the step operation timeout threshold.
3. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, Before calculating the posterior probability of each preset root cause based on the Naive Bayes model, by fusing anomaly type and DOM pop-up features, the Bayesian diagnostic engine is also used for: Predefine N mutually exclusive root causes; Two types of observable evidence are defined, including anomaly types and DOM pop-up features; DOM pop-up features are obtained by performing regular expression pattern matching on DOM snapshots to detect whether they contain target keywords; pop-up features are represented by Boolean values. Construct and store a prior probability configuration file; the prior probability configuration file includes a root cause prior probability table, an anomaly type conditional probability table, and a DOM pop-up conditional probability table; the root cause prior probability table is used to characterize the prior probability of each root cause; the anomaly type conditional probability table is used to characterize the probability of different anomaly types occurring under various root causes; the DOM pop-up conditional probability table is used to characterize the probability of DOM pop-up features being true under various root causes.
4. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 3, characterized in that, The calculation of the posterior probability of each preset root cause based on the Naive Bayes model, which integrates anomaly type and DOM pop-up features, includes: When an exception occurs, based on the exception type and DOM pop-up characteristics, according to the formula... Calculate the score for each preset root cause; where C is the score for a single root cause. It is the prior probability of the root cause. It is the conditional probability of the anomaly type, P(E1|C). It is the conditional probability of the DOM pop-up; where, when E2 is false, P(E2|C) is 1 minus the conditional probability value; The scores of each preset root cause are added together to obtain the total score; Divide the score of each preset root cause by the total score to obtain the posterior probability of each preset root cause.
5. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, During the process of generating repair code by calling the large language model based on the structured prompts, the system prompts in the structured prompts restrict the large language model to use only the specified whitelist of methods of the browser's automated testing framework synchronous application programming interface. The specified method whitelist includes one or more of the following: element location methods, click methods, fill methods, visibility judgment methods, selector wait methods, timeout wait methods, keyboard key methods, page navigation methods, and loading status wait methods. It also explicitly prohibits import statements, file input / output, network calls, class / function definitions, and restricts the output format to plain code text wrapped in Python code block tags, without any explanatory text.
6. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, The prior probability of updating the corresponding root cause using exponential moving average based on the repair results includes: Update the prior probability of diagnosing the root cause using formula (2). ;(2) Use formulas (3), (4) and (5) to update the prior probabilities of other root causes; ;(3) in, 1; (4) ;(5) in, To diagnose the root cause, that is, the diagnostic result, To diagnose the current prior probability of the root cause, The prior probability after the root cause update is used for diagnosis. The outcome is a binary result indicating whether the repair was successful, i.e., the repair result, with 1.0 for success and 0.0 for failure. The learning rate; Let i be the updated prior probability of root cause i. Let be the current prior probability of root cause i; n is the total number of other root causes. The updated prior probabilities of each root cause are persisted to the prior probability configuration file.
7. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, The step of hot-loading and executing the fix code with restricted privileges in a secure sandbox includes: The code for fixing is subjected to syntax compilation verification, restricted global namespace injection, and thread pool timeout control.
8. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 7, characterized in that, The syntax compilation verification, restricted global namespace injection, and thread pool timeout control of the repair code include: Construct a restricted global namespace; wherein the restricted global namespace does not contain dangerous built-in functions, any standard library and third-party module references, and exposes basic data type constructors, safe built-in functions, exception types and browser page objects in a whitelist manner; The repair code is wrapped into a repair function; The repair function is executed in an independent thread within the restricted global namespace. If the repair function does not return a result within a preset time, the control thread is interrupted and the repair is deemed to have failed.
9. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, The system also includes a chaos engineering injection module, used for: Register the Chaos Injector component as a request interception hook in the Flask framework. Before the Hypertext Transfer Protocol (HTTP) response is returned, insert a random pop-up structure before the closing tag of the Hypertext Markup Language (HML) response body with a configurable probability. The random pop-up structure includes N types of pop-ups, which are randomly selected each time. Both the pop-ups and their close buttons are defined using semantic CSS class names.
10. The RPA adaptive repair system based on Bayesian inference and AI code generation according to claim 1, characterized in that, The system also includes a full-link observability module for: For each large model call, a separate JSON file is generated; the JSON file includes a summary of system prompt words, a summary of user prompt words, the model name, the response length, the time taken, and the success or failure status of the fix code generation. Store the posterior probabilities of all root causes obtained from each Bayesian diagnostic operation; When the prior probability configuration file is updated, retain the audit trail of each round of prior probability iteration modification; Output a structured operation log, which includes the step name, success or failure status, step time, fault diagnosis conclusion, repair summary, and repair result.