Automatic process execution method based on large language model

Through the process automation method based on large language models, the problem of low efficiency of traditional tools in understanding natural language and task decomposition is solved, the transformation and dynamic optimization of unstructured instructions to structured tasks are realized, the flexibility and intelligence of the system are enhanced, and the cost of enterprise system integration is reduced.

CN120780437APending Publication Date: 2025-10-14SUZHOU HAIGUANJIA LOGISTICS TECH CO LTD

Patent Information

Application Number
CN202511171262.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional process automation tools cannot effectively understand natural language instructions, have low task decomposition and execution efficiency, suffer from serious data silo problems, lack flexibility and intelligence, and are unable to adapt to business changes in real time.

Method used

It adopts a process automation method based on a large language model, parses user instructions through a multimodal input layer, parses and stratifies task intent through a semantic reasoning layer, executes and monitors through a task execution layer, and integrates and optimizes data through a feedback and optimization layer. It supports text, voice, image, and video input, dynamically optimizes subtask paths, and optimizes task execution using complexity assessment and time-consuming prediction models.

Benefits of technology

It realizes the end-to-end conversion of unstructured instructions to structured tasks, dynamically optimizes the task execution path, enhances system flexibility and intelligence, reduces enterprise system integration costs, provides visual task decomposition and execution time comparison, and improves system credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780437A_ABST
    Figure CN120780437A_ABST
Patent Text Reader

Abstract

The invention discloses a process automation execution method based on a large language model, and belongs to the technical field of artificial intelligence and process automation. User intention is analyzed through multi-modal input, and a structured task definition is constructed; the semantic reasoning layer is used for performing task layering, complexity evaluation and sorting optimization; the task execution layer completes subtask scheduling and execution; and the feedback and optimization layer performs performance evaluation and model updating based on execution data to realize closed loop and continuous optimization of the process, so that the technical problems of dynamically analyzing unstructured instructions, automatically optimizing a complex task dependency relationship and adapting to business changes in real time by a process automation tool are solved; according to the method, end-to-end conversion from an unstructured instruction to a structured task is realized, a subtask execution path is dynamically optimized, cross-platform tool calling is supported, the existing system integration cost of an enterprise is reduced, visual display task decomposition logic and prediction and execution time consumption comparison are provided, and the system credibility is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence and process automation, and particularly relates to a process automation execution method based on a large language model. BACKGROUND

[0002] At present, the rapid development of large language models (LLM) provides a new solution for process automation. Large language models have strong semantic understanding and reasoning capabilities, can directly analyze user natural language input, dynamically decompose tasks and generate execution strategies. At the same time, combined with multi-modal data processing technology, the flexibility and intelligent level of the process automation system can be greatly improved.

[0003] With the acceleration of enterprise digital transformation, the complexity and diversity of business processes have increased significantly. Traditional process automation tools (such as ERP systems, workflow management software) mainly rely on predefined rules and fixed templates, and when facing dynamic, complex and unstructured tasks, the following core problems are exposed: Natural language processing capability is insufficient: traditional tools cannot directly understand user natural language instructions, usually need manual configuration of rules or templates, increasing operation complexity and maintenance cost.

[0004] Task decomposition and execution capability is limited: for complex business processes, traditional tools rely on manual task decomposition, and the system cannot automatically identify the logical dependency relationship of tasks, with low execution efficiency.

[0005] Data island problem is serious: the coordination efficiency between system modules is low, cross-department data integration is difficult, and real-time support for complex scenarios cannot be provided.

[0006] Lack of flexibility and intelligence: traditional tools respond slowly to changes in business requirements, and cannot support real-time task adjustment and dynamic reasoning. SUMMARY

[0007] The purpose of the application is to provide a process automation execution method based on a large language model, which solves the technical problems of process automation tools dynamically analyzing unstructured instructions, automatically optimizing complex task dependency relationships, and real-time adapting to business changes.

[0008] To achieve the above purpose, the application adopts the following technical solutions: A process automation execution method based on a large language model, comprising the following steps: Step 1: Obtain multi-modal data input by the user at the multi-modal input layer, analyze each type of modal data, and generate text data; preprocess and perform semantic analysis on the text data to extract the preliminary intent and corresponding parameter information of the task in the text data; structure the preliminary intent into a standardized task definition and send it to the semantic reasoning layer; Step 2: The semantic reasoning layer receives the task definition, performs intent analysis, template matching, and task layering on the task definition, sorts the subtasks, constructs a subtask complexity evaluation model and a time consumption prediction model, evaluates the complexity and predicts the time consumption of all subtasks, optimizes the sorting of subtasks based on the evaluation and prediction results, obtains a subtask plan, describes the dependency relationship in the subtask plan, and sends the subtask plan to the task execution layer; Step 3: The task execution layer receives the subtask plan, decomposes, tool adapts, schedules, and executes the subtasks in the plan; establishes a real-time monitoring mechanism to record the status of the subtasks and perform exception handling; after the execution of the subtasks is completed, the task execution layer returns the subtask execution results, status records, and exception handling to the feedback and optimization layer; Step 4: The feedback and optimization layer integrates the data sent by the task execution layer, uses the integrated data for user interaction, performance analysis, and exception learning, and embeds the complexity evaluation and time consumption prediction of the subtasks to optimize subsequent task planning.

[0009] Preferably, when performing step 1, the following steps are included: Step 1-1: Deploy an input type support module, a data preprocessing module, a modal data analysis module, and a structured data generation module at the multi-modal input layer; The input type support module obtains multi-modal data input by the user, including text, speech, images, and videos; For text, directly extract the sentences related to the task in the text; For speech, input the speech into a speech recognition model to convert the speech into text; For images or videos, use OCR technology to extract readable text content from the images or videos and convert them into text; The input type support module outputs text data in a unified format; Step 1-2: The data preprocessing module calls the text data output by the input type support module, preprocesses the text data, and obtains pure text; Step 1-3: The modal data analysis module calls the pure text, performs semantic analysis and context fusion processing on the pure text, and generates preliminary structured information containing the preliminary intent and key parameters of the task; Step 1-4: The structured data generation module calls the preliminary structured information, formats the preliminary structured information, obtains standardized structured task data, i.e., task definition, and sends the task definition to the semantic reasoning layer.

[0010] Preferably, when step 2 is performed, the semantic analysis module, the task template matching module, the task layering module, the task prediction analysis module, and the context reasoning module are deployed in the semantic reasoning layer, and the specific steps of step 2 are as follows: Step 2-1: The semantic analysis module obtains the task definition, models the semantics in the task definition, and identifies the task intent and the corresponding parameter field; Step 2-2: The task template matching module obtains the task intent, performs similarity matching and semantic expansion of the semantics, obtains the matched task template, and generates the template instance of the task; Step 2-3: The task layering module obtains the template instance, recursively decomposes the complex task, generates a set of subtasks and a dependency graph, and specifically includes: using the template definition to recursively disassemble the complex subtask, extracting the dependency relationship between the subtasks; using the topological sorting method to construct the execution order graph of the subtasks, generating an initial dependency relationship structure; At the same time, the task prediction analysis module performs complexity evaluation and time consumption prediction on the subtasks, specifically: first, establish a task complexity evaluation model, and perform complexity scoring on each subtask, the scoring indicators include task depth, call chain length and data size; then, use the historical log and time consumption prediction model to predict the execution time of each subtask; then, embed the complexity evaluation result and the time consumption prediction result into the subtask node as attributes, and expand the parameter dimension of the task dependency graph; finally, in the topological sorting stage, the execution path of the subtask is optimized according to the dependency relationship, the complexity score result and the time consumption prediction result, and the subtask plan is generated; Step 2-4: The context reasoning module calls the subtask plan and the dialogue context, checks the consistency of the subtask and the context, updates the parameter field and the execution order in the subtask plan according to the checking result, and sends the subtask plan to the task execution layer.

[0011] Preferably, when step 3 is performed, the task decomposition module, the tool adaptation module, the task scheduling module, and the exception handling module are deployed in the task execution layer, and the specific steps of step 3 are as follows: Step 3-1: The task decomposition module obtains the subtask plan, performs parameter inheritance and verification on the task execution structure, and generates an executable subtask queue; Step 3-2: The tool adaptation module obtains the subtask and its parameter field, encapsulates the interface of the tool API, registers the standardized tool calling request, and generates tool calling information when the subtask calls the tool; Step 3-3: The task scheduling module obtains the sub-tasks and their tool call information, generates a scheduling plan according to the complexity evaluation, time consumption prediction and dependency graph, specifically, on the basis of the DAG structure, the complexity evaluation and time consumption prediction are fused, a heuristic sorting strategy is used to optimize the scheduling path, sub-task parallelization and priority control are performed, and the sub-task parallelization and priority adjustment are performed according to the optimization result; Step 3-4: The abnormality processing module tracks and logs the execution state of the sub-tasks in real time, and automatically retries, alarms or fails over the tasks that fail to execute.

[0012] Preferably, when step 4 is executed, the feedback and optimization layer deployment result integration module, the user interaction module, the performance optimization module and the abnormality recording module are deployed, and the specific steps of step 4 are as follows: Step 4-1: The result integration module merges and formats the execution results of the sub-tasks to generate the final task result output; Step 4-2: The user interaction module calls the task result output and the execution information of the sub-tasks, and performs visual display, and the execution information includes the complexity score, the predicted execution time and the actual time consumption of each sub-task; Step 4-3: The performance optimization module compares the log records of the sub-task execution, and optimizes the complexity evaluation model and the time consumption prediction model of the sub-tasks; Step 4-4: The abnormality recording module clusters the abnormal log data in the log records, trains a self-learning model, and predicts the sub-task execution failure events.

[0013] The process automation execution method based on the large language model solves the technical problems that the process automation tool performs dynamic analysis on unstructured instructions, automatically optimizes the dependency relationship of complex tasks, and adapts to business changes in real time. The application supports multi-modal input of text, voice, image and video, realizes end-to-end conversion of unstructured instructions to structured tasks, dynamically optimizes the sub-task execution path through the complexity evaluation model (task depth, call chain), breaks through the limitations of traditional fixed templates, forms a continuous optimization closed loop of planning-execution-improvement based on the exception learning and model iteration of execution feedback, supports cross-platform tool calling through API encapsulation design of the tool adaptation module, reduces the integration cost of existing enterprise systems, provides visual display of task decomposition logic, prediction and execution time consumption comparison, and enhances the credibility of the system. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is the main flowchart of the application; Figure 2 is the flowchart of step 1 of the application; Figure 3 is the flowchart of step 2 of the application; Figure 4 is a flowchart of step 3 of the present application; Figure 5 is a flowchart of step 4 of the present application. DETAILED DESCRIPTION

[0015] A large language model-based flow automation execution method shown in Figures 1-5 comprises the following steps: Step 1: Obtain the multi-modal data input by the user in the multi-modal input layer, parse each type of modal data to generate text data; pre-process and perform semantic analysis on the text data to extract the preliminary intent of the task and the corresponding parameter information in the text data; structure the preliminary intent into a standardized task definition and send it to the semantic reasoning layer; When performing step 1, the following steps are specifically included: Step 1-1: Deploy an input type support module, a data preprocessing module, a modal data parsing module, and a structured data generation module in the multi-modal input layer; The input type support module obtains multi-modal data input by the user, including text, voice, images, and videos.

[0016] For text, the sentences related to the task in the text are directly extracted.

[0017] For voice, the voice input is converted into text by a voice recognition model. In this embodiment, the voice recognition model can use a model such as Whisper to convert voice into text.

[0018] For images or videos, OCR technology is used to extract readable text content in the images or videos and convert them into text.

[0019] The input type support module outputs text data in a unified format.

[0020] Step 1-2: The data preprocessing module calls the text data output by the input type support module, pre-processes the text data to obtain pure text; In this embodiment, the pre-processing of the text data includes removing redundant characters, correcting spelling errors, and unifying time formats.

[0021] The correction of spelling errors can use a BERT error correction model.

[0022] Step 1-3: The modal data parsing module calls the pure text, performs semantic analysis and context fusion processing on the pure text, and generates preliminary structured information containing the preliminary intent of the task and key parameters; In this embodiment, the intent recognition model can use BERT, RoBERTa, or a lightweight classification model based on LLaMA.

[0023] The extraction of key parameters can be achieved by utilizing the multi-label extraction capability of the BiLSTM-CRF model.

[0024] Steps 1-4: The structured data generation module retrieves the preliminary structured information, formats the preliminary structured information, obtains standardized structured task data, namely the task definition, and sends the task definition to the semantic reasoning layer.

[0025] The structured task data is as follows: Task type: Generate report, Parameters: Time range: Last month; Report type: Sales; Output format: PDF.

[0026] Step 2: The semantic reasoning layer receives the task definition, performs task intent analysis, template matching, and task stratification on the task definition, and sorts the subtasks. It then builds a subtask complexity assessment model and a time consumption prediction model, performs complexity assessment and time consumption prediction on all subtasks, optimizes the subtask sorting based on the assessment and prediction results, and obtains a subtask plan. The subtask plan describes the dependencies and sends the subtask plan to the task execution layer. When executing step 2, the semantic parsing module, task template matching module, task stratification module, task prediction analysis module, and contextual reasoning module are deployed in the semantic reasoning layer. The specific steps of step 2 are as follows: Step 2-1: The semantic parsing module obtains the task definition, models the semantics in the task definition, and identifies the task intent and its corresponding parameter fields; In this example, a pre-trained language model (such as BERT) is combined with fine-tuning technology to identify task intent. The intent classification formula is as follows: ; in, is the hidden layer output of the Transformer model, is the classifier weight matrix.

[0027] In this embodiment, the classifier weight matrix The initialization can adopt the common initialization methods in the field of deep learning, such as Xavier initialization or normal distribution random initialization.

[0028] The dimensionality adaptation is usually determined by the number of model task categories, that is, the hidden layer output is mapped to the task category space through linear transformation, which is a common practice for classification models.

[0029] The parameter field is extracted by a sequence labeling model (such as BERT+CRF), for example, using a BIO labeling mechanism to extract key task parameters, such as generating a sales report and sending it to Zhang Manager can be labeled as [B-Action: generate], [B-Object: sales report], [B-Action: send], [B-Recipient: Zhang Manager].

[0030] In this embodiment, for the BIO labeling mechanism of complex task parameters, a hierarchical BIO labeling method can be used: the main parameters of the task are labeled at the top layer, and the sub-parameters are labeled using a two-level BIO labeling rule within the corresponding span, thereby supporting the labeling of nested and multi-dimensional parameters.

[0031] Step 2-2: The task template matching module obtains the task intent, performs semantic similarity matching and semantic expansion, obtains the matched task template, and generates a template instance of the task; In this embodiment, a matching model is used for semantic similarity matching, specifically using a semantic embedding model Sentence-BERT to calculate the similarity between the user input and the task template: ; wherein, represents the text input by the user, and the pre-defined task template; represents the vector representation of the user input, and the vector representation of the task template, represents the Sentence-BERT model.

[0032] In this embodiment, a method of introducing a domain knowledge graph is used to achieve semantic expansion, for example, "analyze sales trends" can be expanded to match "generate a sales analysis report".

[0033] The task template is as follows: Task type: report generation; Sub-task template: {Sub-task: data pulling, parameters: {...}}; {Sub-task: data cleaning, parameters: {...}}; {Sub-task: chart generation, parameters: {...}}.

[0034] Step 2-3: The task hierarchical module obtains the template instance, recursively decomposes the complex task, generates a set of sub-tasks and a dependency graph, specifically including: using a template definition to recursively disassemble complex sub-tasks, and extracting the dependency relationships between sub-tasks; using a topological sorting method to construct an execution order diagram of the sub-tasks, and generating an initial dependency relationship structure; In this embodiment, the topological sorting method is the Kahn algorithm or the DFS algorithm. Task dependencies are obtained by parsing structured fields in the task template, such as predecessor task IDs and input / output binding relationships. These fields are then mapped into directed edges. Combining the Kahn algorithm or the DFS algorithm generates a dependency graph (DAG).

[0035] At the same time, the task prediction and analysis module performs complexity evaluation and time consumption prediction on subtasks. Specifically, it first establishes a task complexity evaluation model and assigns a complexity score to each subtask. The scoring indicators include task depth, call chain length, and data size. Then, it uses historical logs and the time consumption prediction model to predict the execution time of each subtask. The complexity evaluation results and time consumption prediction results are then embedded as attributes into the subtask nodes, expanding the parameter dimension of the task dependency graph. Finally, in the topological sorting stage, the execution path of the subtask is optimized based on the dependency relationship, complexity scoring results, and time consumption prediction results to generate a subtask plan. In this embodiment, the task complexity evaluation model is used to quantitatively score the execution complexity of each subtask, specifically as follows: ; in, represents the subtask complexity score; Indicates the task depth, that is, the hierarchical depth of the subtask in the entire task hierarchy; Indicates the length of the call chain, that is, the number of tools or APIs required to be connected in series when the subtask is completed; Indicates the data scale, i.e. the number of processed data entries, file size, etc. These are preset weight coefficients used to balance the contributions of different dimensions.

[0036] Weight parameters in task complexity scoring model It is a preset adjustable weight coefficient used to balance the influence of various scoring indicators on complexity. The task complexity scoring model is a linear weighted summation form. It uses weighted summation to score complexity. It has a simple structure, low computational complexity, supports direct parameter setting, and does not require training. The value satisfies a+b+c=1.

[0037] In the embodiment of the present application, based on the characteristic preferences of specific task types, system resource bottlenecks and historical task performance, Adjust the value, such as: For system process tasks with complex structures but simple semantic chains, the impact of D is more significant and should be given a higher weight, while L is not very important. In this case, a>b can be set; for example, a=0.5, b=0.3, c=0.2.

[0038] For natural language tasks that focus on semantic processing, such tasks L are more important and should be given higher weights, at which point b>a; such as a=0.2, b=0.6, c=0.2.

[0039] For tasks with large data volume and high execution time requirements (such as large batch data processing and real-time systems), the weight c of the data scale can be increased, such as a=0.3, b=0.3, c=0.4.

[0040] The value of the weight can be determined by historical task data analysis or small-scale parameter tuning before system deployment. The initial setting can use the proportion combination commonly used in engineering, such as a:b:c=0.3:0.3:0.4, and subsequent fine-tuning can also be performed according to actual operation feedback.

[0041] The time-consuming prediction model is used to predict the expected execution time of each sub-task to assist task scheduling optimization and execution time estimation. The specific formula is as follows: ; Where, represents the predicted time-consuming; represents the historical average time-consuming extracted from the log, represents the error correction term estimated based on the LSTM prediction model, considering the difference in context parameters, system load, and task variety, etc.

[0042] is extracted from the task execution log library (such as the log record in task scheduling), and the log format is as follows: Task type: data pulling; Parameter summary: [data volume: 100MB, source: sales database]; Time-consuming: 2.6; Timestamp: 2025-05-01 14:00; Execution environment: [node: server-A, load: 0.5].

[0043] The extraction process of is as follows: Step S1: Extract historical records similar to the current task (such as the same task type, similar parameters), which specifically includes: Task type (such as chart generation); parameter summary (such as data volume, report dimension number, etc.); environmental conditions (optional: same resource conditions).

[0044] Step S2: At least N samples, N is greater than 5, and the time span is limited to the last 30 days.

[0045] Step S3: Calculate the average time-consuming:​ ; wherein, represents the actual time consumption of the i-th historical task record; N represents the number of selected historical samples, which is 10 to 30 in the embodiment; represents the time decay weight of the i-th historical record; represents the decay coefficient, which is 0.001 to 0.01; represents the current time, which is obtained through the system timestamp; represents the timestamp of the i-th historical task record; is the base of natural logarithm.

[0046] is mainly used to make individualized correction to the historical average time consumption to adapt to the specific differences of the current task parameters, environment or context.

[0047] ; ; ; wherein, represents the sensitivity coefficient obtained by the model when adjusting the parameters, represents the context time consumption weight coefficient, represents the hyperparameter, represents the load factor obtained when monitoring the CPU / memory / IO status of the real-time monitoring system, which is [0, 1].

[0048] represents the sensitivity of the model to parameter adjustment, which is a parameter automatically learned through error back propagation in the training process of the deep learning model.

[0049] is a dynamically generated weight according to the task context, which is used to reflect the influence of the context on the time consumption.

[0050] As a hyperparameter in model training, it mainly controls the model complexity or the strength of the regular term in the training process, and is generally selected and verified through common parameter tuning methods (such as cross-validation, grid search, and Bayesian optimization).

[0051] In the embodiment, a context classifier (such as a BERT-based model) is used to label the task context, and then the context time consumption weight coefficient is obtained.

[0052] respectively represent the parameter difference item, the context state item, and the system load adjustment item.

[0053] In this embodiment, the task complexity and time consumption are embedded, and the total execution cost is minimized by using dynamic programming, and the specific calculation formula is as follows: ; wherein, is a task weight, represents a predicted time consumption.

[0054] Step 2-4: The context reasoning module calls the subtask plan and the dialogue context, checks the consistency of the subtask and the context, updates the parameter field and the execution order in the subtask plan according to the checking result, and sends the subtask plan to the task execution layer.

[0055] In this embodiment, when checking the consistency of the subtask and the context, the dialogue state tracking (DST) technology is used for checking, specifically, the context state is maintained to ensure the parameter consistency, the dynamic adjustment of the task in the multi-round interaction is supported, the long-term memory of the context state is realized by using the Transformer or the LSTM, the BERT-based context classifier is introduced to label the task context, such as [time range change], [task type addition], etc.

[0056] In this embodiment, the state storage structure of the DST can store the dialogue parameters in the form of key-value pairs (slot-value), or store the hidden state of the sequence model, and the state update condition is to trigger the update when the user input or the system response is confirmed, and the checking rule is to trigger the re-adjustment when the parameter difference exceeds the preset threshold (such as the numerical deviation exceeds 10%); if there is a context conflict, it is solved according to the time sequence or the preset priority.

[0057] The specific structure of the finally output subtask plan is as follows: Subtask plan-task graph-node: [ID:T1, task name: data pulling, complexity: 3.2, predicted time consumption: 2.4]; [ID:T2, task name: data cleaning, complexity: 4.8, predicted time consumption: 3.1]; Dependency relationship: [{from: T1, to: T2}].

[0058] Step 3: The task execution layer receives the subtask plan, decomposes, tool adapts, schedules and executes the subtasks therein; establishes a real-time monitoring mechanism to record the state of the subtasks and performs exception handling; after the execution of the subtasks is completed, the task execution layer returns the subtask execution result, the state record and the exception handling to the feedback and optimization layer; When executing step 3, the task decomposition module, tool adaptation module, task scheduling module, and exception handling module are deployed in the task execution layer. The specific steps of step 3 are as follows: Step 3-1: The task decomposition module obtains the subtask plan, inherits and verifies the parameters of the task execution structure, and generates an executable subtask queue; Step 3-2: The tool adapter module obtains the subtask and its parameter fields, connects to the tool API for interface encapsulation (such as RESTful API or gRPC), registers standardized tool call requests, and generates tool call information when the subtask calls the tool; Step 3-3: The task scheduling module obtains subtasks and their tool call information, and generates a scheduling plan based on complexity assessment, time consumption prediction, and dependency graph. Specifically, based on the DAG structure, it integrates complexity assessment and time consumption prediction, uses a heuristic sorting strategy to optimize the scheduling path, performs subtask parallelization and priority control, and adjusts the subtask parallelism and priority based on the optimization results. The scheduling strategy combines task complexity scores, time consumption predictions, and priority functions to optimize the scheduling path. The priority function takes into account factors such as task importance, resource costs, and time constraints. The scheduling process supports parallel scheduling and dynamic adjustment based on the DAG structure (such as adjusting the order when some tasks are delayed). At the implementation level, a distributed scheduling framework (such as Apache Airflow or Celery) can be used for task management and scheduling execution.

[0059] Step 3-4: The exception handling module tracks and logs the execution status of subtasks in real time, and automatically retries, issues alarms, or performs failover processing on failed tasks.

[0060] The exception handling module uses an event-driven state machine to track task status, including the following: ["pending", "executing", "completed", "failed"]. It supports an automatic retry mechanism (with set retry intervals and times) and calls a backup tool solution when a tool fails. When the automatic mechanism cannot recover, the system will request user intervention through a notification mechanism (such as message push and email) to ensure that the task process is not interrupted.

[0061] In this embodiment, the state transition rule of the event-driven state machine is that the task can be transferred from the "executing" state to the "success / failure" state. If it fails, it enters the "retry / alarm" state. The state persistence method can store state information through a database (such as MySQL / Redis) or log file, and maintain consistency through a synchronization mechanism. The backup tool selection logic is that when the main tool fails, the system gives priority to the backup tool according to the performance indicators (such as average time consumption) and success rate. The tool switching mechanism is to release the original tool handle and initialize the backup tool when switching to ensure that resources do not conflict.

[0062] The design of the event-driven state machine is the prior art in the field of distributed scheduling, and thus is not described in detail.

[0063] Step 4: The feedback and optimization layer integrates the data sent by the task execution layer, and the integrated data is used for user interaction, performance analysis and anomaly learning, and is embedded in the complexity evaluation and time consumption prediction of the subtasks, thereby optimizing subsequent task planning.

[0064] When performing step 4, the feedback and optimization layer deploys a result integration module, a user interaction module, a performance optimization module and an anomaly recording module, and the specific steps of step 4 are as follows: Step 4-1: The result integration module merges and formats the execution results of each subtask to generate the final task result output; In this embodiment, the result integration module merges the distributed subtask execution results, uses a message queue (such as Kafka) to collect the execution output, and uses a distributed computing framework (such as Apache Spark) for classification processing; a rule engine is introduced in the integration process for data consistency checking (such as timestamp uniform format, field integrity verification), and a structured task output is generated according to user needs, supporting multiple formats such as JSON, PDF, Excel, charts, etc.

[0065] Step 4-2: The user interaction module calls the task result output and the execution information of the subtasks, and performs visual display, and the execution information includes the complexity score, the predicted execution time and the actual time consumption of each subtask; The user interaction module displays the task execution results through a visualization framework (such as D3.js, Plotly), including the complexity score, the predicted time consumption and the actual time consumption of each subtask; the module supports a real-time interaction mechanism based on WebSocket, and the user can adjust the task parameters (such as time range, analysis dimension, etc.), and the system updates the task planning in real time according to the changes, dynamically generates a new dependency relationship graph and an execution path.

[0066] Step 4-3: The performance optimization module compares the log records of the subtask execution, and optimizes the complexity evaluation model and the time consumption prediction model of the subtasks; The performance optimization module can analyze the execution bottlenecks based on the subtask log records, and the monitoring indicators include execution time consumption, tool response time, resource usage (such as CPU, memory); the indicators are stored and analyzed through a time series database (such as InfluxDB) and an OLAP analysis tool (such as ClickHouse). The module optimizes the task decomposition strategy through a genetic algorithm, and uses a reinforcement learning mechanism (such as selecting the optimal tool based on the task context) to continuously improve the task execution efficiency and success rate.

[0067] In this embodiment, the parameters of the genetic algorithm can be set as follows: the population size is generally set to 50-200, the crossover rate and mutation rate are adjusted in the range of 0.1-0.3.

[0068] The parameter settings of the reinforcement learning mechanism are as follows: the learning rate, discount factor, etc. can adopt the commonly used range (such as learning rate 0.001-0.01).

[0069] The state-action-reward samples are constructed using the historical task logs, and then the training set is constructed, wherein the reward value can be defined by the task execution time reduction and success rate improvement.

[0070] Step 4-4: The anomaly record module clusters the abnormal log data in the log record, trains the self-learning model, and predicts the sub-task execution failure event.

[0071] When the anomaly record module collects and analyzes the abnormal logs during execution, the abnormal types include tool call timeout, dependent task not completed, insufficient resources, etc. The module uses the ELK stack (Elasticsearch, Logstash, Kibana) for log collection and visual analysis, and identifies abnormal patterns based on clustering algorithms (such as K-Means). Combined with the prediction model (such as XGBoost), the potential failure event is predicted, and the model inference strategy is fine-tuned through the self-learning mechanism to optimize the task execution planning.

[0072] In this embodiment, by establishing a correlation model (such as linear regression, correlation analysis) between the tool call response time recorded in the task log and the task parameter size, the complexity and time consumption are associated.

[0073] In the analysis of execution bottlenecks, clustering analysis or association rule mining algorithms can be used to locate the bottlenecks. Both clustering analysis and association rule mining algorithms are prior art, so they will not be described in detail.

[0074] In this embodiment, the self-learning model is set to periodic self-learning scheduling, such as once a week or triggered after a certain number of samples are met.

[0075] The process automation execution method based on the large language model solves the technical problems that the process automation tool analyzes unstructured instructions dynamically, optimizes complex task dependency relations automatically, and adapts to business changes in real time, supports multi-modal input of text, voice, image and video, realizes end-to-end conversion of unstructured instructions to structured tasks, dynamically optimizes the sub-task execution path through the complexity evaluation model (task depth, call chain), breaks through the limitation of the traditional fixed template, learns abnormity based on execution feedback and iterates the model, forms a continuous optimization closed loop of planning-execution-improvement, supports cross-platform tool calling through API encapsulation design of the tool adaptation module, reduces the integration cost of the existing system of enterprises, provides visual display of task decomposition logic, comparison of prediction and execution time consumption, and enhances the credibility of the system.

Claims

1. A process automation execution method based on a large language model, characterized by: The steps include: Step 1: Obtain multimodal data input by the user at the multimodal input layer, parse the various modal data, and generate text data; perform preprocessing and semantic analysis on the text data to extract the preliminary intent of the task and corresponding parameter information from the text data; Structuring preliminary intent into standardized task definitions and sending them to the semantic reasoning layer; Step 2: The semantic reasoning layer receives the task definition, performs task intent analysis, template matching, and task stratification on the task definition, and sorts the subtasks. It then builds a subtask complexity assessment model and a time consumption prediction model, performs complexity assessment and time consumption prediction on all subtasks, optimizes the subtask sorting based on the assessment and prediction results, and obtains a subtask plan. The subtask plan describes the dependencies and sends the subtask plan to the task execution layer. The task complexity evaluation model is used to quantify the execution complexity of each subtask, as shown in the following formula: ; in, represents the subtask complexity score; Indicates the task depth, that is, the hierarchical depth of the subtask in the entire task hierarchy; Indicates the length of the call chain, that is, the number of tools or APIs required to be connected in series when the subtask is completed; Indicates the data scale, i.e. the number of processed data entries and file size; They are all preset weight coefficients used to balance the contributions of different dimensions; The time consumption prediction model is used to predict the expected execution time of each subtask to assist in task scheduling optimization and execution time estimation. The specific formula is as follows: ; in, Indicates the prediction time; represents the historical average time taken to extract data from logs. represents the error correction term estimated based on the LSTM prediction model, taking into account context parameter differences, system load, and task variations; It is extracted from the task execution log library; The extraction process is as follows: Step S1: Extract historical records similar to the current task, including: Mission type; parameter summary; environmental conditions; Step S2: at least N samples, N is greater than 5; Step S3: Calculate the average time: ; in, represents the actual time taken by the i-th historical task record; N represents the number of selected historical samples; represents the time decay weight of the i-th historical record; Represents the attenuation coefficient, ranging from 0.001 to 0.01; Indicates the current time, obtained through the system timestamp; Indicates the timestamp of the i-th historical task record; is the base of natural logarithms; Mainly used to average the historical time Make personalized modifications to accommodate specific differences in the current task parameters, environment, or context; ; ; ; in, It represents the sensitivity coefficient obtained when the model adjusts the parameters. represents the context time weight coefficient, represents the hyperparameter, Indicates that the load factor is obtained when the system CPU / memory / IO status is monitored in real time, and the value is [0,1]; Use a context classifier to label the task context and then obtain the context time weight coefficient ; They represent parameter difference item, context state item, and system load adjustment item respectively; Embed the task complexity and time consumption, and use dynamic programming to minimize the total execution cost. The specific calculation formula is as follows: ; in, is the task weight, Indicates the prediction time; Step 3: The task execution layer receives the subtask plan, decomposes the subtasks, adapts tools, schedules, and executes them. It also establishes a real-time monitoring mechanism to record the status of the subtasks and handle exceptions. After the subtasks are completed, the task execution layer returns the subtask execution results, status records, and exception handling to the feedback and optimization layer. Step 4: In the feedback and optimization layer, the data sent by the task execution layer is integrated. The integrated data is used for user interaction, performance analysis, and anomaly learning, and is embedded in the complexity evaluation and time consumption prediction of subtasks to optimize subsequent task planning.

2. The process automation execution method based on a large language model according to claim 1, characterized in that: When executing step 1, the specific steps include: Step 1-1: Deploy the input type support module, data preprocessing module, modal data parsing module, and structured data generation module in the multimodal input layer; The input type supports the module to obtain multimodal data input by the user, including text, voice, image and video; Directly extract the sentences related to the task from the text; For speech, the speech is fed into a speech recognition model to transcribe the speech into text; For images or videos, use OCR technology to extract readable text content in the image or video and convert it into text; The input type supports text data with a unified module output format; Step 1-2: The data preprocessing module retrieves the text data output by the input type support module and preprocesses the text data to obtain pure text; Steps 1-3: The modal data parsing module retrieves the clean text, performs semantic analysis and context fusion processing on the clean text, and generates preliminary structured information containing the preliminary task intent and key parameters; Steps 1-4: The structured data generation module retrieves the preliminary structured information, formats the preliminary structured information, obtains standardized structured task data, namely the task definition, and sends the task definition to the semantic reasoning layer.

3. The process automation execution method based on a large language model according to claim 1, characterized in that: When executing step 2, the semantic parsing module, task template matching module, task stratification module, task prediction analysis module, and contextual reasoning module are deployed in the semantic reasoning layer. The specific steps of step 2 are as follows: Step 2-1: The semantic parsing module obtains the task definition, models the semantics in the task definition, and identifies the task intent and its corresponding parameter fields; Step 2-2: The task template matching module obtains the task intent, performs semantic similarity matching and semantic expansion, obtains the matching task template, and generates the task template instance; Step 2-3: The task layering module obtains the template instance, recursively decomposes the complex task, and generates a subtask set and dependency graph. Specifically, it uses the template definition to recursively decompose the complex subtasks and extract the dependency relationships between the subtasks. Use the topological sorting method to construct the execution sequence graph of subtasks and generate the initial dependency structure; At the same time, the task prediction and analysis module performs complexity evaluation and time consumption prediction on subtasks. Specifically, it first establishes a task complexity evaluation model and assigns a complexity score to each subtask. The scoring indicators include task depth, call chain length, and data size. Then, it uses historical logs and the time consumption prediction model to predict the execution time of each subtask. The complexity evaluation results and time consumption prediction results are then embedded as attributes into the subtask nodes, expanding the parameter dimension of the task dependency graph. Finally, in the topological sorting stage, the execution path of the subtask is optimized based on the dependency relationship, complexity scoring results, and time consumption prediction results to generate a subtask plan. Step 2-4: The contextual reasoning module retrieves the subtask plan and the dialogue context book, checks the consistency between the subtask and the context book, updates the parameter fields and execution order in the subtask plan based on the verification results, and sends the subtask plan to the task execution layer.

4. The process automation execution method based on a large language model according to claim 1, characterized in that: When executing step 3, the task decomposition module, tool adaptation module, task scheduling module, and exception handling module are deployed in the task execution layer. The specific steps of step 3 are as follows: Step 3-1: The task decomposition module obtains the subtask plan, inherits and verifies the parameters of the task execution structure, and generates an executable subtask queue; Step 3-2: The tool adapter module obtains the subtask and its parameter fields, connects to the tool API for interface encapsulation, registers the standardized tool call request, and generates tool call information when the subtask calls the tool; Step 3-3: The task scheduling module obtains subtasks and their tool call information, and generates a scheduling plan based on complexity assessment, time consumption prediction, and dependency graph. Specifically, based on the DAG structure, it integrates complexity assessment and time consumption prediction, uses a heuristic sorting strategy to optimize the scheduling path, performs subtask parallelization and priority control, and adjusts the subtask parallelism and priority based on the optimization results. Step 3-4: The exception handling module tracks and logs the execution status of subtasks in real time, and automatically retries, issues alarms, or performs failover processing for failed tasks.

5. The process automation execution method based on a large language model according to claim 1, characterized in that: When executing step 4, deploy the result integration module, user interaction module, performance optimization module, and exception recording module in the feedback and optimization layer. The specific steps of step 4 are as follows: Step 4-1: The result integration module merges and formats the execution results of each subtask to generate the final task result output; Step 4-2: The user interaction module retrieves the task result output and subtask execution information and displays them visually. The execution information includes the complexity score, estimated execution time, and actual execution time of each subtask; Step 4-3: The performance optimization module compares the subtask execution log records and optimizes the subtask complexity evaluation model and time consumption prediction model; Step 4-4: The exception record module performs cluster modeling on the abnormal log data in the log records, trains the self-learning model, and predicts subtask execution failure events.

Citation Information

Patent Citations

  • Multi-task cooperative processing method and system based on digital employees

    CN119179563A

  • Complex task decomposition and dynamic optimization method and device based on large language model

    CN119883549A

  • Multi-agent interactive efficient data analysis system

    CN119988421A

  • Task understanding and executing system based on multi-mode intelligent agent

    CN120337977A

Cited By

  • Task complexity driven graph semantic multi-agent collaborative decision-making method and system

    CN120950220A

  • Large language model low-delay reasoning method based on dynamic reasoning graph optimization

    CN121072787A

  • A large language model low-latency inference method based on dynamic inference graph optimization

    CN121072787B

  • Bio-computing multi-tool unified calling method and system based on large language model

    CN121233270A

  • Unified calling method and system for multiple tools of biological computing based on large language model

    CN121233270B