A language-driven intelligent workflow planning method and system
Patent Information
- Application Number
- CN202610851165.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]现有专利CN120542818A公开了一种基于工作流引擎算法的智能指挥及其数据处理方法,其通过任务管理模块接收任务信息并生成任务,再由工作流引擎模块采用最早截止时间优先算法进行任务调度,然而,任务被创建之后的调度、资源分配和路径优化,其默认已经获得了用于计算和调度的结构化任务信息,对于用户提出的复合型需求,并未考虑从用户表达中识别任务意图,例如,当用户输入“收集数据、清洗数据、生成图表、撰写摘要并发送邮件”时,该类需求并不是单一任务,而是包含数据获取、数据预处理、图表生成、文本撰写和消息发送等多个隐含步骤,并且各步骤之间还存在先后依赖、数据传递和执行状态衔接关系,由于缺少将用户复合需求转化为结构化工作流的规划机制,也缺少对多个执行节点之间依赖关系的构建能力,降低了工作流规划的效率和智能化程度,难以满足语言驱动场景下对自动规划、编排和执行的需求
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: By segmenting natural language instructions into sentences and forming task items based on action words, object words, and result words, it can convert complex user needs expressed in natural language into task units with processing objects, processing methods, and output content, avoiding the problem of only processing structured task information. Furthermore, by merging and judging task items based on processing objects, processing methods, and output content, it can distinguish between related processing in the same execution stage and continuous processing in different execution stages. By establishing data transfer relationships based on the correspondence between output content and processing objects, and combining sequential words, parallel words, and conditional words to determine execution dependencies, it can enable data succession, execution order, parallel hierarchy, and conditional triggering relationships among candidate execution items. By identifying missing prerequisite items and breakpoint items, and backtracking and completing missing prerequisite items and re-accepting or confirming breakpoint items, it can reduce workflow breaks caused by missing input sources or unaccepted intermediate outputs, improving the efficiency and intelligence of workflow planning, and meeting the needs for automatic planning, orchestration, and execution in language-driven scenarios.
Smart Images

Figure CN122672775A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of workflow planning technology, and more specifically, to a language-driven intelligent workflow planning method and system. Background Technology
[0002] With the development of technologies such as natural language interaction and large-scale intelligent agents, intelligent workflow planning systems are gradually being applied to scenarios such as data processing, content generation, business approval, software operation, intelligent office and cross-system task collaboration.
[0003] Existing patent CN120542818A discloses an intelligent command and data processing method based on a workflow engine algorithm. This method receives task information and generates tasks through a task management module, and then the workflow engine module schedules tasks using an earliest deadline first algorithm. However, the scheduling, resource allocation, and path optimization of tasks after creation already utilize structured task information for computation and scheduling. For complex user requests, it does not consider identifying task intent from user expressions. For example, when a user inputs "collect data, clean data, generate charts, write summaries, and send emails," this type of request is not a single task but includes multiple implicit steps such as data acquisition, data preprocessing, chart generation, text writing, and message sending. Furthermore, there are sequential dependencies, data transfer, and execution status connections between these steps. The lack of a planning mechanism to transform complex user requests into structured workflows, and the lack of ability to construct dependencies between multiple execution nodes, reduces the efficiency and intelligence of workflow planning, making it difficult to meet the needs of automatic planning, orchestration, and execution in language-driven scenarios.
[0004] Therefore, it is necessary to design a language-driven intelligent workflow planning method and system to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a language-driven intelligent workflow planning method and system, which aims to solve the problems of existing technologies lacking a planning mechanism to transform complex user needs into structured workflows, as well as lacking the ability to construct dependencies between multiple execution nodes, thus reducing the efficiency and intelligence of workflow planning and making it difficult to meet the needs of automatic planning, orchestration and execution in language-driven scenarios.
[0006] In one aspect, this invention proposes a language-driven intelligent workflow planning method, comprising: The system receives input natural language instructions, segments the natural language instructions into sentences, determines several language segments, determines task items based on action words, object words, and result words in each language segment, and determines the processing object, processing method, and output content of each task item. Tasks with the same processing object are merged to determine several candidate execution items. Based on the correspondence between the output content of each candidate execution item and the processing object of other candidate execution items, the data transmission relationship of each candidate execution item is determined. The order of the data transmission relationship is corrected based on the order words, parallel words and condition words of the natural language instructions to determine the execution dependency relationship. Based on the execution dependency relationship, each candidate execution item is judged. Candidate execution items that do not have an input source and do not correspond to the starting content of the natural language instruction are judged as missing preconditions. Candidate execution items whose output content is not inherited by the candidate execution item and do not correspond to the result content of the natural language instruction are judged as breakpoint items. Based on the processing object corresponding to the missing prerequisite, the natural language instruction is backtracked, the data transmission relationship is re-established based on the output content corresponding to the breakpoint, and a structured workflow corresponding to the natural language instruction is generated based on the candidate execution items after backtracking and re-establishing the data transmission relationship.
[0007] Furthermore, when segmenting the natural language instructions into sentences, the process includes: The segmentation position is determined based on punctuation marks and sequential words. The content of sentences located at adjacent segmentation positions is identified as language segments. Action words for processing behaviors are extracted from each language segment, and nouns that have semantic association with the action words are identified as object words. The words representing the processing result, output form, or target state are used as result words. If the previous language segment contains an object word, the next language segment omits the object word, and the action word in the next language segment acts on the object word in the previous language segment, then the object word in the previous language segment is filled back into the next language segment.
[0008] Furthermore, when defining the tasks and their respective processing objects, processing methods, and outputs, the following should be included: When a language segment includes an action word, an object word, and a result word, the corresponding processing method, processing object, and output content are taken as task items. When a language segment contains multiple action words, the object word and result word corresponding to each action word are extracted based on the order in which each action word appears in the language segment. When an action word corresponds to multiple object words, corresponding task items are generated based on each object word.
[0009] Furthermore, when defining the tasks and their respective processing objects, processing methods, and outputs, the following also applies: The action words in each language segment are used as the processing method, the object words that have a dominant relationship with the action words are used as the processing objects, and the result words that have a result correspondence with the action words or object words are used as the output content.
[0010] Furthermore, in determining the data transfer relationships for each candidate execution item, the following is included: The output content of each candidate execution item is matched with the processing objects of other candidate execution items. Based on the term consistency relationship and hierarchical inclusion relationship between the output content and the processing object, the succession matching result of each candidate execution item is determined. Based on the succession matching result, the data transmission relationship of each candidate execution item is determined.
[0011] Furthermore, when determining execution dependencies, this includes: When there are precedence terms between the candidate execution items at both ends of the data transmission relationship, the corresponding candidate execution items are determined as serial dependencies based on the execution order indicated by the precedence terms. When there are parallel terms in the candidate execution items at both ends of the data transmission relationship, the corresponding candidate execution items are determined as sibling dependencies; When there are condition words in the candidate execution items at both ends of the data transmission relationship, the judgment content corresponding to the condition word is written into the dependency condition of the corresponding candidate execution item.
[0012] Furthermore, when determining the absence of prerequisites, the following should be considered: Determine the input source and language location of each candidate execution item, and compare the candidate execution items without an input source with the starting content of the natural language instruction; When the processing object of the candidate execution item does not appear in the beginning content of the natural language instruction, the candidate execution item is determined to be a missing prerequisite item, and the language position of the missing prerequisite item is recorded.
[0013] Furthermore, when the processing object corresponding to the missing precondition is backtracked in the natural language instruction, it includes: Using the language location of the missing preceding item as the starting point for backtracking, search for language fragments that contain the same processing object, superior processing object, or source description; Based on the search results, the pre-processing object and pre-processing method are determined, and the pre-processing candidate execution items are determined based on the pre-processing object and pre-processing method. The output content of the pre-processing candidate execution items is used as the input source of the missing pre-processing item.
[0014] Furthermore, when generating a structured workflow corresponding to the natural language instructions based on candidate execution items after backtracking and re-establishing data transfer relationships, the process includes: Based on the output content of the breakpoint event, determine whether there is a candidate execution event that receives the output content; When there is a candidate execution item that can receive the output content, the data transfer relationship is re-established between the breakpoint item and the candidate execution item; If there is no candidate execution item to receive the output content, the output content of the breakpoint item is compared with the natural language instruction. When the output content of the breakpoint item is consistent with the natural language instruction, the breakpoint item is determined as a termination node; otherwise, an output confirmation node is generated based on the output content of the breakpoint item, and the output content of the breakpoint item is used as the input source of the output confirmation node. Based on the serial dependencies, sibling dependencies, and dependency conditions, each candidate execution item is connected to determine the structured workflow.
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: By segmenting natural language instructions into sentences and forming task items based on action words, object words, and result words, it can convert complex user needs expressed in natural language into task units with processing objects, processing methods, and output content, avoiding the problem of only processing structured task information. Furthermore, by merging and judging task items based on processing objects, processing methods, and output content, it can distinguish between related processing in the same execution stage and continuous processing in different execution stages. By establishing data transfer relationships based on the correspondence between output content and processing objects, and combining sequential words, parallel words, and conditional words to determine execution dependencies, it can enable data succession, execution order, parallel hierarchy, and conditional triggering relationships among candidate execution items. By identifying missing prerequisite items and breakpoint items, and backtracking and completing missing prerequisite items and re-accepting or confirming breakpoint items, it can reduce workflow breaks caused by missing input sources or unaccepted intermediate outputs, improving the efficiency and intelligence of workflow planning, and meeting the needs for automatic planning, orchestration, and execution in language-driven scenarios.
[0016] On the other hand, this application also provides a language-driven intelligent workflow planning system for applying the above-mentioned language-driven intelligent workflow planning method, including: The planning receiving unit is configured to receive input natural language instructions, segment the natural language instructions into sentences, determine several language segments, determine task items based on action words, object words and result words in each language segment, and determine the processing object, processing method and output content of each task item. The first planning unit is configured to merge tasks with the same processing object, determine several candidate execution items, determine the data transmission relationship of each candidate execution item based on the correspondence between the output content of each candidate execution item and the processing object of other candidate execution items, and perform order correction on the data transmission relationship based on the order words, parallel words and condition words of the natural language instructions to determine the execution dependency relationship. The second planning unit is configured to determine each candidate execution item based on the execution dependency relationship, and to determine the candidate execution item that has no input source and does not correspond to the start content of the natural language instruction as a missing prerequisite item, and to determine the candidate execution item whose output content is not inherited by the candidate execution item and does not correspond to the result content of the natural language instruction as a breakpoint item. The planning and processing unit is configured to backtrack in the natural language instruction based on the processing object corresponding to the missing prerequisite, re-establish the data transmission relationship based on the output content corresponding to the breakpoint, and generate a structured workflow corresponding to the natural language instruction based on the candidate execution items after backtracking and re-establishing the data transmission relationship.
[0017] It is understandable that the above-mentioned language-driven intelligent workflow planning method and system have the same beneficial effects, and will not be elaborated further here. Attached Figure Description
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating a language-driven intelligent workflow planning method provided in an embodiment of the present invention; Figure 2 A logical flowchart for determining execution dependencies provided in an embodiment of the present invention; Figure 3 This is a functional block diagram of a language-driven intelligent workflow planning system provided in an embodiment of the present invention. Detailed Implementation
[0019] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] See Figure 1-2 As shown in some embodiments of this application, a language-driven intelligent workflow planning method includes: S100: Receives input natural language instructions, segments the natural language instructions into sentences, determines several language segments, determines the task items based on the action words, object words and result words in each language segment, and determines the processing object, processing method and output content of each task item. S200: Merge tasks with the same processing object, determine several candidate execution items, determine the data transmission relationship of each candidate execution item based on the correspondence between the output content of each candidate execution item and the processing object of other candidate execution items, and perform order correction on the data transmission relationship based on the order words, parallel words and condition words of natural language instructions to determine the execution dependency relationship; S300: Based on execution dependencies, each candidate execution item is judged. Candidate execution items that do not have an input source and do not correspond to the starting content of the natural language instruction are judged as missing predecessor items. Candidate execution items whose output content is not inherited by the candidate execution item and do not correspond to the result content of the natural language instruction are judged as breakpoint items. S400: Backtracking is performed in natural language instructions based on the processing objects corresponding to the missing preceding items, and the data transmission relationship is re-established based on the output content corresponding to the breakpoint items. Based on the backtracking and the re-established data transmission relationship, the candidate execution items are generated to form a structured workflow corresponding to the natural language instructions.
[0021] Specifically, in S100, a natural language instruction is a complex processing request expressed by the user in natural language. It can include a single processing action or multiple consecutive, parallel, or conditionally triggered processing actions. For example, if a user inputs "First read the uploaded sales data, clean outliers, and generate bar charts and line charts. If the data is abnormal, generate a review report and send the charts and reports to the person in charge," this natural language instruction simultaneously contains multiple processing intentions such as reading, cleaning, generating, judging, and sending.
[0022] Upon receiving a natural language instruction, the instruction is first segmented. Sentence segmentation breaks down continuous expressions into language fragments that can individually carry processing intent. For example, a natural language instruction might be segmented into language fragments such as "clean outliers" or "generate a review report." This allows subsequent processing to identify the action, object, and result within each fragment. During segmentation, the segmentation position is determined based on punctuation marks and sequence words. Punctuation marks include natural pauses such as commas, semicolons, and periods, while sequence words include words indicating execution order such as "first," "then," "afterwards," and "finally."
[0023] After determining the segmentation points, the content of statements between adjacent segmentation points is identified as language fragments. For example, the natural language instruction "First read the uploaded sales data, clean outliers, and generate bar charts and line charts" can be segmented into language fragments such as "read the uploaded sales data," "clean outliers," and "generate bar charts and line charts." Each language fragment is recorded in the order of its appearance in the natural language instruction to facilitate subsequent determination of execution order, input source, and backtracking starting point.
[0024] Extract action words from each language segment to describe the processing behavior. Action words are words that indicate the processing behavior requested by the user, including reading, importing, collecting, cleaning, filtering, transforming, analyzing, generating, composing, sending, archiving, and judging. Action words are not limited to single verbs in traditional grammar; they can also be phrases with task execution meanings, such as "extract fields," "generate reports," and "send emails."
[0025] After extracting the action words, the noun phrases that are semantically related to the action words are identified as object words. Object words represent the objects that the action words act upon, process, read, generate, or send; these can be data, files, tables, charts, reports, emails, responsible persons, fields, or business objects. For example, in "Read uploaded sales data," "sales data" is the object word; in "Send email," "email" is the object word; and in "Generate bar charts and line charts," "bar chart" and "line chart" are the object words.
[0026] Terms describing the processing result, output form, or target state are used as result terms. Result terms characterize the output content, target state, or result form formed after the action is completed. For example, "cleaned sales data," "review report," and "sent email" can all be used as result terms. If a result term does not appear directly in the language fragment, it can be deduced based on the correspondence between action terms and object terms. For example, the result term for "reading sales data" can be determined as "sales data set," and the result term for "cleaning outliers" can be determined as "cleaned sales data."
[0027] During sentence segmentation, if a preceding language segment contains an object word, and the following language segment omits the object word, and the action word in the following language segment acts on the object word in the preceding language segment, then the object word from the preceding language segment is filled back into the following language segment. For example, in the sentence "Import sales data, clean and generate charts," the object word for "clean" is omitted, but its action can act on "sales data" in the preceding language segment. Therefore, "sales data" is filled back into the language segment corresponding to "clean," making that segment "clean sales data."
[0028] Object word backfilling does not process all adjacent language segments. Backfilling only occurs when the object word of the preceding language segment can be reasonably governed by the action word of the following language segment. If the preceding language segment is "send email" and the following language segment is "generate chart", the object to be generated for "chart" cannot be directly backfilled from "email". In this case, object word backfilling is not performed to avoid subsequent tasks being unable to determine the input source due to object omission, and also to avoid introducing irrelevant objects into subsequent processing.
[0029] A task is the smallest programmable processing unit extracted from a natural language segment. It consists of a processing object, a processing method, and an output content. The processing object comes from object words, the processing method comes from action words, and the output content comes from result words or result content derived from action words and object words.
[0030] When a language fragment includes an action word, an object word, and a result word, the processing method corresponding to the action word, the processing object corresponding to the object word, and the output content corresponding to the result word are considered as a single task. For example, in the phrase "Generate bar charts and line charts based on sales data," "generate" is the processing method, "sales data" or "cleaned sales data" output from a preceding candidate task is the processing object, and "bar chart" and "line chart" are the output content. When the language fragment is only expressed as "generate bar charts" without directly specifying the underlying data used to generate the bar charts, then "bar charts" can be identified as the output content. The processing object or input source of this generation task can then be determined based on the output content of the preceding candidate tasks, the source description in the natural language instruction, or the contextual relationship.
[0031] When a language segment contains multiple action words, the object word and result word corresponding to each action word are extracted based on the order in which they appear in the language segment, and then these are used to form separate task items. For example, in the task "Read sales data and generate statistical charts", the processing object "sales data" and the output content "sales data set" corresponding to "read" are extracted first, and then the processing object "sales data set" or "statistical chart generation data" and the output content "statistical chart" corresponding to "generate" are extracted, thus forming two task items: one for reading and one for generating.
[0032] When an action term corresponds to multiple object terms, corresponding task items are generated based on each object term. For example, in "cleaning sales data and inventory data," "cleaning" is the same processing method, while "sales data" and "inventory data" are two processing objects. Therefore, two task items, "cleaning sales data" and "cleaning inventory data," are generated respectively. This allows the input sources and output content of the two processing objects to be recorded separately in subsequent workflows, and a common relationship can be established during the aggregation.
[0033] When defining tasks, action words in each language segment are used as processing methods, object words with a dominance relationship to the action words are used as processing objects, and result words with a result correspondence relationship to either the action word or the object word are used as output content. A dominance relationship indicates that the action word can directly act on the object word; for example, in "cleaning sales data," "cleaning" dominates "sales data," and in "sending email," "sending" dominates "email." A result correspondence indicates that the result word describes the output form after the action is completed; for example, "generating a chart" corresponds to the output "chart." Each task includes language location, processing object, processing method, and output content, serving as the basis for merging candidate execution items, determining data transfer relationships, and correcting execution dependencies.
[0034] In S200, tasks with the same processing object are merged to determine several candidate execution items. Candidate execution items are merged executable workflow nodes, used to prevent the same processing object from being repeatedly split into multiple isolated nodes in the same execution phase. During merging, the processing object is used as the main thread, while also considering the processing method and output content to determine whether multiple tasks belong to the same execution phase.
[0035] The same execution phase refers to multiple tasks that, although broken down into multiple actions in language, share the same or substantially the same processing object, have a continuous and consistent processing purpose, and do not produce a phased output that can be independently inherited by subsequent candidate execution tasks. They can be considered a continuous processing process within an execution node. For example, when performing "reading," "format validation," and "field normalization" on the same sales data, if their common purpose is to obtain sales data usable for subsequent processing, and the intermediate results do not serve as input sources for other candidate execution tasks, then these multiple tasks can be identified as the same execution phase and merged into a single candidate execution task. Conversely, if the output of a previous task changes the data state upon which subsequent processing depends, or if the output can be independently inherited by subsequent tasks, then the multiple tasks have formed different execution phases and should not be merged solely because they share the same processing object.
[0036] When multiple tasks share the same processing object and their processing methods belong to the same execution stage, or when a later task is only used to supplement, standardize, or improve the processing results of a previous task and does not form an independent stage output, these multiple tasks can be merged into a single candidate execution task. For example, "reading sales data and validating the sales data format" can be merged into a single candidate execution task focusing on data reading and format validation related to "sales data".
[0037] When multiple tasks involve the same object of processing but their processing methods belong to different execution stages, or when the output of a previous task can serve as the input source for a subsequent task, they should not be merged but rather identified as separate candidate tasks. For example, although "collecting sales data, cleaning sales data, and analyzing sales data" all involve "sales data," their outputs are, in order, "sales data set," "cleaned sales data," and "data analysis results." These three tasks form consecutive processing stages and should be identified as separate candidate tasks.
[0038] After identifying candidate execution items, the data transfer relationships between them are determined based on the correspondence between the output content of each candidate execution item and the processing objects of other candidate execution items. A data transfer relationship indicates that the output content of one candidate execution item can serve as the input source required for the execution of another candidate execution item. This relationship is not determined by language order, but rather by whether the output content and the processing object can be seamlessly connected.
[0039] When determining data transfer relationships, the output content of each candidate execution item is matched with the processing objects of other candidate execution items. Based on the terminology consistency and hierarchical inclusion relationships between the output content and the processing objects, the matching result of each candidate execution item is determined. Terminology consistency indicates that the two contain the same or equivalent core terms. For example, there is a terminology consistency relationship between "sales data set" and "sales data". Hierarchical inclusion relationships indicate that there is a scope inclusion, pre-processing / post-processing, or category subordination relationship between the two. For example, "cleaned sales data" is the processed result of "sales data". Terminology consistency relationships include the same core terms, corresponding synonyms, correspondence between abbreviations and full names, correspondence between field names, or correspondence between business object names. Hierarchical inclusion relationships include the inclusion relationship between raw data and processed data, data sets and data subsets, files and file content, charts and chart data, and reports and report sections. In cases where the language expressions are not completely consistent, it can be used to determine whether there is a data transfer relationship between candidate execution items.
[0040] When the output of one candidate execution item has a terminological agreement or hierarchical inclusion relationship with the processing object of another candidate execution item, the candidate execution item that generates the output is determined as the front end of the data transfer relationship, and the candidate execution item that receives the processing object is determined as the back end of the data transfer relationship. For example, if "clean sales data" outputs "cleaned sales data" and "generate bar chart" requires sales data as input, then a data transfer relationship is established between "clean sales data" and "generate bar chart".
[0041] After establishing the initial data transfer relationships, the order of these relationships is further corrected based on the sequence words, parallel words, and condition words in natural language instructions to determine execution dependencies. Execution dependencies include not only data succession relationships but also execution order, sibling relationships, and conditional triggering relationships, enabling the structured workflow to simultaneously reflect "where the data comes from," "how the tasks are executed sequentially," and "under what conditions the tasks are executed." When candidate execution tasks at both ends of a data transfer relationship contain sequence words, the corresponding candidate execution tasks are determined as serial dependencies based on the execution order indicated by these sequence words. A serial dependency means that one candidate execution task needs to be executed after another. For example, in the task "First read sales data, then clean the data, and then generate a chart," the words "first," "then," and "after" specify the order in which the data is read, cleaned, and the chart generated; therefore, the three candidate execution tasks form a serial dependency.
[0042] When candidate execution items at both ends of a data transfer relationship contain parallel terms, such as "and," "with," or "as well as," the corresponding candidate execution items are identified as sibling dependencies. Sibling dependencies indicate that multiple candidate execution items exist side-by-side at the same execution level, can share the same preceding output, and can be executed in parallel or at the same level without any mutual waiting relationship. For example, in the task "After cleaning the data, simultaneously generate a bar chart and a line chart," "generating a bar chart" and "generating a line chart" both receive "the cleaned data," thus forming a sibling dependency between them.
[0043] When both ends of a data transfer relationship contain conditional terms, the judgment content corresponding to the conditional term is written into the dependency condition of the corresponding candidate execution item. Conditional terms include "if," "when," "when," etc., and the judgment content is the triggering content limited by the conditional term. For example, in "If the data is abnormal, generate a review report," "data abnormal" is the judgment content, and "generate a review report" is the candidate execution item controlled by this judgment content. Therefore, "data abnormal" is written into the dependency condition of "generate a review report."
[0044] In S300, each candidate execution item is determined based on execution dependencies. Before determination, the input source and language position of each candidate execution item are first determined. The input source refers to the data, files, text, charts, reports, or output content of the previous candidate execution item required for the execution of the candidate execution item. The language position refers to the position of the corresponding language segment of the candidate execution item in the natural language instruction. When determining the input source, it is determined whether the processing object of the candidate execution item can be provided by the output content of other candidate execution items based on the established data transfer relationship. If it can be provided by the front-end candidate execution item, then the output content of the front-end candidate execution item is its input source. If it cannot be provided by any front-end candidate execution item, it is necessary to further determine whether the candidate execution item corresponds to the starting content of the natural language instruction.
[0045] The initial content of a natural language instruction refers to the object, file, data source, or initial processing content directly given by the user in the instruction. For example, in the instruction "Generate a chart based on the uploaded sales table," the "uploaded sales table" is the initial content of the natural language instruction and can serve as the input source for subsequent candidate execution items. If a candidate execution item does not have an input source, but its processing object has already appeared in the initial content of the natural language instruction, it will not be considered a missing prerequisite item.
[0046] When a candidate execution item lacks a processing object with no input source and is not present in the starting content of a natural language instruction, the candidate execution item is determined to be a missing prerequisite, and its language position is recorded. A missing prerequisite indicates that the candidate execution item requires prerequisite input in the current workflow, but no candidate execution item providing that input has been found, nor has any object that can be directly used as input been found in the starting content of the natural language instruction. For example, if a user only enters "Generate chart and send email," where "Generate chart" requires data, tables, or statistical results as input, but the starting content of the natural language instruction does not provide a data source, and there are no candidate execution items such as "Read data" or "Import file" providing input, then "Generate chart" is determined to be a missing prerequisite, and its corresponding language position is recorded as the starting point for subsequent backtracking.
[0047] Candidate execution items whose output is not inherited by any subsequent candidate execution item and does not correspond to the result of a natural language instruction are identified as breakpoint items. Output not inherited by a candidate execution item means that the output generated by that candidate execution item is not processed or input by any subsequent candidate execution item; output not corresponding to the result of a natural language instruction means that the output is not the result requested by the user. For example, in the process of "reading sales data, cleaning outliers, writing a summary, and sending an email," if the output of "cleaning outliers" is "cleaned sales data," but the subsequent "writing a summary" is not recognized as needing to receive this output, and the user's final result is "email sent," then "cleaning outliers" will be temporarily designated as a breakpoint item. The purpose of breakpoint items is to prompt subsequent reassessment of whether a data inheritance relationship should be established or whether it should be used as the final output. Missing antecedent items are used to trigger backtracking and completion, while breakpoint items are used to trigger the re-establishment of data transmission relationships.
[0048] In S400, backtracking is performed on the natural language instructions based on the processing object corresponding to the missing precondition. Backtracking starts from the language position of the missing precondition and searches along the preceding segments of the natural language instructions for language segments containing the same processing object, a superior processing object, or a source description. The same processing object refers to an object in the preceding language segment that is the same as or equivalent to the processing object required by the missing precondition. A superior processing object refers to an object with a larger scope in the preceding language segment. A source description refers to a file source, data source, upload source, or business source appearing in the preceding language segment. For example, if the missing precondition is "cleaning outliers," and its surface processing object is "outliers," if the preceding language segment contains "reading sales data," then "sales data" can be considered the superior processing object of "outliers." If the preceding language segment contains "based on the uploaded sales table," then "uploaded sales table" can be considered the source description. By searching for the same processing object, superior processing object, and source description, missing input can be supplemented from the context already provided by the user.
[0049] After finding relevant language fragments, the pre-processing objects and pre-processing methods are determined based on the results. Pre-processing objects are those that can provide input for missing pre-processing items, such as "sales data," "uploaded sales forms," and "customer records." Pre-processing methods are the actions required to obtain the pre-processing object, such as "read," "import," "extract," and "collect."
[0050] Based on the pre-processing object and pre-processing method, candidate pre-processing items are determined, and their outputs are used as the input source for missing pre-processing items. For example, if the missing pre-processing item is "generate chart", and "uploaded sales table" is found, the candidate pre-processing item can be determined as "read uploaded sales table", whose output is "sales data table", and the "sales data table" is used as the input source for "generate chart".
[0051] After backtracking the missing prerequisites, the data transfer relationship is re-established based on the output content corresponding to the breakpoint. The output content of the breakpoint is used to determine if there are any candidate execution items that can receive that output. This determination involves not only comparing the literal terms for consistency but also considering the processing method of the candidate execution item to determine if it actually requires that output content as input. When a candidate execution item exists that can receive the output content, the data transfer relationship is re-established between the breakpoint and the candidate execution item. For example, "cleaning outliers" outputs "cleaned sales data," while "generating a chart" is identified as having the term "chart," but chart generation actually requires data as input. Therefore, it can be determined that "generating a chart" can receive "cleaned sales data," and the data transfer relationship between the two is re-established.
[0052] When no candidate execution item is found to receive the output, the output of the breakpoint item is compared with the result of the natural language instruction. The result of the natural language instruction is the final result that the user explicitly expects or desires in the natural language instruction, such as "report," "email," "chart," "export file," "archive result," etc. When the output of the breakpoint item matches the result of the natural language instruction, the breakpoint item is designated as the termination node.
[0053] When the output of a breakpoint item is inconsistent with the result of a natural language command, an output confirmation node is generated based on the output of the breakpoint item, and the output of the breakpoint item is used as the input source for the output confirmation node. The output confirmation node is used to retain this intermediate output in the structured workflow and prompt subsequent confirmation of its processing direction, avoiding the omission of intermediate results or their unfounded use as final results.
[0054] After missing prerequisites are backtracked and completed, and breakpoints are re-accepted or confirmed, a structured workflow is determined by connecting candidate execution items based on serial dependencies, sibling dependencies, and dependency conditions. Specifically, candidate execution items with serial dependencies are connected sequentially, candidate execution items with sibling dependencies are set at the same execution level, candidate execution items with dependency conditions are attached after their corresponding judgment conditions, and the output of each candidate execution item is connected to the input source of the corresponding subsequent candidate execution item. For example, for a natural language instruction such as "first read the uploaded sales data, clean outliers, and generate bar charts and line charts; if the data is abnormal, generate a review report and send the charts and report to the person in charge," the final structured workflow can be as follows: read the uploaded sales data as the starting input, output a sales data table, clean outliers receive the sales data table and output the cleaned sales data, generate bar charts and line charts receive the cleaned sales data and form sibling dependencies, generate a review report when the dependency condition for data abnormality is met, and finally execute the sending item using the charts and review report as input sources.
[0055] This implementation method generates task items from natural language instructions, then candidate execution items from these task items, subsequently determines data transfer relationships and execution dependencies from these candidate execution items, and finally repairs missing prerequisite items and breakpoint items to generate a structured workflow. Each processing result is passed down level by level, maintaining consistent logic throughout, thereby reducing workflow breaks caused by natural language omissions, inconsistent expression order, or unclaimed intermediate results.
[0056] In summary, by segmenting natural language instructions into sentences and forming task items based on action words, object words, and result words, complex user needs expressed in natural language can be transformed into task units with processing objects, processing methods, and output content. This avoids the problem of only processing structured task information. Furthermore, by merging and judging task items based on processing objects, processing methods, and output content, it is possible to distinguish between related processing in the same execution stage and continuous processing in different execution stages. By establishing data transfer relationships based on the correspondence between output content and processing objects, and combining sequence words, parallel words, and condition words to determine execution dependencies, data succession, execution order, parallel hierarchy, and conditional triggering relationships can be formed among candidate execution items. By identifying missing prerequisite items and breakpoint items, and backtracking and completing missing prerequisite items and re-accepting or confirming breakpoint items, workflow breaks caused by missing input sources or unaccepted intermediate outputs can be reduced, improving the efficiency and intelligence of workflow planning, and meeting the needs for automatic planning, orchestration, and execution in language-driven scenarios.
[0057] In another preferred embodiment based on the above embodiments, see [reference] Figure 3 As shown, this embodiment provides a language-driven intelligent workflow planning system for applying a language-driven intelligent workflow planning method, including: The planning receiving unit is configured to receive input natural language instructions, segment the natural language instructions into sentences, determine several language segments, determine the task items based on the action words, object words and result words in each language segment, and determine the processing object, processing method and output content of each task item. The first planning unit is configured to merge tasks with the same processing object, determine several candidate execution items, determine the data transmission relationship of each candidate execution item based on the correspondence between the output content of each candidate execution item and the processing object of other candidate execution items, and perform order correction on the data transmission relationship based on the order words, parallel words and condition words of natural language instructions to determine the execution dependency relationship. The second planning unit is configured to judge each candidate execution item based on execution dependencies. Candidate execution items that have no input source and do not correspond to the starting content of the natural language instruction are judged as missing preconditions. Candidate execution items whose output content is not inherited by the candidate execution item and do not correspond to the result content of the natural language instruction are judged as breakpoint items. The planning and processing unit is configured to backtrack in natural language instructions based on the processing objects corresponding to the missing prerequisites, re-establish data transmission relationships based on the output content corresponding to the breakpoints, and generate a structured workflow corresponding to the natural language instructions based on the candidate execution items after backtracking and re-establishing the data transmission relationships.
[0058] It is understandable that the above-mentioned language-driven intelligent workflow planning method and system have the same beneficial effects, and will not be elaborated further here.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A language-driven intelligent workflow planning method, characterized in that, include: The system receives input natural language instructions, segments the natural language instructions into sentences, determines several language segments, determines task items based on action words, object words, and result words in each language segment, and determines the processing object, processing method, and output content of each task item. Tasks with the same processing object are merged to determine several candidate execution items. Based on the correspondence between the output content of each candidate execution item and the processing object of other candidate execution items, the data transmission relationship of each candidate execution item is determined. The order of the data transmission relationship is corrected based on the order words, parallel words and condition words of the natural language instructions to determine the execution dependency relationship. Based on the execution dependency relationship, each candidate execution item is judged. Candidate execution items that do not have an input source and do not correspond to the starting content of the natural language instruction are judged as missing preconditions. Candidate execution items whose output content is not inherited by the candidate execution item and do not correspond to the result content of the natural language instruction are judged as breakpoint items. Based on the processing object corresponding to the missing prerequisite, the natural language instruction is backtracked, the data transmission relationship is re-established based on the output content corresponding to the breakpoint, and a structured workflow corresponding to the natural language instruction is generated based on the candidate execution items after backtracking and re-establishing the data transmission relationship.
2. The language-driven intelligent workflow planning method according to claim 1, characterized in that, When segmenting the natural language instructions into sentences, the following steps are included: The segmentation position is determined based on punctuation marks and sequential words. The content of sentences located at adjacent segmentation positions is identified as language segments. Action words for processing behaviors are extracted from each language segment, and nouns that have semantic association with the action words are identified as object words. The words representing the processing result, output form, or target state are used as result words. If the previous language segment contains an object word, the next language segment omits the object word, and the action word in the next language segment acts on the object word in the previous language segment, then the object word in the previous language segment is filled back into the next language segment.
3. The language-driven intelligent workflow planning method according to claim 2, characterized in that, When defining tasks and determining the objects, methods, and outputs of each task, the following should be included: When a language segment includes an action word, an object word, and a result word, the corresponding processing method, processing object, and output content are taken as task items. When a language segment contains multiple action words, the object word and result word corresponding to each action word are extracted based on the order in which each action word appears in the language segment. When an action word corresponds to multiple object words, corresponding task items are generated based on each object word.
4. The language-driven intelligent workflow planning method according to claim 3, characterized in that, When defining the tasks and determining the objects, methods, and outputs of each task, the following is also included: The action words in each language segment are used as the processing method, the object words that have a dominant relationship with the action words are used as the processing objects, and the result words that have a result correspondence with the action words or object words are used as the output content.
5. The language-driven intelligent workflow planning method according to claim 4, characterized in that, When determining the data transfer relationships for each candidate execution item, the following should be included: The output content of each candidate execution item is matched with the processing objects of other candidate execution items. Based on the term consistency relationship and hierarchical inclusion relationship between the output content and the processing object, the succession matching result of each candidate execution item is determined. Based on the succession matching result, the data transmission relationship of each candidate execution item is determined.
6. The language-driven intelligent workflow planning method according to claim 5, characterized in that, When determining execution dependencies, the following are included: When there are precedence terms between the candidate execution items at both ends of the data transmission relationship, the corresponding candidate execution items are determined as serial dependencies based on the execution order indicated by the precedence terms. When there are parallel terms in the candidate execution items at both ends of the data transmission relationship, the corresponding candidate execution items are determined as sibling dependencies; When there are condition words in the candidate execution items at both ends of the data transmission relationship, the judgment content corresponding to the condition word is written into the dependency condition of the corresponding candidate execution item.
7. The language-driven intelligent workflow planning method according to claim 6, characterized in that, When determining the absence of prerequisites, the following should be included: Determine the input source and language location of each candidate execution item, and compare the candidate execution items without an input source with the starting content of the natural language instruction; When the processing object of the candidate execution item does not appear in the beginning content of the natural language instruction, the candidate execution item is determined to be a missing prerequisite item, and the language position of the missing prerequisite item is recorded.
8. The language-driven intelligent workflow planning method according to claim 7, characterized in that, When backtracking in the natural language instruction based on the processing object corresponding to the missing antecedent, it includes: Using the language location of the missing preceding item as the starting point for backtracking, search for language fragments that contain the same processing object, superior processing object, or source description; Based on the search results, the pre-processing object and pre-processing method are determined, and the pre-processing candidate execution items are determined based on the pre-processing object and pre-processing method. The output content of the pre-processing candidate execution items is used as the input source of the missing pre-processing item.
9. The language-driven intelligent workflow planning method according to claim 8, characterized in that, When generating a structured workflow corresponding to the natural language instructions based on candidate execution items after backtracking and re-establishing data transfer relationships, the process includes: Based on the output content of the breakpoint event, determine whether there is a candidate execution event that receives the output content; When there is a candidate execution item that can receive the output content, the data transfer relationship is re-established between the breakpoint item and the candidate execution item; If there is no candidate execution item to receive the output content, the output content of the breakpoint item is compared with the natural language instruction; When the output content of the breakpoint item is consistent with the natural language instruction, the breakpoint item is determined as a termination node; otherwise, an output confirmation node is generated based on the output content of the breakpoint item, and the output content of the breakpoint item is used as the input source of the output confirmation node. Based on the serial dependencies, sibling dependencies, and dependency conditions, each candidate execution item is connected to determine the structured workflow.
10. A language-driven intelligent workflow planning system, used to apply the language-driven intelligent workflow planning method as described in any one of claims 1-9, characterized in that, include: The planning receiving unit is configured to receive input natural language instructions, segment the natural language instructions into sentences, determine several language segments, determine task items based on action words, object words and result words in each language segment, and determine the processing object, processing method and output content of each task item. The first planning unit is configured to merge tasks with the same processing object, determine several candidate execution items, determine the data transmission relationship of each candidate execution item based on the correspondence between the output content of each candidate execution item and the processing object of other candidate execution items, and perform order correction on the data transmission relationship based on the order words, parallel words and condition words of the natural language instructions to determine the execution dependency relationship. The second planning unit is configured to determine each candidate execution item based on the execution dependency relationship, and to determine the candidate execution item that has no input source and does not correspond to the start content of the natural language instruction as a missing prerequisite item, and to determine the candidate execution item whose output content is not inherited by the candidate execution item and does not correspond to the result content of the natural language instruction as a breakpoint item. The planning and processing unit is configured to backtrack in the natural language instruction based on the processing object corresponding to the missing prerequisite, re-establish the data transmission relationship based on the output content corresponding to the breakpoint, and generate a structured workflow corresponding to the natural language instruction based on the candidate execution items after backtracking and re-establishing the data transmission relationship.
Citation Information
Patent Citations
Intelligent command and data processing method based on workflow engine algorithm
CN120542818A