Complex query-oriented automatic interactive large language model pipeline arrangement method, system and equipment and storage medium
Through the automated interactive large-language model pipeline orchestration method, combined with large-language model and cost model, the execution pipeline is automatically generated and optimized, which solves the problem of insufficient complex query processing capabilities in the data lake, and realizes efficient, automated and intelligent query processing.
Patent Information
- Application Number
- CN202510100472.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art does not perform well when processing complex queries in data lakes, especially when multi-skip semantic retrieval and linking, multi-step logical reasoning, and multi-stage semantic analysis across data types, it is difficult to meet the needs of complex queries.
An automated interactive large-language model pipeline orchestration method for complex queries is adopted. By combining large-language models, predefined operator sets and cost models, the optimal execution pipeline is automatically generated and dynamically optimized and adjusted, so as to achieve efficient processing of complex queries in the data lake.
It significantly reduces the labor cost of complex query processing of data lakes, improves the standardization and execution efficiency of pipeline construction, and enhances query accuracy and system response speed.
Smart Images

Figure CN120030048A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information retrieval technology, and in particular to an automated interactive large-scale language model pipeline orchestration method, system, device and storage medium for complex queries. Background Art
[0002] With the advent of the big data era, data lakes, as a solution for centralized storage of large amounts of raw data, have become a core component of enterprise data management. Data lakes can store structured, semi-structured, and unstructured data, providing rich data resources for data analysis and business intelligence. However, despite the advantages of data lakes in data storage, they face a series of challenges in data analysis.
[0003] First, one of the data analysis problems in data lakes is the insufficient processing capabilities for complex queries. Existing data lake technologies mainly provide basic access operations for unstructured data and analytical queries for structured data, but they perform poorly in processing complex queries that require multi-hop semantic retrieval and linking across data types, multi-step logical reasoning, and multi-stage semantic analysis. These queries often require a deep understanding of data content, complex logical reasoning, and the ability to convert and integrate between different data formats. This requirement is far beyond the capabilities of traditional structured query languages (such as SQL), so even the advanced natural language to SQL (NL2SQL) method cannot meet these requirements.
[0004] Secondly, the emergence of large language models (LLMs) has brought new opportunities for data search and analysis. With their powerful semantic understanding and reasoning capabilities, they can theoretically handle complex data analysis tasks. However, large language models have also encountered challenges in practical applications, especially when dealing with complex queries that require complex task decomposition, pipeline orchestration, pipeline optimization, interactive execution, and self-reflection. These tasks are beyond the current reasoning capabilities of large language models because they require complex multi-step reasoning or "thinking chain" processes and in-depth understanding of the data in the data lake, which large language models themselves lack.
[0005] In addition, as described in the official documents of current large language model execution frameworks, such as Llamaindex and LangChain, existing large language model application methods rely on static, manually orchestrated execution pipelines when processing complex queries. Although these pipelines can decompose queries into subtasks and combine pre-programmed steps, retrieval steps, and prompt-based subtasks to obtain accurate answers, they have obvious limitations. Manually orchestrated pipelines rely on the user's professional skills and are usually very complex, involving hundreds of steps, which increases labor costs. At the same time, these pipelines are static and cannot be dynamically adjusted to cope with the failure of intermediate operations. For example, if the retrieval step produces irrelevant results, the subsequent steps need to be adjusted accordingly, otherwise the final answer may be wrong based on irrelevant information. In addition, although traditional retrieval-augmented generation (RAG) methods can improve the accuracy of information retrieval, they still face the problem of insufficient accuracy when processing queries that require multi-step reasoning and complex logical analysis. This is because these methods still rely on preset retrieval paths and fixed execution processes, and cannot adapt to changes and uncertainties that may occur during the query process. Moreover, these methods cannot effectively support aggregate analysis of large amounts of data, not just point queries.
[0006] Existing methods are also limited in functionality. Human-designed pipelines are usually only targeted at specific queries and cannot adapt to other complex queries. As a result, many queries are not covered by existing pipelines and can only use relatively basic pipelines, resulting in reduced accuracy.
[0007] In summary, there is an urgent need for a technical solution to efficiently process complex queries in a data lake. Summary of the invention
[0008] The purpose of the present disclosure is to provide an automated interactive large-scale language model pipeline orchestration solution for complex queries. By combining a large-scale language model, a predefined set of operators and a cost model, the optimal execution pipeline is automatically generated and dynamically optimized and adjusted according to the actual execution situation, so as to achieve efficient processing of complex queries in the data lake.
[0009] According to an embodiment of the present disclosure, an automated interactive large-scale language model pipeline orchestration method for complex queries is proposed, including:
[0010] receiving a query input by a user in natural language;
[0011] Extracting operators from a predefined set of operators using a large language model, and constructing a plurality of chained candidate pipelines with different reasoning paths corresponding to the query;
[0012] rewriting each chained candidate pipeline into a directed acyclic graph (DAG) structure, and identifying parallel opportunities in the rewriting process to write corresponding operators into parallel structures;
[0013] Use the cost model to estimate the cost of each operator in the DAG structure, and select at least one DAG candidate pipeline with the lowest overall cost as the optimal execution pipeline;
[0014] Execute the selected at least one optimal execution pipeline, and monitor the intermediate results and execution status in real time to dynamically and adaptively adjust the execution pipeline;
[0015] The outputs of each execution pipeline are integrated and the integrated result is returned to the user as the final query result.
[0016] In some implementations, the predefined set of operators includes some or all of the following operators: retrieve, scan, filter, sort, summarize, generate, refine, classify, translate, convert, evaluate, interpret, integrate, conceptualize, extract, plan, link, aggregate, verify, group, compare, and cluster.
[0017] In some embodiments, the cost model is used to estimate the cost of each operator based on the size of the input data and the cost function of the operator, wherein, when estimating the cost of the operator, the cardinality estimate is used as the size of the input data, and the cardinality estimate approximates the selectivity through a random sampling method.
[0018] In some embodiments, the method further comprises:
[0019] According to the sampled workload, the parameters of the cost function of the operators in the cost model are adjusted.
[0020] In some implementations, real-time monitoring of intermediate results and execution status to dynamically and adaptively adjust the execution pipeline includes:
[0021] Real-time monitoring of the intermediate results and current execution status generated during the execution process;
[0022] Based on the intermediate results and the execution status, evaluating the execution quality of the currently executed pipeline;
[0023] If the execution quality does not meet the expected quality requirement, stop executing the currently executed pipeline, regenerate at least one reference execution pipeline based on the collected intermediate results and using the large language model, and determine an adjusted execution pipeline from the at least one reference execution pipeline using the cost model.
[0024] In some implementations, during execution, a layer-by-layer execution strategy is adopted to execute the pipeline of the DAG structure.
[0025] In some embodiments, the method further comprises employing the following pre-fetching technique during execution:
[0026] When the operator corresponding to the search is executed, a primary search and multiple backup searches are generated, and different backup searches correspond to different search scopes directly or indirectly related to the query intention;
[0027] The plurality of backup searches are performed in parallel while processing the results returned in response to the primary search, and the results returned in response to the backup searches are received and stored.
[0028] In some implementations, integrating the outputs of various execution pipelines includes:
[0029] After each execution pipeline is completed, the output of each execution pipeline is explained using a large language model;
[0030] Analyze and merge the interpretation results of the outputs of each execution pipeline, and use the merged results as the integrated results.
[0031] In some embodiments, the method further comprises performing feedback optimization using a reward model:
[0032] Evaluate the relevance and impact of operators based on their output and execution status during execution;
[0033] Rewards are distributed based on the relevance and impact of the evaluation, and the reward information is used to fine-tune a large language model that extracts relevant operators from the query to enhance the chain thinking ability of the large language model.
[0034] According to an embodiment of the present disclosure, an automated interactive large-scale language model pipeline orchestration system for complex queries is proposed, wherein the system is used to implement the above method, and the system includes:
[0035] A query interface, used to receive a query input by a user in natural language and identify the user's query intent;
[0036] An operator set module, used to store predefined operator sets;
[0037] A pipeline generator, configured to extract operators from an operator set using a large language model, and construct a plurality of chained candidate pipelines with different reasoning paths corresponding to the query;
[0038] a pipeline rewriter for rewriting the chained candidate pipeline into a directed acyclic graph (DAG) structure and identifying parallel opportunities during the rewriting process to write the corresponding operators into a parallel structure;
[0039] An optimizer, for estimating the cost of each operator in the DAG structure using a cost model, and selecting at least one DAG candidate pipeline with the lowest overall cost as the optimal execution pipeline;
[0040] A pipeline executor, configured to execute at least one selected optimal execution pipeline according to a layer-by-layer execution strategy, and monitor intermediate results and execution status in real time to dynamically and adaptively adjust the execution pipeline;
[0041] The pipeline integrator is used to evaluate and integrate the output results of various execution pipelines;
[0042] A context manager to manage intermediate results in execution to make the information coherent and avoid exceeding the context length limit for large language models;
[0043] Indexing and storage modules for storing and indexing structured, semi-structured, and unstructured data in the data lake; and
[0044] A reward model for distributing rewards based on the relevance and impact of operators in a given state, fine-tuning large language models to enhance chain thinking capabilities.
[0045] According to one embodiment of the present disclosure, an electronic device is provided, the device comprising a memory and a processor, the memory being used to store computer instructions executable on the processor, the processor being used to implement any of the above methods when executing the computer instructions.
[0046] According to an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, any of the above methods is implemented.
[0047] The automated interactive large-scale language model pipeline orchestration scheme for complex queries proposed in the present disclosure enables the system to cope with various complex query scenarios based on automated pipeline orchestration and optimization mechanisms, greatly reducing the labor cost of complex query processing in the data lake. By introducing predefined operators as standardized building blocks, not only the standardization of pipeline construction is improved, but also unified optimization and management are facilitated. By rewriting the chain pipeline into a DAG structure that supports parallelism, the execution efficiency is further improved. Pipeline selection and adjustment based on the cost model, as well as dynamic and adaptive adjustment of the execution path through real-time monitoring of intermediate results and execution status, these multi-level execution and optimization mechanisms are conducive to significantly improving query efficiency and query accuracy, and reducing resource consumption. The present disclosure also further reduces execution delays and improves system response speed through a pre-fetch optimization mechanism, and uses a reward model for feedback optimization to continuously enhance the chain thinking ability of large language models, so that its performance in processing complex queries continues to improve.
[0048] The application of the automated interactive large-scale language model pipeline orchestration solution for complex queries proposed in this disclosure is conducive to the automation, intelligence and efficiency of complex query processing in the data lake.
[0049] Other features and advantages of the technical solution proposed in the present disclosure are described in detail below. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0051] Figure 1 A flowchart of an automated interactive large-scale language model pipeline orchestration method for complex queries according to an embodiment of the present disclosure is shown.
[0052] Figure 2 A schematic diagram of rewriting candidate pipeline 1 into a DAG structure according to an embodiment of the present disclosure is shown.
[0053] Figure 3 FIG. 4 is a schematic diagram showing rewriting the candidate pipeline 2 into a DAG structure according to an embodiment of the present disclosure.
[0054] Figure 4 A schematic diagram of dynamic adaptive adjustment of pipelines according to an embodiment of the present disclosure is shown.
[0055] Figure 5 A schematic diagram of a prefetch optimization technology according to an exemplary embodiment of the present disclosure is shown.
[0056] Figure 6 A context management schematic diagram according to an exemplary embodiment of the present disclosure is shown.
[0057] Figure 7 A schematic diagram of an automated interactive large-scale language model pipeline orchestration system for complex queries according to an exemplary embodiment of the present disclosure is shown.
[0058] Figure 8 A schematic diagram of a processing flow of a complex query according to an exemplary embodiment of the present disclosure is shown.
[0059] Fig. 9 A schematic diagram of the execution process of rewriting a chain structure execution pipeline into an R1 path in a DAG structure execution pipeline according to an embodiment of the present disclosure is shown.
[0060] Fig.10 A schematic diagram of the execution process of rewriting a chain structure execution pipeline into an R2 path in a DAG structure execution pipeline according to an embodiment of the present disclosure is shown.
[0061] Fig.11 It is a schematic diagram of the structure of an electronic device shown in at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0062] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0063] The disclosed embodiments may be applied to a computer system / server that may operate with numerous other general or special computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with a computer system / server include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, networked personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and the like.
[0064] Computer systems / servers may be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. In general, program modules may include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules may be located on local or remote computing system storage media including storage devices.
[0065] The core concept of the automated interactive large-scale language model pipeline orchestration solution for complex queries proposed in this disclosure is: using a predefined standardized operator set as the basic building block, using a large-scale language model to automatically extract relevant operators from the query and build an initial chain execution pipeline; rewriting the chain pipeline into a DAG structure to support parallel execution, and using a cost model to select the optimal execution path; real-time monitoring of intermediate results during execution, dynamically adjusting the pipeline according to the execution status, and improving efficiency through pre-fetching mechanisms and parallel execution; using a large-scale language model to assist in result interpretation and integration, and continuously optimizing system performance through a reward model. This design realizes the automation of the entire process of complex query processing from pipeline construction to execution to optimization, overcomes the limitations of static manual orchestration in existing technologies, and can improve the automation, intelligence, and efficiency of complex query processing in data lakes.
[0066] Figure 1 A flowchart of an automated interactive large-scale language model pipeline orchestration method for complex queries according to an embodiment of the present disclosure is shown. The method can be applied to complex queries of structured, semi-structured and unstructured data in a data lake. As shown in the figure, the method includes steps 1 to 6.
[0067] Step 1: Receive a query input by a user in natural language.
[0068] It can accept users to express query requirements in everyday language rather than structured query statements (such as SQL). Queries can involve multiple data sources, require multi-step reasoning, or analyze across different types of data. For example, "Analyze the per capita expenditure of each department in the fourth quarter of 2024 and find abnormal expenditure items that exceed the historical average" is a typical natural language query.
[0069] In this embodiment, semantic understanding technology based on pre-trained language models such as GPT and BERT can be used to perform intent recognition and semantic analysis on queries. Supported query types include but are not limited to: multi-hop semantic retrieval queries, multi-step logical reasoning queries, and multi-stage semantic analysis queries.
[0070] The natural language-based query method lowers the user threshold, allowing non-technical users to easily perform complex data analysis. At the same time, the flexibility of natural language also provides the system with more semantic information, which is helpful for subsequent intelligent processing.
[0071] Step 2: extract operators from a predefined set of operators using a large language model, and construct multiple chained candidate pipelines with different reasoning paths corresponding to the query.
[0072] After obtaining the user query, operators can be extracted according to the query intent and multiple chained candidate pipelines can be constructed. This embodiment achieves standardization and modularization of query processing by pre-defining operators, which facilitates subsequent pipeline construction and optimization. The combination of operators can cope with various complex query scenarios, and its standardized characteristics facilitate unified optimization and management. The chained pipeline is a linear structure formed by connecting operators in sequence. The randomness of large language models can be used to generate multiple candidate pipelines with different reasoning paths.
[0073] In some implementations, the predefined operator set includes some or all of the following operators:
[0074] Retrieval operators: Fetch relevant data from external sources to answer queries, enhancing the timeliness or specific details of responses from large language models;
[0075] Scan operators: load documents or tables and enumerate their elements, often applying filtering or initial transformations while scanning;
[0076] Filter operators: remove irrelevant information from retrieved or generated data based on specific criteria, similar to the select operator in SQL;
[0077] Sorting operator: Sorts the retrieved documents or information according to a given criterion, similar to the SQL ORDER BY operator, but with semantic sorting capabilities;
[0078] Summarize operator: condenses longer text into shorter summaries to optimize context consumption and readability;
[0079] Generative operators: generate coherent text based on input, usually used to retrieve relevant information and perform a series of analyses before a large language model generates the final response;
[0080] Refinement Operator: Improve or adjust the input text to better meet requirements or improve coherence;
[0081] Classification operators: classify or label input text or entities in it, and perform simple classification tasks;
[0082] Translation Operator: Converts text from one language to another;
[0083] Transformation operators: convert data from one format to another, such as converting a structured table to a text description;
[0084] Evaluation operator: evaluates the quality of different inputs according to a criterion, usually through the output or internal log-probability of the LLM;
[0085] Explanation operator: provides explanation or reasoning for the model's response or decision, enabling the model's self-reflection;
[0086] Integrator: Receives information from multiple sources, judges their accuracy, and integrates them into a coherent response;
[0087] Conceptualization operators: extract and identify key concepts described in the text, simplifying complex information by locating the main ideas;
[0088] Extraction operator: Identifies and extracts specific information from the data, similar to the projection operator in SQL;
[0089] Plan Operators: Automatically orchestrate a series of operators to execute complex queries;
[0090] Link operators: process heterogeneous data sets, identify and link data containing relevant information for the current query;
[0091] Set operators: perform set operations on data sets, such as union, intersection, and complement;
[0092] Verification Operator: Verifies the accuracy and reliability of generated or retrieved information;
[0093] Grouping operators: divide the data into subgroups based on specific attributes, allowing summary statistics to be calculated for each group;
[0094] Comparison: Evaluates two input values and returns a Boolean result that satisfies the comparison criteria;
[0095] Aggregation operator: Calculates aggregate results of data, such as count, sum, maximum value, etc.
[0096] Those skilled in the art may select some or all of the operators to form a predefined operator set as needed, and may also design other operators as needed.
[0097] Implementations of operators can include large language model-based implementations and pre-programmed implementations.
[0098] Implementation based on large language models: Executed by designing specific prompts for large language models. For example, the execution of overview and extraction operators includes inputting the operator and its description as prompts, guiding the large language model to perform the operation and return the result. For operators that require data retrieval, such as semantic retrieval or semantic clustering, data is retrieved first, and then the large language model is guided to perform the operation.
[0099] Pre-programmed implementations: Operators that do not involve large language models can be directly executed using pre-defined programs. For example, a retrieval operator might perform a nearest neighbor search of document embeddings or a keyword search using the BM25 algorithm, and a filtering operator might apply programmed rules to implement, for example, filtering by attribute value.
[0100] Large language models can identify the required types of operations from queries through semantic understanding. For example, for the query "Analyze the per capita expenditure of each department in the fourth quarter of 2024, and find out the abnormal expenditure items that exceed the historical average", the large language model can identify that "the fourth quarter of 2024" requires a filter operator to filter the time range, that "the per capita expenditure of each department" requires a search operator to obtain expenditure data, and an aggregation operator to calculate the per capita value, and that "find the abnormal expenditure items that exceed the historical average" requires a search for historical data, an aggregation operator to calculate the average, and a comparison operator to identify anomalies.
[0101] Taking the query "Analyze the per capita expenditure of each department in the fourth quarter of 2024 and find abnormal expenditure items that exceed the historical average" as an example, the following two chain candidate pipelines may be generated:
[0102] Candidate Pipeline 1:
[0103] Retrieve (department expenditure data) -> Filter (2024Q4) -> Extract (per capita expenditure) -> Retrieve (historical data) -> Aggregate (calculate mean) -> Compare (identify anomalies)
[0104] Candidate Pipeline 2:
[0105] Retrieve (2024Q4 expenditure data) -> Filter (by department) -> Extract (per capita expenditure) -> Retrieve (historical mean) -> Compare (identify anomalies)
[0106] The randomness of large language models can be utilized to generate multiple candidate pipelines with different inference paths. The semantic understanding of large language models is beneficial for constructing reasonable operation sequences, and each candidate pipeline contains a complete operation chain required to solve the query. The operators in each pipeline are executed in topological order, that is, the output of the previous operator serves as the input of the next operator.
[0107] By generating simple and intuitive initial chained pipelines, it provides a basis for subsequent DAG rewriting and optimization.
[0108] Step 3, rewrite each chained candidate pipeline into a directed acyclic graph (DAG) structure, and identify parallel opportunities during the rewriting process to write the corresponding operators as parallel structures.
[0109] The chained candidate pipeline can be rewritten into a DAG structure to support parallel execution. The rewriting process can be guided by a large language model through prompts, and identify the operations that can be executed in parallel during the conversion process. For example, the retrieval operators without dependencies can be executed in parallel, while operations such as extraction and aggregation need to wait for the previous operations to complete.
[0110] For the query "Analyze the per capita expenditure of each department in the fourth quarter of 2024 and find out the abnormal expenditure items that exceed the historical average level", the two candidate pipelines given above can be rewritten into a DAG.
[0111] Figure 2 FIG. shows a schematic diagram of rewriting Candidate Pipeline 1 into a DAG structure according to an embodiment of the present disclosure. Figure 3 FIG. shows a schematic diagram of rewriting Candidate Pipeline 1 into a DAG structure according to an embodiment of the present disclosure.
[0112] The above DAG structure supports parallel execution while maintaining operation dependencies, can improve execution efficiency, and provides a basis for subsequent cost evaluation and dynamic adjustment.
[0113] Step 4: Use the cost model to estimate the cost of each operator in the DAG structure, and select at least one DAG candidate pipeline with the lowest overall cost as the optimal execution pipeline.
[0114] In this embodiment, a cost model is used to evaluate candidate pipelines of a DAG structure. Cost evaluation can take into account the computational characteristics of operators, data size, resource consumption, etc., to achieve accurate cost estimation.
[0115] In some embodiments, the cost model is used to estimate the cost of each operator based on the size of the input data and the cost function of the operator, wherein, when estimating the cost of the operator, the cardinality estimate is used as the size of the input data, and the cardinality estimate approximates the selectivity through a random sampling method.
[0116] The cost model can define a resource-specific cost function for each operator, which can be based on two core parameters: the computational complexity of the operator and the input data size (cardinality). Different types of operators consume different resources, where structured data operators mainly involve I / O resource consumption, while semi-structured and unstructured data operators mainly consume CPU / GPU resources for performing embedding calculations and attention calculations. Those skilled in the art can design a specific cost function form as needed.
[0117] The parameterization of the cost function depends on the operator type. For pre-programmed operators, it has a fixed computational complexity coefficient; for operators based on large language models, the cost is approximately linear in the output size.
[0118] Constructing a cost model is equivalent to determining specific parameter values for the cost function of each operator, including computational complexity coefficients, resource consumption weights, etc.
[0119] In data lake analysis or unstructured data analysis, the scale of intermediate results may vary significantly, so cardinality estimation is of great significance to the cost model. According to this embodiment, a random sampling method (for example, sampling 1% of the data) can be used to approximate the selectivity and thus estimate the cardinality of the results. By executing queries on samples, the system is able to estimate the cardinality of the results with relatively little time consumption. This embodiment is particularly suitable for semantically related operators and unstructured data, because these scenarios are difficult to use histograms in traditional databases or learning-based data distribution methods. When multiple candidate pipelines are given, the cost of each pipeline can be estimated, and the cardinality of the intermediate data can be predicted to select the optimal execution path.
[0120] According to some embodiments of the present application, the parameters of the cost function of the operator in the cost model can be adjusted according to the sampled workload to ensure that the cost can be accurately estimated in actual execution. See the relevant introduction below for details.
[0121] The number of optimal execution pipelines finally selected can be set. For example, if K=1 is set, a candidate pipeline with the lowest overall cost can be selected as the optimal execution pipeline. If K=2 is set, two candidate pipelines with the lowest overall cost can be selected as the optimal execution pipelines.
[0122] The cost model also allows the execution order of operators in the pipeline to be dynamically adjusted without affecting the accuracy of the results, thereby further improving execution efficiency. Intelligent data lake analysis usually involves multiple optimization goals, such as minimizing query latency, maximizing result accuracy, or minimizing the use of large language models (LLMs). In order to meet different optimization goals, the cost model can be adjusted to adapt to different optimization criteria, and a composite cost model can be used, that is, a weighted average of the costs from different sources is performed to achieve multi-objective optimization.
[0123] With the support of the unified cost model, operators can be dynamically scheduled based on resource thresholds and cost estimation results, so as to achieve parallel execution of operators as much as possible without exceeding resource constraints.
[0124] In addition, in order to reduce the delay caused by the materialization of intermediate results of operators, operators located between multiple pipeline processing stages can be given priority during execution, so that downstream operators can start processing some outputs earlier and improve execution efficiency.
[0125] Through the above-mentioned detailed pipeline optimization strategy, this embodiment can maximize the efficiency of query execution while ensuring the accuracy of the results.
[0126] Step 5: execute at least one selected optimal execution pipeline, and monitor the intermediate results and execution status in real time to dynamically and adaptively adjust the execution pipeline.
[0127] If the number of optimal execution pipelines finally selected in step 4 is greater than or equal to 2, these optimal execution pipelines are executed in parallel during execution.
[0128] In some embodiments, a layer-by-layer execution strategy is used to execute the pipeline of the DAG structure. According to this embodiment, it can be ensured that the operators are executed in the topological order of the DAG, and the operator will only start to execute when all the pre-dependencies of an operator are completed. The layer-by-layer execution strategy is conducive to ensuring the correctness and reliability of the data flow.
[0129] In some implementations, pre-fetching technology may be used to further improve query execution efficiency. The pre-fetching technology is described in detail below.
[0130] During execution, the execution results and status of each operator can be continuously tracked, and dynamic adjustments can be made when necessary to ensure the accuracy and efficiency of the query results.
[0131] In some implementations, real-time monitoring of intermediate results and execution status to dynamically and adaptively adjust the execution pipeline includes:
[0132] Real-time monitoring of the intermediate results and current execution status generated during the execution process;
[0133] Based on the intermediate results and the execution status, evaluating the execution quality of the currently executed pipeline;
[0134] If the execution quality does not meet the expected quality requirement, stop executing the currently executed pipeline, regenerate at least one reference execution pipeline based on the collected intermediate results and using the large language model, and determine an adjusted execution pipeline from the at least one reference execution pipeline using the cost model.
[0135] According to this embodiment, the intermediate results and execution status generated during the execution process can be continuously monitored. The intermediate results refer to the data content and scale of the operator output, for example, whether the data returned by the retrieval operation is relevant, and whether the amount of data after the filtering operation is reasonable; the execution status usually includes whether the operator is successfully completed, resource consumption, execution time, etc. The relevance, accuracy and execution status of the intermediate results can be comprehensively considered to evaluate whether the execution quality of the current path meets the expected quality requirements. For example, when the retrieval operator does not return the expected relevant data, the amount of data returned by the filtering operator is abnormal, the operator execution fails or times out, and / or the resource consumption exceeds expectations, it can be judged that the current execution quality does not meet the expected quality requirements, then the current execution path is stopped, and a new execution path exploration is started based on the collected intermediate results and using a large language model.
[0136] The following DAG structure of the complex query "Analyze the per capita expenditure of each department in the fourth quarter of 2024 and find abnormal expenditure items that exceed the historical average" is used as an example to illustrate dynamic adaptive adjustment. Figure 4 shown.
[0137] First, execute layer by layer according to the DAG structure. First, execute "Search (department expenditure data)" and "Search (historical data)" in parallel. After executing "Search (department expenditure data)", the system will pass the data to the "Filter (2024Q4)" operator. During the process, the system monitors the intermediate results and execution status in real time.
[0138] Through real-time monitoring, it can be monitored that "Retrieve (department expenditure data)" returns a large amount of data, and it can also be monitored that the execution speed of the "Filter (2024Q4)" operator is abnormally slow, consumes a large amount of CPU and memory resources, and is in a "waiting" state for a long time.
[0139] Based on the above monitored information, it can be considered that the data quality is low and the execution efficiency is low, and it is judged that the execution quality of the current execution path does not meet the expected quality requirements.
[0140] At this time, dynamic adjustment can be triggered to stop executing the "Filter (2024Q4)" operator and its subsequent "Extract (Per Capita Expenditure)" operator, and the collected intermediate results and execution status can be passed as input to the large language model. The input information can include the output results of "Retrieve (Department Expenditure Data)", the error information generated when executing the "Filter (2024Q4)" operator, resource consumption information, execution time information, etc.
[0141] Based on this information and the user's initial query intent, large language models can perform semantic analysis and reasoning, determine that the current execution path has data quality problems, and generate multiple reference execution paths. For example, it is believed that the "verification" operator can be used to verify the data before filtering, or that the parameters of the "retrieval" operator can be adjusted to specify the data field to be retrieved to avoid retrieving invalid data, or that the "classification" operator can be used to classify departments, and then filter the classified results, etc. The cost model can estimate the multiple reference execution paths generated by the large language model, and estimate the overall cost of each reference execution path based on the operator type, input data size, and cost function.
[0142] Next, the adjustment target can be determined from multiple reference execution paths based on the estimation results of the cost model. Suppose that according to the estimation results, the search (department expenditure data) -> verification -> filtering (2024Q4) -> extraction (per capita expenditure) is updated to the execution path, and the "search (historical data)" and its downstream operators in the original execution path are still executed in parallel.
[0143] In dynamic adjustment, a large language model can analyze the problem and generate multiple reference execution pipelines based on query intent semantic analysis. The cost model can estimate the overall cost of the reference execution pipeline and determine the adjustment plan.
[0144] After adjusting the pipeline to continue execution, you can continue to monitor the intermediate results and execution status and trigger dynamic adjustments again if necessary.
[0145] Through the dynamic adjustment mechanism that combines a large language model with a cost model, it is possible to promptly detect and handle execution anomalies. For example, when the retrieval results are irrelevant, the pipeline can be adjusted in a timely manner, and the downstream steps of the path can be pruned early to avoid continuing execution based on erroneous information, thereby saving resources and improving efficiency.
[0146] Step 6: Integrate the outputs of each execution pipeline and return the integrated result to the user as the final query result.
[0147] In some implementations, integrating the outputs of various execution pipelines includes:
[0148] After each execution pipeline is completed, the output of each execution pipeline is explained using a large language model;
[0149] Analyze and merge the interpretation results of the outputs of each execution pipeline, and use the merged results as the integrated results.
[0150] For example, the output of each execution pipeline can be input to the interpretation operator, the output of the interpretation operator passes through the comparison, verification, summary and merger operators in sequence, and then the output of the merger operator is returned to the user as the integrated result.
[0151] For example, for the query "Analyze the per capita expenditure of each department in the fourth quarter of 2024 and find abnormal expenditure items that exceed the historical average", the query results can be obtained through the following two different channels:
[0152] Execution pipeline 1: Calculate abnormal expenditure items based on department expenditure data;
[0153] Execution pipeline 2: Analyze the quarterly expenditure data to find abnormal expenditure items.
[0154] A large language model can be used to interpret the output of each pipeline, compare the similarities and differences between the two results, verify the rationality of the results and integrate them into a consistent final answer. Depending on the query intent, this final result may be an aggregated dataset, a calculated value, a piece of text, or other forms of answers.
[0155] Through integration, the accuracy and reliability of query results are further improved, ensuring the integrity and consistency of the final output.
[0156] The automated interactive large-scale language model pipeline orchestration method for complex queries proposed in this embodiment is based on an automated pipeline orchestration and optimization mechanism, can cope with various complex query scenarios, and greatly reduces the labor cost of complex query processing in the data lake. By introducing predefined operators as standardized building blocks, not only the standardization of pipeline construction is improved, but also unified optimization and management are facilitated. By rewriting the chained execution pipeline into a DAG structure that supports parallelism, the execution efficiency is further improved. Pipeline selection and adjustment based on the cost model, as well as dynamic and adaptive adjustment of the execution path through real-time monitoring of intermediate results and execution status, improves query efficiency and query accuracy, and reduces resource consumption.
[0157] In some embodiments, the following pre-fetching technology may be used during execution: when executing the operator corresponding to the search, a main search and multiple backup searches are generated, and different backup searches correspond to different search scopes that are directly or indirectly related to the query intent; when processing the results returned in response to the main search, the multiple backup searches are executed in parallel, and the results returned in response to the backup searches are received and saved.
[0158] This implementation method introduces pre-fetch optimization technology, which is beneficial to reducing delays caused by retrieval failures and improving overall execution efficiency.
[0159] During the execution process, the search operator may not be able to retrieve the expected results. At this time, re-planning and retrieval are required, resulting in delays in the query processing flow. In order to reduce delays caused by retrieval failures, according to this embodiment, idle computing resources can be used to pre-execute in the background and obtain potentially useful information. When the search operator generates a main search, multiple backup searches can be generated at the same time. Backup searches can cover search ranges of multiple granularities, such as exact match searches directly related to the query intent, such as "2024 Q4 department expenditure data", and indirectly related extended searches such as "2024 department expenditure data", "department quarterly expenditure summary" and "department expenditure details", providing support for possible pipeline adjustments through multi-level retrieval strategies.
[0160] While processing the results returned by the main search, these backup searches can be executed in parallel in the background. If the initial results are not ideal and the execution pipeline needs to be adjusted, the results of the backup searches can be obtained directly to avoid additional waiting time. This pre-fetching mechanism significantly reduces the delay caused by failed searches, improves query processing efficiency, and effectively utilizes computing resources that may be idle during large language model analysis and reasoning.
[0161] Figure 5 FIG. 1 is a schematic diagram of a prefetch optimization technique according to an exemplary embodiment of the present disclosure. In order to more clearly demonstrate the effect of the prefetch optimization technique in this embodiment, Figure 5 A comparison of the execution flow with and without prefetch optimization is shown. Figure 5 The upper middle portion corresponds to the execution process without prefetch optimization, and the lower portion corresponds to the execution process with prefetch optimization.
[0162] In the execution without prefetch optimization, when executing the main search O 4 After receiving the returned results, the system needs to wait for the intermediate results and execution status returned by LLM (Large Language Model) analysis. LLM analysis believes that the main search O 4 If the search fails, the pipeline is guided to adjust to perform the search O 3 Next, execute O 3 Retrieve and receive the returned results, and continue to wait for LLM to analyze and guide subsequent pipeline adjustments.2 The retrieval was successful. The whole process was delayed several times due to multiple retrieval failures.
[0163] During execution with prefetch optimization, when executing the main search O 3 After receiving the returned results, LLM is used to process the results and adjust the pipeline, while performing alternate retrieval in parallel. 3 and O 2 When LLM finishes processing the main search 4 After the results are returned and the pipeline is adjusted, the target pipeline O 3 and O 2 The execution has been completed, so there is no need to wait for the search to be executed, but the backup search O can be obtained directly 3 and O 2 The whole process fully utilizes computing resources and significantly reduces execution delay.
[0164] This embodiment significantly improves query execution efficiency and result accuracy through a multi-level execution monitoring and adjustment mechanism combined with pre-fetch optimization. The system can quickly respond to execution anomalies, dynamically adjust execution strategies, and reduce delays caused by adjustments through the pre-fetch mechanism.
[0165] In some implementations, parameters of the cost function of the operator in the cost model may be adjusted based on the sampled workload to ensure accurate cost estimation in actual execution.
[0166] According to this embodiment, a dynamic parameter adjustment mechanism based on sampled workload is introduced to improve the accuracy and adaptability of the cost model. Specifically, a portion of execution data can be randomly extracted from the actual operation process to obtain a sampled workload. These sampled data can reflect the performance characteristics of the system in real scenarios. By analyzing the sampled workload, deviations in the current cost model can be found, for example, the cost of certain operators is overestimated or underestimated, so that the parameters of the cost model can be adjusted in time. In one example, the actual cost can be calculated based on the collected sampled workload and compared with the cost predicted by the cost model to calculate the prediction error of the model, and further adjust the parameters of the model through machine learning. By dynamically adjusting the cost model according to the actual operating conditions, the cost model can better adapt to different data characteristics and execution environments.
[0167] In some embodiments, the method further comprises performing feedback optimization using a reward model:
[0168] Evaluate the relevance and accuracy of each operator based on its output and execution status during execution;
[0169] Rewards are distributed based on the relevance and accuracy of the evaluation, and the reward information is used to fine-tune a large language model that extracts relevant operators from the query to enhance the chain thinking ability of the large language model.
[0170] According to this embodiment, the reward model can be used to continuously optimize the performance of the system. Taking the query "Analyze the per capita expenditure of each department in the fourth quarter of 2024, and find out the abnormal expenditure items that exceed the historical average" as an example, the execution of each operator can be evaluated. After the retrieval operator is executed, the relevance of the retrieval result to the query is evaluated. If accurate expenditure records are obtained when retrieving "department expenditure data", a higher reward is assigned; if irrelevant data is retrieved, a lower reward is assigned. After the filter operator is executed, the accuracy of the filter result can be evaluated. If the "fourth quarter of 2024" data is correctly filtered out, a higher reward is assigned; if relevant data is omitted or irrelevant data is included, a lower reward is assigned.
[0171] Based on these evaluation results, the large language model used to extract operators from queries can be updated. For example, if the model is found to frequently extract poor operator combinations, the model's selection strategy for operators can be adjusted through feedback optimization. This continuous optimization mechanism can enhance the chain thinking ability of the large language model and improve the accuracy of subsequent query processing.
[0172] According to some embodiments, context management techniques can be used for the context length limit of large language models (e.g., 8,192 tokens) to ensure that only context information closely related to the current state is retained during execution. During execution, context can be managed in the following ways: using summary operators to compress and summarize intermediate execution information; using explanation operators to compress the context of each execution path; and only retaining information related to the current execution state.
[0173] For example, when executing the query "Analyze the per capita expenditure of each department in the fourth quarter of 2024, and find out the abnormal expenditure items that exceed the historical average", after using the retrieval operator to obtain a large amount of raw data, the overview operator can be used to generate a data summary; the intermediate results produced by multiple execution paths are compressed through the explanation operator; when integrating the results, only key information is retained to avoid exceeding the context restrictions.
[0174] The context management mechanism can effectively reduce the execution delay caused by context restrictions, enabling the system to efficiently process complex queries that require multi-step reasoning and large amounts of data processing.
[0175] Figure 6A context management schematic diagram according to an exemplary embodiment of the present disclosure is shown. Multiple context units can be maintained during execution, and these units are used to store intermediate results and state information generated during the execution of operators. When executing to a step that requires a large language model for reasoning (for example, the "interpretation" and "integration" operators in the figure), the large language model can read the information in the current context unit and analyze, reason or perform corresponding operations based on this information. In order to avoid the input length of the large language model exceeding the limit, the system can use the "interpretation" operator to compress and summarize the context information and only pass the key information to the subsequent steps. In addition, the context unit will also be passed between different operators as the execution process progresses, thereby ensuring the consistency and integrity of the information. When the execution process requires result integration, the system can read the information in multiple context units and use the integration operator to integrate information from multiple sources into a consistent response.
[0176] This embodiment uses a large language model in multiple processing steps. For different tasks such as extracting operators from queries, generating execution pipelines, and interpreting results, the same large language model can be used, or different large language models can be used according to the characteristics of the tasks. The solution proposed in this embodiment supports flexible configuration to meet the needs of different scenarios.
[0177] The present disclosure also proposes an automated interactive large-scale language model pipeline orchestration system for complex queries, the system being used to implement the method described above, the system comprising:
[0178] A query interface, used to receive a query input by a user in natural language and identify the user's query intent;
[0179] An operator set module, used to store predefined operator sets;
[0180] A pipeline generator, configured to extract operators from an operator set using a large language model, and construct a plurality of chained candidate pipelines with different reasoning paths corresponding to the query;
[0181] a pipeline rewriter for rewriting the chained candidate pipeline into a directed acyclic graph (DAG) structure and identifying parallel opportunities during the rewriting process to write the corresponding operators into a parallel structure;
[0182] An optimizer, for estimating the cost of each operator in the DAG structure using a cost model, and selecting at least one DAG candidate pipeline with the lowest overall cost as the optimal execution pipeline;
[0183] A pipeline executor, configured to execute at least one selected optimal execution pipeline according to a layer-by-layer execution strategy, and monitor intermediate results and execution status in real time to dynamically and adaptively adjust the execution pipeline;
[0184] The pipeline integrator is used to evaluate and integrate the output results of various execution pipelines;
[0185] A context manager to manage intermediate results in execution to make the information coherent and avoid exceeding the context length limit for large language models;
[0186] Indexing and storage modules for storing and indexing structured, semi-structured, and unstructured data in the data lake; and
[0187] A reward model for distributing rewards based on the relevance and impact of operators in a given state, fine-tuning large language models to enhance chain thinking capabilities.
[0188] Figure 7 A schematic diagram of an automated interactive large-scale language model pipeline orchestration system for complex queries according to an exemplary embodiment of the present disclosure is shown. As shown in the figure, the system mainly includes modules such as application layer, query interface, planner, executor, data index, reward model and data storage.
[0189] Applications supported by the system include data analysis, data discovery, information extraction, retrieval enhanced generation (RAG) and machine learning, which can explore data value from multiple angles.
[0190] The query interface can receive complex queries in natural language format. The planner includes two types of operators: semantic operators and pre-programmed operators. Operators refer to the operators mentioned above, and pipeline planning is performed through the optimizer (including pipeline generator, rewriter and integrator).
[0191] The executor consists of two parts: the operator executor and the interactive executor. The operator executor implements the physical execution of the operator and improves efficiency through parallel execution. The interactive executor handles unexpected intermediate results through the pipeline adjuster and context manager.
[0192] Data indexing is divided into vector indexing (supporting text and image embedding) and obtaining external data (through retrieval and scanning operators). The reward model provides feedback based on the accuracy of the execution results for fine-tuning large language models.
[0193] Data storage supports structured (tables), semi-structured (JSON, XML, etc.) and unstructured (documents, images, videos, etc.) data.
[0194] Figure 7 The system shown supports flexible complex queries, multi-step planning based on predefined operators, pipeline transformation from chain to DAG, optimization based on cost model, and the ability to process heterogeneous data.
[0195] Figure 8FIG. 2 shows a schematic diagram of a processing flow of a complex query according to an exemplary embodiment of the present disclosure. Figure 8 As shown, for the complex query in natural language form "What is the average height of New York Knicks players attending Villanova University?", the system performs the following processing.
[0196] First, the system generates the initial pipeline - the chain structure execution pipeline:
[0197] R1 path: retrieve->gather->verify->generate;
[0198] R2 path: Retrieval->Extraction->Retrieval->Verification->Aggregation.
[0199] The system rewrites the chain structure execution pipeline into the execution process of the R1 path in the DAG structure execution pipeline as follows: Fig. 9 As shown, the execution process of the R2 path is as follows Fig.10 shown.
[0200] In the layer-by-layer execution stage, the search for the R1 path ("New York Knicks players who attended Villanova University") failed to obtain the target result, and the system pruned its subsequent steps early. The search for the R2 path ("New York Knicks players") successfully obtained data, and after executing the entire process, the average height of the players was found to be 1.91 meters.
[0201] The above examples demonstrate the system's automated pipeline orchestration and interactive execution capabilities when processing complex queries, including features such as multi-path parallel processing, dynamic adjustment, and result integration.
[0202] Fig.11 An electronic device provided for at least one embodiment of the present disclosure includes a memory and a processor, wherein the memory is used to store computer instructions executable on the processor, and the processor is used to implement the automated interactive large-scale language model pipeline orchestration method for complex queries described in any embodiment or implementation of the present disclosure when executing the computer instructions.
[0203] At least one embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the automated interactive large-scale language model pipeline orchestration method for complex queries as described in any embodiment or implementation of the present disclosure.
[0204] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0205] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0206] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0207] The embodiments of the subject matter and functional operations described in this specification may be implemented in the following: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules in computer program instructions encoded on a tangible non-temporary program carrier to be executed by a data processing device or to control the operation of the data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information and transmit it to a suitable receiver device for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0208] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits), and the apparatus can also be implemented as special purpose logic circuits.
[0209] Computers suitable for executing computer programs include, for example, general and / or special microprocessors, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, the computer will also include one or more large-capacity storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to this large-capacity storage device to receive data from it or to transmit data to it, or both. However, the computer does not necessarily have such a device. In addition, the computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0210] Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0211] Although this specification includes many specific implementation details, these should not be interpreted as limiting the scope of any invention or the scope of protection claimed, but are mainly used to describe the features of the specific embodiments of specific inventions. Certain features described in multiple embodiments in this specification may also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may work in certain combinations as described above and even initially claim protection, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may point to a sub-combination or a variation of a sub-combination.
[0212] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or requiring that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
[0213] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In some implementations, multitasking and parallel processing may be advantageous.
[0214] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. An automated interactive large-scale language model pipeline orchestration method for complex queries, characterized in that: include: receiving a query input by a user in natural language; Extracting operators from a predefined set of operators using a large language model, and constructing a plurality of chained candidate pipelines with different reasoning paths corresponding to the query; rewriting each chained candidate pipeline into a directed acyclic graph (DAG) structure, and identifying parallel opportunities in the rewriting process to write corresponding operators into parallel structures; Use the cost model to estimate the cost of each operator in the DAG structure, and select at least one DAG candidate pipeline with the lowest overall cost as the optimal execution pipeline; Execute the selected at least one optimal execution pipeline, and monitor the intermediate results and execution status in real time to dynamically and adaptively adjust the execution pipeline; The outputs of each execution pipeline are integrated and the integrated result is returned to the user as the final query result.
2. The method according to claim 1, characterized in that: The predefined set of operators includes some or all of the following operators: retrieve, scan, filter, sort, summarize, generate, refine, classify, translate, convert, evaluate, interpret, integrate, conceptualize, extract, plan, link, aggregate, verify, group, compare, and cluster.
3. The method according to claim 1, characterized in that: The cost model is used to estimate the cost of each operator according to the size of input data and the cost function of the operator, wherein when estimating the cost of the operator, the cardinality estimate is used as the size of the input data, and the cardinality estimate approximates the selectivity through a random sampling method.
4. The method according to claim 1, characterized in that: The method further comprises: According to the sampled workload, the parameters of the cost function of the operators in the cost model are adjusted.
5. The method according to claim 1, characterized in that Real-time monitoring of intermediate results and execution status to dynamically and adaptively adjust the execution pipeline, including: Real-time monitoring of the intermediate results and current execution status generated during the execution process; Based on the intermediate results and the execution status, evaluating the execution quality of the currently executed pipeline; If the execution quality does not meet the expected quality requirement, stop executing the currently executed pipeline, regenerate at least one reference execution pipeline based on the collected intermediate results and using the large language model, and determine an adjusted execution pipeline from the at least one reference execution pipeline using the cost model.
6. The method according to claim 1, characterized in that During execution, a layer-by-layer execution strategy is adopted to execute the pipeline of the DAG structure.
7. The method according to claim 1, characterized in that The method also includes using the following pre-fetching technique during execution: When the operator corresponding to the search is executed, a primary search and multiple backup searches are generated, and different backup searches correspond to different search scopes directly or indirectly related to the query intention; The plurality of backup searches are performed in parallel while processing the results returned in response to the primary search, and the results returned in response to the backup searches are received and stored.
8. The method according to claim 1, characterized in that Integrate the outputs of various execution pipelines, including: After each execution pipeline is completed, the output of each execution pipeline is explained using a large language model; Analyze and merge the interpretation results of the outputs of each execution pipeline, and use the merged results as the integrated results.
9. The method according to claim 1, characterized in that: The method also includes using a reward model for feedback optimization: Evaluate the relevance and impact of operators based on their output and execution status during execution; Rewards are distributed based on the relevance and impact of the evaluation, and the reward information is used to fine-tune a large language model that extracts relevant operators from the query to enhance the chain thinking ability of the large language model.
10. An automated interactive large-scale language model pipeline orchestration system for complex queries, characterized in that: The system is used to implement the method according to any one of claims 1 to 9, and the system comprises: A query interface, used to receive a query input by a user in natural language and identify the user's query intent; An operator set module, used to store predefined operator sets; A pipeline generator, configured to extract operators from an operator set using a large language model, and construct a plurality of chained candidate pipelines with different reasoning paths corresponding to the query; a pipeline rewriter for rewriting the chained candidate pipeline into a directed acyclic graph (DAG) structure and identifying parallel opportunities during the rewriting process to write the corresponding operators into a parallel structure; An optimizer, for estimating the cost of each operator in the DAG structure using a cost model, and selecting at least one DAG candidate pipeline with the lowest overall cost as the optimal execution pipeline; A pipeline executor, configured to execute at least one selected optimal execution pipeline according to a layer-by-layer execution strategy, and monitor intermediate results and execution status in real time to dynamically and adaptively adjust the execution pipeline; The pipeline integrator is used to evaluate and integrate the output results of various execution pipelines; A context manager to manage intermediate results in execution to make the information coherent and avoid exceeding the context length limit for large language models; Indexing and storage modules for storing and indexing structured, semi-structured, and unstructured data in the data lake; and A reward model for distributing rewards based on the relevance and impact of operators in a given state, fine-tuning large language models to enhance chain thinking capabilities.
11. An electronic device, characterized in that: The device comprises a memory and a processor, wherein the memory is used to store computer instructions executable on the processor, and the processor is used to implement the method according to any one of claims 1 to 9 when executing the computer instructions.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Database query optimization method, medium and equipment
CN108874954A
Query condition analysis method and device and electronic equipment
CN114265960A
Systems and methods for decoupling search processing language and machine learning analytics from storage of accessed data
US11500871B1
Artificial intelligence sandbox for automating development of AI models
US12204565B1
System and Methodology for Parallel Query Optimization Using Semantic-Based Partitioning
US20060218123A1
Cited By
Method and device for analyzing unstructured data and storage medium
CN120578740A
Alarm information processing method, system and device and computer storage medium
CN121193586A